Method and apparatus for generating deep ensemble model, and computer device
By generating a deep ensemble model from the search space sampling of classic CNN models and a multi-objective optimization strategy based on proxy models, the problem of insufficient generalization ability of deep learning models in image processing tasks is solved, and better image classification performance is achieved.
Patent Information
- Application Number
- PCT/CN2024/101506
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2024-06-26
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art lacks generalization capabilities of deep learning models in image processing tasks, resulting in far inferior performance on real images than on training data sets.
By randomly sampling multiple CNN models from the search space of the classical convolutional neural network model for image classification, collecting representative image samples based on scenes of image classification tasks, building training data sets, training and evaluation of CNN models, and generating the structure of the deep ensemble model based on proxy models and multi-objective optimization strategies, including shared blocks.
The performance of the deep ensemble learning model is improved, and better generalization ability and classification accuracy are achieved on real images.
Smart Images

Figure CN2024101506_08052025_PF_FP_ABST
Abstract
Description
Method, device and computer equipment for generating deep integration model
[0001] This disclosure claims priority to a Chinese patent application filed with the State Intellectual Property Office of China on November 1, 2023, with application number 202311436678.X and invention name “Method, device and computer equipment for generating deep integration models”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to the field of image classification technology, and in particular to methods, devices, and computer equipment for generating deep integrated models for image classification. The methods, models, devices, and equipment disclosed herein can be used for image classification tasks in fields such as visual search, image labeling, content filtering, medical image analysis, security monitoring, agriculture, and environmental testing. Background Art
[0003] With the development of artificial intelligence (AI), its applications are becoming increasingly widespread, placing increasing demands on machine learning and deep learning. Deep learning and machine learning have achieved fruitful results in image processing. However, some current deep learning models still suffer from insufficient generalization capabilities in image processing tasks, resulting in performance on real images far inferior to that on training datasets.
[0004] Typically, a deep ensemble model is constructed from multiple heterogeneous neural networks (NNs). The neural architecture of a deep ensemble model is more complex than that of a single NN. Because the neural architecture is closely related to model performance, deep learning requires a well-defined neural architecture, specifically the schema of each layer and the connections between them. Currently, deep ensemble models can be designed manually or automatically.
[0005] However, manual design requires designers to have professional knowledge and rich experience, and the existing methods for automatically designing deep integration model structures have problems of low efficiency and poor diversity. Therefore, there is an urgent need for a method that can quickly generate deep integration models to improve the performance of deep integration learning models.
[0006] Summary of the Invention
[0007] To overcome the problems existing in the related art, the present disclosure provides a method, apparatus and computer device for generating a deep integration model for image classification.
[0008] According to a first aspect of an embodiment of the present disclosure, a method for generating a deep integration model is provided, the method comprising: randomly sampling a first group of CNN models from a search space of a classic CNN model for image classification; collecting representative image samples based on a scenario of an image classification task, annotating the collected image samples with classification labels, and constructing a training data set based on the annotated image samples; the image classification task comprises at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; training and evaluating the image classification accuracy of the sampled first group of CNN models based on the training data set; generating a second group of CNN models based on a proxy model and the trained first group of CNN models; constructing a structure of a deep integration model for image classification based on the second group of CNN models and a multi-objective optimization strategy, the structure of the deep integration model comprising a shared block.
[0009] Optionally, a second set of CNN models is generated based on the single-objective differential evolution operation, the proxy model, and the trained first set of CNN models.
[0010] Optionally, the first group of CNN models after training is encoded and clustered, and the cluster centers are extracted to construct a first parent population; a performance comparator is trained using the sampled first group of CNN models; when the termination condition of the training is not met, a first child population is generated using a single-objective differential evolution operation; the first parent population and the first child population are merged into a new first parent population; it is determined whether to use a performance comparator; when the performance comparator is used, the new first parent population is sorted using a merge sort method of a proxy model, and a next generation population is constructed through tournament selection; when the termination condition of the training is met, the last generation population is output, and the last generation population is the second group of CNN models.
[0011] Optionally, after determining whether to use the performance comparator, the method also includes: decoding and training the CNN model in the first child population without using the performance comparator; sorting the new first parent population according to accuracy; using the first child population to train the performance comparator; evaluating and sorting the performance of the CNN model of the new first parent population without using the trained performance comparator every T iterations, training the CNN model of the new first parent population using the training data set, testing the true accuracy of the CNN model of the new first parent population using the validation data set, sorting the candidate solutions according to the trained performance comparator and selecting high-quality candidate solutions; and using the neural architecture encoding and corresponding accuracy ranking of the CNN model of the new first parent population as training data to perform incremental training on the trained performance comparator; or, in other cases, using the trained performance comparator to evaluate and sort the performance of the CNN model of the new first parent population.
[0012] Optionally, a structure of a deep ensemble model for image classification is constructed based on a dual-objective differential evolution operation, a second set of CNN models, and a multi-objective optimization strategy.
[0013] Optionally, a second parent population is randomly constructed based on the second group of CNN models; if the termination condition is not met, a dual-objective differential evolution operation is used on the second group of CNN models to generate a second child population; the second child population is evaluated, and the second child population and the second parent population are merged into a new second parent population; a non-dominated solution of the current multi-objective optimization problem is selected from the new second parent population; wherein, for any non-dominated solution, there is no solution in the new second parent population that is simultaneously better than the non-dominated solution in all optimization objectives; an external archive is updated to save the non-dominated solutions; the next generation population is constructed through tournament selection; if the termination condition is met, the solution with the highest accuracy in the external archive is decoded; and the constructed CNN ensemble model is output, which is a deep ensemble model for image classification.
[0014] Optionally, each CNN model in the first group of trained CNN models and the error rate of the image classification corresponding to each CNN model are encoded into real vectors, and the encoded real vectors are clustered into multiple classes; wherein the error rate N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 if yes, otherwise returns 0.
[0015] Optionally, the search space of the CNN model includes 4 convolution blocks and 1 pooling layer, the structure of each convolution block is determined by the hyperparameters of the convolution block, and the hyperparameters of the convolution block include at least one of the following: convolution unit type, convolution layer channel expansion factor and number of convolution layer repetitions; the method also includes: encoding each CNN model into a fixed-length integer array according to the hyperparameters of each convolution block.
[0016] Optionally, each overall CNN model structure in the first group of CNN models after training is changed based on the first mutation operator; or a single convolution block in each CNN model in the first group of CNN models after training is changed based on the second mutation operator.
[0017] Optionally, the first mutation operator and the second mutation operator are different values of the target mutation operator; the target mutation operator is: Among them, j1 represents the index of the j1th CNN model in the population, represents the mutated intermediate obtained by the j1th CNN model in the population after mutation in the tth generation, represents the j1th CNN model in the population, and for A random neighbor CNN model randomly selected from the neighborhood, F represents the factor that controls the range of the variant CNN model, and r represents a real number between [0,1] randomly generated by the random number generator rand(0,1). A vector indicating the direction of change of the optimal solution and The optimal solutions corresponding to the t-1 and t-2 generations, t is the current iteration round.
[0018] Optionally, the original CNN model and the convolution blocks of the original CNN model are exchanged based on a first crossover operator; or random bits of the original CNN model are exchanged based on a second crossover operator; wherein the first crossover operator is: r i2 =randI(0,2);r i2 is the i2th integer randomly generated by the random integer generator randI(0,2) in the range [0,2], CR represents the preset crossover probability factor, and j2 represents the index of the j2th variable dimension; represents the result of the j2th first crossover operator; represents the j2th dimension of the mutation solution of the target mutation operator, represents the j2th dimension of the original candidate solution of the target mutation operator; the second crossover operator is: represents the result of the j3-th second crossover operator, represents the j3rd dimension of the mutation solution of the target mutation operator, Represents the j3-th dimension of the original candidate solution of the target mutation operator.
[0019] Optionally, a second set of CNN models is evaluated based on a first objective function and a second objective function to construct a deep ensemble model for image classification; wherein the first objective function is used to determine the accuracy of the deep ensemble model, and the second objective function is used to determine the diversity of the deep ensemble model; wherein the first objective function is: Where N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 when , otherwise returns 0; the second objective function is: represents the output of the i4th head of the deep ensemble model, represents the architecture of the kth dimension of the i4th head, represents the architecture of the kth dimension of the j4th head, express On the pth test sample s p The output of , d represents the dimension of the encoding of the head, represents the Euclidean distance between the i4th head and the j4th head, Indicates the number of test samples with different classification results.
[0020] Optionally, an external archive stores non-dominated solutions that satisfy a preset objective function; among the candidate solutions of the preset objective function, if a first candidate solution satisfies: No other candidate solution is superior to the first candidate solution in both the first and second objective functions, then the first candidate solution is a non-dominated solution, and the first candidate solution is any candidate solution; wherein the preset objective function is: represents the j5th candidate solution, W represents the set of all candidate solutions, represents the i5th candidate solution The kth dimension of represents the j5th candidate solution The kth dimension of .
[0021] Optionally, the first group of CNN models after training is encoded based on an integer group; wherein each element of the integer group is an index of the first group of CNN models after training, the length of the array is m, the first element of the integer group represents the index of the CNN model contributing the shared layer, and the second element to the mth element of the integer group represent the index of the CNN model for constructing the head architecture.
[0022] According to a second aspect of an embodiment of the present disclosure, a device for generating a deep integration model is provided, which includes: a sampling module, a dataset construction module, a training and evaluation module, a generation module and a model construction module; the sampling module is used to randomly sample the neural architecture of a first group of CNN models from the search space of a classic convolutional neural network (CNN) model for image classification; the dataset construction module is used to collect representative image samples based on the scene of the image classification task, annotate the classification labels of the collected image samples, and construct a training dataset based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; the training and evaluation module is used to train and evaluate the image classification performance of the sampled first group of CNN models based on the training dataset; the generation module is used to generate a second group of CNN models based on the proxy model and the trained first group of CNN models; the model construction module is used to construct the structure of a deep integration model for image classification based on the second group of CNN models and a multi-objective optimization strategy, and the structure of the deep integration model includes a shared block.
[0023] According to a third aspect of an embodiment of the present disclosure, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect above when executing the program.
[0024] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0025] In the disclosed embodiment, a first group of CNN models is first randomly sampled from the search space of the classic CNN model for image classification; then, representative image samples are collected based on the scene of the image classification task, the classification labels of the collected image samples are annotated, and a training data set is constructed based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; then the sampled first group of CNN models are trained and evaluated based on the training data set; then a second group of CNN models are generated based on the proxy model and the trained first group of CNN models; finally, based on the second group of CNN models and a multi-objective optimization strategy, a structure of a deep integration model for image classification is constructed, and the structure of the deep integration model includes a shared block. That is, in the first stage, a group of high-precision CNN models are first trained and generated, and then the high-precision CNN models are used to construct the deep integration model of the second stage. The first stage is a single-objective optimization process, so it has a relatively fast convergence speed, and the second stage is a combination with a small-scale search space, so computing resources can be saved.
[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0028] FIG1 is a flowchart of a method for generating a neural integration module according to one or more embodiments of the present disclosure;
[0029] FIG2 is a schematic diagram of the structure of a convolution unit according to one or more embodiments of the present disclosure;
[0030] FIG3 is a schematic diagram of encoding of a CNN model backbone architecture according to one or more embodiments of the present disclosure;
[0031] FIG4 is a flowchart of a method for generating a deep integration model according to one or more embodiments of the present disclosure;
[0032] FIG5 is a flowchart of a method for generating a deep integration model according to one or more embodiments of the present disclosure;
[0033] FIG6 is a schematic diagram of an execution logic of a method for generating a deep integration model according to one or more embodiments of the present disclosure;
[0034] FIG7 is a schematic diagram showing an error rate trend according to one or more embodiments of the present disclosure;
[0035] FIG8 is a schematic diagram showing a comparison of error rate and search time according to one or more embodiments of the present disclosure;
[0036] FIG9 is a schematic diagram showing a comparison of error rates according to one or more embodiments of the present disclosure;
[0037] FIG10 is a schematic diagram illustrating the effect of the number of basic classifiers on generalization capability according to one or more embodiments of the present disclosure;
[0038] FIG11 is a schematic diagram of a possible structure of a device for generating a deep integration model according to one or more embodiments of the present disclosure;
[0039] FIG12 is a hardware structure diagram of a computer device where the apparatus for generating a deep integration model according to one or more embodiments of the present disclosure is located. DETAILED DESCRIPTION
[0040] The following first introduces technical terms in the technology related to the present disclosure.
[0041] 1. Deep Ensemble Model
[0042] A deep ensemble model is a deep learning model constructed from multiple heterogeneous neural networks. Each element of a deep ensemble model is called a base model. The diversity among the base models ensures the robustness of the deep ensemble model, meaning that it maintains good performance when the input data type differs from the training data type.
[0043] Compared to neural networks with linear architectures, deep ensemble models have stronger robust generalization capabilities. The neural structure of a deep ensemble model is more complex than that of a single neural network, and an appropriate neural structure must be designed before the model is applied. This neural structure includes the specific patterns of each layer and the connection relationships between layers.
[0044] 2. CNN model
[0045] The CNN model is a deep learning method that has excellent performance in image processing problems. A large number of advanced image processing models are built based on the CNN model. Due to the parameter sharing of the convolution kernel and the sparse connections between the convolution layers, the CNN model can process grid data such as time series data, images, and audio at a lower computational cost.
[0046] 3. Generalization
[0047] Generalization refers to a machine learning algorithm's ability to adapt to new samples. The goal of learning is to learn the patterns underlying the data. This ability allows a trained network to produce appropriate outputs for data outside of the learning set that exhibits the same patterns.
[0048] 4. Image Classification
[0049] Image classification is a key application scenario for deep ensemble models. These models must be able to correctly classify images based on various features, including point features, local features, regional features, and overall features. Image classification has a wide range of applications in data storage and processing, social media, digital healthcare, agriculture, and environmental protection. It serves as a foundational technology for applications such as content filtering, disease diagnosis, defect detection, and disaster monitoring.
[0050] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0051] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0052] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."
[0053] Next, the embodiments of the present disclosure are described in detail.
[0054] For the sake of ease of description, the method for generating a neural integration module provided in the embodiments of the present disclosure is named DENE in the embodiments of the present disclosure.
[0055] As shown in FIG1 , FIG1 is a flowchart of a method for generating a neural integration module according to an exemplary embodiment of the present disclosure, comprising the following steps:
[0056] Step 100: Randomly sample a first set of CNN models from the search space of classic CNN models for image classification.
[0057] Among them, the search space of the CNN model includes 4 convolution blocks and 1 pooling layer; the structure of each convolution block is determined by the hyperparameters of the convolution block, and the hyperparameters of the convolution block include at least one of the following: convolution unit type, convolution layer channel expansion factor and number of convolution layer repetitions; furthermore, each CNN model can be encoded into a fixed-length integer array according to the hyperparameters of each convolution block.
[0058] Exemplarily, in the embodiment of the present disclosure, the architecture of the candidate CNN model in the search space is a backbone architecture based on a wide residual network (WRN), which includes: 4 convolution blocks and 1 pooling layer. The 4 convolution blocks are respectively recorded as conv1, conv2, conv3, and conv4, where the filter size of each convolution block is 3*3.
[0059] Table 1 is a schematic table of a search space provided by an embodiment of the present disclosure.
[0060] Table 1
[0061] In the disclosed embodiment, input data is processed in the following order: conv1, conv2, conv3, conv4, and pooling layers. Convolutional block 1 includes one convolutional layer with 16 channels. Each of convolutional blocks 2, 3, and 4 may include multiple repeated convolutional units, which may be either type a or type b. Type a is a base-width convolutional unit, while type b is an enhanced-width convolutional unit.
[0062] Figure 2 is a structural schematic diagram of a convolution unit provided by an embodiment of the present disclosure. As shown in Figure 2, the convolution unit (a) includes two convolution layers, each layer is a 3*3 convolution block, and the convolution unit (b) includes two convolution layers, each layer is a 3*3 convolution block, and a drop operation is included between the two convolution layers.
[0063] It should be noted that in the disclosed embodiments, a fixed-length integer encoding method can be used to encode the backbone architecture of the candidate CNN model. For example, a candidate CNN model can be represented by an integer array of length 9, where every 3 bits of the array represent the convolution unit type, convolution layer channel expansion factor, and number of convolution layer repetitions of the second to fourth convolution blocks.
[0064] Figure 3 is a coding diagram of a CNN model backbone architecture provided by an embodiment of the present disclosure. As shown in Figure 3, conv2 uses type a convolution unit, the convolution layer channel expansion factor k1 is 2, and the number of convolution layer repetitions is 2, then conv2 can be expressed as 0-2-2; conv3 uses type b convolution unit, the convolution layer channel expansion factor k1 is 4, and the number of convolution layer repetitions is 3, then conv3 can be expressed as 1-4-3; conv4 uses type b convolution unit, the convolution layer channel expansion factor k1 is 2, and the number of convolution layer repetitions is 2, then conv4 can be expressed as 1-2-2, and thus the CNN model illustrated in Figure 3 can be expressed as 0-2-2-1-4-3-1-2-2.
[0065] Step 101: Collect representative image samples based on the scene of the image classification task, annotate the classification labels of the collected image samples, and construct a training dataset based on the annotated image samples.
[0066] Optionally, the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection.
[0067] Step 102: Train and evaluate the image classification performance of the first sampled CNN model based on the training dataset.
[0068] Step 103: Generate a second group of CNN models based on the proxy model and the sampled first group of CNN models.
[0069] It should be noted that the disclosed embodiments provide a proxy model-based performance ranking strategy that can predict the performance ranking of candidate solutions without determining the accuracy of each solution. This allows for stable performance estimation based on the neural architecture, independent of the quality of the candidate solutions. Applied in the selection phase, the proxy model is used to estimate the performance of the candidate CNN model, rather than training all CNN models in the search space on the training set. Since the training portion is the most time-consuming process in the method, this can improve the efficiency of generating the neural ensemble.
[0070] Step 104: Based on the second group of CNN models and the multi-objective optimization strategy, a structure of a deep integration model for image classification is constructed.
[0071] Among them, the structure of the deep integration model includes shared blocks.
[0072] It should be noted that in the embodiment of the present disclosure, a sampling layer sharing strategy is adopted to reduce the number of parameters of the CNN model set. When constructing a deep integration model, three convolution blocks are fixed as shared layers. For a CNN model set with M heads, M different convolution blocks and pooling layers are selected from the fourth convolution block and pooling layer of the candidate CNN model.
[0073] The disclosed embodiment provides a method for generating a deep ensemble model, which first randomly samples a first group of CNN models from the search space of a classic CNN model for image classification; then, based on the scene of the image classification task, representative image samples are collected, the classification labels of the image samples are annotated, and a training data set is constructed based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; then, the sampled first group of CNN models are trained and evaluated based on the training data set; then, based on the proxy model and the trained first group of CNN models, a second group of CNN models are generated; finally, based on the second group of CNN models and a multi-objective optimization strategy, a structure of a deep ensemble model for image classification is constructed, and the structure of the deep ensemble model includes a shared block. That is, in the first stage, a group of high-precision CNN models are first trained and generated, and then the high-precision CNN models are used to construct the deep ensemble model of the second stage. The first stage is a single-objective optimization process, so it has a relatively fast convergence speed, and the second stage is a combination with a small-scale search space, so it can save computing resources.
[0074] Optionally, embodiments of the present disclosure provide a method for generating a deep ensemble model. In step 103, a second set of CNN models may be generated based on a single-objective differential evolution operation, a proxy model, and a trained first set of CNN models. In step 104, a structure of a deep ensemble model for image classification may be constructed based on a dual-objective differential evolution operation, the second set of CNN models, and a multi-objective optimization strategy.
[0075] Specifically, step 103 may include the following steps:
[0076] Step 01: Encode and cluster the first group of trained CNN models, extract the cluster centers to construct the first parent population, and use the sampled first group of CNN models to train the performance comparator.
[0077] Specifically, each CNN model in the first group of trained CNN models and the error rate corresponding to each CNN model are encoded as a real vector, and the encoded real vectors are clustered into multiple classes.
[0078] The error rate can be determined based on formula (1).
[0079] Where N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 if yes, otherwise returns 0.
[0080] In step 01, the first set of CNN models trained is encoded based on the integer group.
[0081] Each element of the integer array is the index of the first group of CNN models after training. The length of the array is m. The first element of the integer array represents the index of the CNN model that contributes to the shared layer. The second element to the mth element of the integer array represent the index of the CNN model that constructs the head architecture.
[0082] That is, in the embodiment of the present disclosure, a data-driven clustering-based CNN model initialization strategy is adopted, and multiple CNN models are first randomly sampled. Then, the multiple CNN models are trained on a training set and tested on a validation set. Then, the CNN models and the error rates corresponding to the CNN models are encoded as real vectors, and a clustering method is used to divide the encoding into multiple classes. The center of each cluster constitutes the initial CNN model for the next differential evolution operation.
[0083] Step 02: When the training termination condition is not met, use the single-objective differential evolution operation to generate the first offspring population; merge the first parent population and the first offspring population into a new first parent population.
[0084] Step 03: Determine whether to use a performance comparator.
[0085] Among them, the performance comparator is provided for the proxy model.
[0086] For example, a light GBM (Light Gradient Boosting Machine) can be selected as the proxy model, where light GBM is an efficient algorithm framework based on the gradient boosting decision tree (GBDT), so it has fast training speed, low memory consumption and high accuracy.
[0087] Exemplarily, it may be determined whether the current iteration satisfies a preset fixed number of iterations to determine whether to use the performance comparator.
[0088] It should be noted that if the performance comparator is used, the following step 04a is executed; if the performance comparator is not used, the following step 04b is executed, and the above step 03 is executed again.
[0089] Step 04a: Sort the new first-parent population using the merge sort method of the surrogate model using the performance comparator, and construct the next generation population through tournament selection.
[0090] It should be noted that based on the proxy model and using the merge sort method, a sorting method with a time complexity of O(nlog2n) and a space complexity of O(n) can be obtained, and this sorting method is stable.
[0091] Step 04b: Decode and train the CNN model in the first generation population without using the performance comparator; sort the new first parent population by accuracy; and train the performance comparator using the first child population.
[0092] Among them, every T iterations, the trained performance comparator is not used to evaluate and rank the performance of the CNN model of the new first parent population, the training dataset is used to train the CNN model of the new first parent population, the validation dataset is used to test the true accuracy of the CNN model of the new first parent population, the candidate solutions are ranked according to the trained performance comparator and high-quality candidate solutions are selected; and the neural architecture encoding and corresponding accuracy ranking of the CNN model of the new first parent population are used as training data to perform incremental training on the trained performance comparator.
[0093] Alternatively, in other cases, the performance of the CNN models of the new first parent population is evaluated and ranked using the trained performance comparator.
[0094] It should be noted that the above T iterations refer to the iterations of optimizing the neural architecture of the CNN model using the evolutionary optimization method described in step 02.
[0095] Exemplarily, the input of the proxy model is an integer string, half of which is the encoding of the first CNN model and the other half is the encoding of the second CNN model, and the output is a bool value indicating whether the first CNN model is better than the second CNN model.
[0096] For ease of explanation, the process of generating the first set of CNN models is called the CNN model initialization stage. In the disclosed embodiment, the candidate CNN models generated in the CNN model initialization stage are used to train the proxy model from scratch, and the proxy model is updated at regular intervals of generations, that is, the proxy model is trained using the newly generated candidate solutions.
[0097] Step 05: When the training stop condition is met, the last generation of population is output, which is the second group of CNN models mentioned above.
[0098] Specifically, in step 104, the following steps may be included:
[0099] Step 11: Randomly construct the second parent population based on the second group of CNN models.
[0100] Step 12: If the termination condition is not met, a dual-objective differential evolution operation is performed on the second set of CNN models to generate a second offspring population; the second offspring population is evaluated, and the second offspring population and the second parent population are merged into a new second parent population; a non-dominated solution to the current multi-objective optimization problem is selected from the new second parent population, and the external archive is updated to save the non-dominated solution; and the next generation population is constructed through tournament selection.
[0101] Exemplarily, the termination condition may be that the number of iterations reaches a set maximum number of iterations, the accuracy reaches a set maximum accuracy, etc., which is not specifically limited in the embodiments of the present disclosure.
[0102] Among them, for any non-dominated solution, there is no solution in the new second parent population that is better than the non-dominated solution in all optimization objectives.
[0103] Step 13: If the termination condition is met, decode the solution with the highest accuracy in the external archive.
[0104] Step 14: Output the constructed CNN ensemble model.
[0105] The CNN ensemble model is a deep ensemble model for image classification.
[0106] Optionally, in a method for generating a deep ensemble model provided by an embodiment of the present disclosure, two modified mutation operators and two modified crossover operators may be used when performing a differential operation. The operators can balance the exploration and development of the method, so that complex neuron ensemble architecture problems can be handled.
[0107] Furthermore, the above step 02 may include the following steps:
[0108] Step 21: Change the structure of each overall CNN model in the first group of trained CNN models based on the first mutation operator.
[0109] It should be noted that the first mutation operator is more inclined to global search.
[0110] The first mutation operator can be described by formula (2).
[0111] Among them, j1 represents the index of the j1th CNN model in the population, represents the mutated intermediate obtained by the j1th CNN model in the population after mutation in the tth generation, represents the j1th CNN model in the population, and for A random neighbor CNN model randomly selected from the neighborhood, F represents the factor that controls the range of the variant CNN model, and r represents a real number between [0,1] randomly generated by the random number generator rand(0,1). A vector indicating the direction of change of the optimal solution and The optimal solutions corresponding to the t-1 and t-2 generations, t is the current iteration round.
[0112] Step 22: Change a single convolutional block in each CNN model in the trained first group of CNN models based on the second mutation operator.
[0113] Specifically, we randomly select a bit of a CNN model and change the bit to a feasible value.
[0114] The first mutation operator and the second mutation operator are different values of the target mutation operator.
[0115] Step 23: Exchange the original CNN model and the convolution blocks of the original CNN model based on the first cross operator.
[0116] Step 24: Swap random bits of the original CNN model based on the second crossover operator.
[0117] The first crossover operator is represented by the following formula (3).
[0118] Among them, r i2 =randI(0,2);r i2 is the i2th integer randomly generated by the random integer generator randI(0,2) in the range [0,2], CR represents the preset crossover probability factor, and j2 represents the index of the j2th variable dimension; represents the result of the j2th first crossover operator; represents the j2th dimension of the mutation solution of the target mutation operator, Represents the j2-th dimension of the original candidate solution of the target mutation operator.
[0119] It should be noted that in the embodiment of the present disclosure, CR is a parameter that controls the crossover strategy. If the random number is less than CR, the crossover result is the first term of the formula. If the random number is greater than or equal to CR, the crossover result is the second term. That is, CR is used for exploration and search of the balance method.
[0120] The second crossover operator can be expressed by the following formula (4):
[0121] in, represents the result of the j3-th second crossover operator, represents the j3rd dimension of the mutation solution of the target mutation operator, Represents the j3-th dimension of the original candidate solution of the target mutation operator.
[0122] Optionally, in the method for generating a deep integration model provided in the embodiment of the present disclosure, the above-mentioned step 104 may include the following steps:
[0123] Step 31: Evaluate the second set of CNN models based on the first objective function and the second objective function to construct a deep integration model for image classification.
[0124] It can be understood that in the embodiment of the present disclosure, the task of constructing a deep integration model is modeled as a dual-objective optimization problem, wherein the first objective function is used to determine the accuracy of the deep integration model, and the second objective function is used to determine the diversity of the deep integration model.
[0125] The first objective function can be expressed by the following formula (5).
[0126] Among them, N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 if yes, otherwise returns 0.
[0127] The second objective function can be expressed by the following formula (6).
[0128] in, represents the output of the i4th head of the deep ensemble model, represents the architecture of the kth dimension of the i4th head, represents the architecture of the kth dimension of the j4th head, express On the pth test sample s p The output of , d represents the dimension of the encoding of the head, represents the Euclidean distance between the i4th head and the j4th head, Indicates the number of test samples with different classification results.
[0129] In an embodiment of the present disclosure, multi-objective optimization can be performed, wherein, for each CNN model, a CNN model is randomly selected from the second group of CNN models as a shared layer, and then a neighborhood-based DE and binomial crossover operator are used to generate an offspring population, wherein an external archive stores non-dominated solutions that satisfy a preset objective function; if a first candidate solution among the candidate solutions of the preset objective function satisfies: there is no other candidate solution that is better than the first candidate solution in both the first objective function and the second objective function optimization objectives, then the first candidate solution is a non-dominated solution, and the first candidate solution is any candidate solution.
[0130] The preset objective function (external archive) satisfies the non-dominated solution of the following formula (7).
[0131] in, represents the j5th candidate solution, W represents the set of all candidate solutions, represents the i5th candidate solution The kth dimension of represents the j5th candidate solution The kth dimension of .
[0132] It can be understood that for each iteration, the next parent population is constructed using non-dominated sorting based on crowding, and finally the one with the highest classification accuracy in the external archive is selected to be decoded into the CNN model set, and the CNN model set is output, which indicates the CNN model of the deep integration model.
[0133] For example, the method for generating a deep integration model provided by an embodiment of the present disclosure includes two main stages, wherein the first stage is an initialization stage, which uses a data-driven clustering-based CNN model initialization strategy. In conjunction with FIG4 , the following steps may be included:
[0134] Step 401: Randomly sample from the search space to obtain a first set of CNN models.
[0135] Step 402: Train and evaluate the first set of CNN models.
[0136] Step 403: Encode and cluster the trained first group of CNN models.
[0137] Step 404: extract the center point CNN model of each type of clustered CNN model.
[0138] Step 405: Use the center points of various CNN models to construct a parent CNN population P.
[0139] Step 406: Use the sampled first set of CNN models to train a performance comparator.
[0140] Step 407: Determine whether the termination condition is met.
[0141] Step 408: If the termination condition is not met, generate a child population Q based on the parent population P and the DE operator.
[0142] Step 409: Merge the offspring population Q and the parent population P into a new parent population P.
[0143] Step 410: Determine whether to use a performance comparator.
[0144] Step 411: If a performance comparator is used, the new parent population is sorted.
[0145] Step 412: Use tournament selection to construct P t+1 .
[0146] Step 413: If the termination condition is met, output the last generation.
[0147] Step 414: If the performance comparator is not used, decode and train the CNN model in the offspring population Q.
[0148] Step 415: Sort the CNN models in the parent population P based on accuracy.
[0149] Step 416: Use the offspring population Q to train the performance comparator.
[0150] After the training is completed, the process returns to step 410 to determine whether to use the trained performance comparator.
[0151] In the deep integration model generation method provided by the embodiment of the present disclosure, multiple CNN models are first randomly sampled, and then the multiple CNN models are trained using a training data set. Then, the trained multiple CNN models are tested using a verification data set. Then, the verified multiple CNN models and their corresponding error rates are encoded into real vectors. Then, a clustering method is used to divide the encoded vectors into N classes, and the center of each class in the N classes constitutes the initial CNN model of the next DE.
[0152] It should be noted that the above population initialization method can find promising candidate region solutions, thereby improving the efficiency of initialization. By clustering the encoded vectors, that is, clustering CNNs, the diversity of the parent population can be maintained and the problem of local optimality can be avoided.
[0153] It should be noted that in the method for generating a deep integration model provided in the embodiment of the present disclosure, the population size is denoted as N, and the time complexity of population generation is denoted as O(N). When GBDT is selected as the proxy model, the time complexity of training the model is O(2KN log N), K is the number of trees, and the time complexity of performance evaluation is O(N log N); the time complexity of generating the second group of CNN models is O(2TKN log N), T is the number of iterations, and the time complexity of generating a deep integration model based on the second group of CNN models is O(TN log N), so the total time complexity is O(2TKN log N).
[0154] The second stage generates a deep ensemble model based on a set of CNN models selected in the first stage. Combined with Figure 5, it includes the following steps:
[0155] Step 501: randomly construct an initialized population H according to uniform distribution.
[0156] In the embodiment of the present disclosure, the parent population may also be referred to as the initialization population.
[0157] Step 502: Generate a descendant population W based on the initialized population H and the DE operator.
[0158] Step 503: Use the objective function to evaluate the offspring population W.
[0159] Step 504: The initialization population H and the offspring population W are merged into a new initialization population H.
[0160] Step 505: Determine the non-dominated solutions in the new initialized population H.
[0161] Step 506: Update the external set using non-dominated solutions based on Pareto dominance.
[0162] Step 507: Construct the next generation population H based on the new initial population H through the tournament selection method. t+1 .
[0163] Step 508: Determine whether the termination condition is met.
[0164] Step 509: When the termination condition is met, decode the solution with the highest accuracy in the external set.
[0165] Step 510: Output the deep integration model.
[0166] When the termination condition is not met, based on the next generation population H t+1 Re-execute the above step 502.
[0167] The following is a code example for implementing the method for generating a deep integration model provided by an embodiment of the present disclosure, wherein the first part is the code for generating a second group of CNN models, and the second part is the code for constructing a deep integration model based on the second group of CNN models.
[0168] Part 1:
[0169] Part II:
[0170] Figures 4 to 6 are schematic diagrams of the execution logic of the method according to the above method embodiment, including initialization, step 1 and step 2.
[0171] The following is an example of a deep integration model for an image classification task generated based on an embodiment of the present disclosure:
[0172] CIFAR-10 and CIFAR-100 are used as benchmark datasets.
[0173] CIFAR-10 includes 60,000 images in 10 categories, and the pixels of each image are 32*32; among them, the training set includes 48,000 image samples, the test set includes 10,000 image samples, and the validation set includes 2,000 image samples.
[0174] CIFAR-100 includes 100 categories, each category includes 600 image samples; among them, the training set includes 450 image samples in each category, the test set includes 100 image samples in each category, and the validation set includes 50 image samples in each category.
[0175] The experimental verification can be carried out based on PyTorch 1.6.0 on a workstation equipped with an Intel(R) Core(TM) i7-9700K CPU, an NVIDIA RTX2080 GPU, and 32GB of memory.
[0176] Table 2 is an exemplary table of parameter configurations of the DE differential evolution operator in an embodiment of the present disclosure.
[0177] Table 2
[0178] The above parameter values are set based on previous experience and typical research by researchers in this field. Due to the complexity of the architecture search problem, the crossover rate is greater than 0.5 to obtain stronger global search capabilities.
[0179] Table 3 shows the clustering method and parameter settings for CNN model training provided by the present disclosure.
[0180] Table 3
[0181] It should be noted that the embodiment of the present disclosure provides an example illustration of the impact of the number of heads on the method. When the number of heads m is set to 3, 5 and 10, the method for generating the deep integration model provided by the embodiment of the present disclosure is executed 10 times, and the execution results are shown in Table 4, where Table 4 shows the impact of the number of heads on the classification error rate and the number of parameters of the CNN set. It can be determined from the table that the parameters increase with the increase of the number of heads, and the error rate is negatively correlated with the number of heads.
[0182] Table 4
[0183] As can be seen from Table 4, the error rate is the highest when the number of heads is 3, and the number of parameters when the number of heads is 5 is more than twice that of when the number of heads is 3, but the error rates are close in these two cases.
[0184] Figure 7 is a schematic diagram of the error rate trend provided by an embodiment of the present disclosure, in which the line represents the changing trend of the error rate, and the filled area represents the distribution range of the experimental results. On CIFAR-10 and CIFAR-100, when m = 5 and 10, the experimental results are more stable and concentrated than when m = 3. When m = 10, the error rate is lower than the error rate when m = 5. Compared with the case of m = 3, this difference is not obvious. Considering the huge number of parameters when m = 10, in subsequent experiments, m is set to 5 unless otherwise specified.
[0185] This experiment compares a hierarchical shared CNN ensemble model with an ensemble model constructed directly from candidate CNN models. Table 5 is an exemplary table of hierarchical sharing effects provided by an embodiment of the present disclosure.
[0186] Table 5
[0187] Among them, Table 5 shows the effect of layer sharing. The parameter mean values of each model are listed. For the error rate, the numbers in brackets represent the mean and the best results. This disclosure uses the Wilcoxon signed rank test to study the significant differences between different methods. The last row of Table 6 shows the results of the Wilcoxon test. The confidence level is set to 0.05 in this experiment. When the p-value is lower than 0.05, it is considered that there is a significant difference in the result distribution of different methods. Without the layer-by-layer sharing strategy, the number of parameters is 6.96 times larger than the layer-by-layer sharing CNN model set on CIFAR-10 and 4.84 times larger than the layer-by-layer sharing on CIFAR-100. At the same time, using layer-by-layer sharing on CIFAR-10 and CIFAR-100 can obtain significantly better error rates. Therefore, the adoption of the layer-by-layer sharing strategy can improve the performance of DENE in terms of classification accuracy and efficiency.
[0188] In the embodiment of the present disclosure, a data-driven clustering-based CNN model initialization strategy is proposed. Table 6 is an exemplary table of the impact of CNN model initialization provided by the embodiment of the present disclosure.
[0189] Table 6
[0190] Table 6 compares the error rates of the CNN ensemble model generated by random initialization and DENE. The experimental results show that when using the proposed clustering-based CNN model strategy, the method provided by the embodiment of the present disclosure achieves better average and best error rates.
[0191] Table 7 is an exemplary table of the impact of a performance ranking strategy provided by an embodiment of the present disclosure.
[0192] Table 7
[0193] Table 7 shows the effectiveness of the proxy model-based performance ranking strategy, comparing search time and error rate. The performance ranking strategy can save 13.70% of model search time on CIFAR-10 and 18.10% on CIFAR-100. In addition, on CIFAR-10 and CIFAR-100, the Wilcoxon test shows that there is no significant difference in error rate between the proposed method and the method in which all candidate CNN models are fully trained. Therefore, the performance ranking strategy can save model search time without affecting the performance of the ensemble model.
[0194] Table 8 is an exemplary table of classification accuracy comparison provided by an embodiment of the present disclosure.
[0195] Table 8
[0196] Table 8 compares the error rate and number of parameters of the present disclosure with the most advanced evolutionary NAS algorithm and neuron set search algorithm. In each cell of the fifth and sixth columns, the numbers before and after the brackets represent the average error rate and the best error rate in repeated independent runs, respectively. The numbers in the curly brackets are the p-values of the Wilcoxon signed rank test. The best results in the comparison are highlighted in bold. Some algorithms only report one error rate result. This study regards this result as the best error rate. The symbol "-" indicates that the indicator is not reported in the relevant literature. The proposed algorithm is compared with six single neural structure algorithms (NSGANet, DeepMaker, EPSOCNN, EEEA-Net, SaMuNet and MFENAS) and four NEAS algorithms (DeepEns, NES-RS, HyperDeepEns and MH-NES). Since the results of all algorithms were not obtained in this study, some algorithms were not included in the Wilcoxon test comparison. Table 8 shows that the embodiment of the present disclosure generates a small model on CIFAR-10 and a relatively large model on CIFAR-100. The ensemble model generated by the disclosed embodiment is larger than DeepEns, NES-RS, HyperDeepEns, and MH-NES. Accordingly, the classification error rate of the disclosed embodiment on CIFAR-10 and CIFAR-100 is lower than that of other ensemble models. Compared with the ENAS algorithm, the disclosed embodiment has the best average error rate and the lowest error rate on CIFAR-10, and the lowest average error rate on CIFAR-100. According to the p-value, the disclosed embodiment significantly outperforms the compared NEAS algorithm on CIFAR-10 and CIFAR-100. In addition, the error rate of the disclosed embodiment is better than the three ENAS methods. Although the error rate of NSGANet is close to that of the disclosed embodiment, and the best error rate of EEEA-Net is the lowest among all algorithms, the lack of independent replication experimental statistics makes it difficult to evaluate the performance and stability of the algorithms. Therefore, the disclosed embodiments generally have competitive or better classification accuracy.
[0197] Table 9 is an exemplary table for comparing search times provided by an embodiment of the present disclosure.
[0198] Table 9
[0199] Table 9 compares the execution time of the embodiment of the present disclosure with DeepEns, NES-RS, and MH-NES. All algorithms are implemented on the same device. The present disclosure ran each algorithm ten times independently. The table shows the average results of the search time. The embodiment of the present disclosure consumed 17.0 GPU hours to generate the CNN ensemble model on CIFAR-10 and 18.1 hours on CIFAR-100. DENE's time consumption is lower than MH-NES, but higher than DeepEns and NES-RS. According to Table 9, the number of model parameters generated by the embodiment of the present disclosure is greater than that of DeepEns and NES-RS.
[0200] The experimental results indicate that the method for generating a deep integration model provided by the embodiments of the present disclosure has competitive performance compared to the currently advanced evolutionary NAS algorithm and NES algorithm in automatically constructing a deep integration model for image classification tasks.
[0201] The clustering-based processing method in the embodiment of the present disclosure can improve the classification accuracy of the CNN integration model. Through the multi-head architecture and shared layer processing method, computing resources can be saved and the number of parameters of the generated deep integration model can be reduced; the performance ranking method based on the proxy model can save the search time of the deep integration model. Figure 8 is a comparative diagram of the error rate and search time provided by the embodiment of the present disclosure, which specifically compares the error rate and model search time of the method proposed in the embodiment of the present disclosure with the typical algorithm. The size of each point represents the size of the generated deep model. The closer the coordinate axis is, the better the processing performance. As shown in Figure 8, the embodiment of the present disclosure achieves a good balance between classification accuracy and execution time. NSGANet, DeepMaker and SaMuNet are not included in the figure because their execution time is much longer than other algorithms. However, the embodiment of the present disclosure does not have an advantage in the number of parameters, which limits the application scenarios of the model.
[0202] Images in CIFAR-100 are divided into twenty superclasses, each containing five image categories. This paper uses three categories from each superclass as the training set, and the other two categories as the test set. Therefore, the training and test datasets share some common features, but also some differences. In this case, the present disclosure can use CIFAR-100 to study the performance of ensemble classifiers in the presence of data variation, that is, when the observed data distribution differs from the training data, and use the performance to reflect the generalization ability of the model.
[0203] Figure 9 shows a comparative error rate diagram of DENE, comparing it with DeepMaker, NSGANet, DeepEns, NES-RS, and MH-NES. The box corresponding to DENE is narrower, indicating that the experimental results generated by DENE are more concentrated. Furthermore, the box corresponding to DENE is positioned lower than the other boxes, indicating that the CNN model set generated by DENE has a lower error rate. Therefore, DENE has superior generalization capabilities compared to the other tested algorithms.
[0204] Figure 10 illustrates the effect of the number of base classifiers on generalization ability. This publication compares the error rates when m equals 5 and 10. As shown in Figure 9, when m = 10, the boxes of DENE and DeepEns are narrower than when m = 5. Furthermore, when m = 10, the error rates of all algorithms are lower. Therefore, increasing the number of base classifiers helps improve generalization ability. Of all the algorithms tested, DENE has the lowest and most stable classification error rate on offset data.
[0205] Experimental results show that DENE can generate an ensemble of CNN models with competitive or better classification accuracy within a single GPU day. Furthermore, DENE achieves stable results on offset data. Therefore, DENE can generate an ensemble of CNN models with high accuracy and strong generalization capabilities.
[0206] Ablation experiments show that the layer sharing strategy reduces the number of parameters in the neural ensemble. The proposed performance ranking strategy reduces the execution time of DENE without significantly affecting the classification accuracy of the CNN model ensemble. Furthermore, the proposed population initialization strategy significantly improves classification accuracy on CIFAR-10 and CIFAR-100. Therefore, building on the DE framework, by improving the key operations and strategies of traditional DE, DENE improves the efficiency and performance of the original DE on the NEAS problem.
[0207] The DENE provided in the embodiments of the present disclosure uses two stages to search for the neural structure of the CNN model set. When a single search process is used, shared layers and head structures are searched simultaneously. At the same time, in order to maintain the diversity of the neural set, the search process becomes complicated in this case. DENE uses the first stage to generate a set of CNN models with high precision for constructing the neural set of the second stage. Since the first stage is a single-objective optimization process, the algorithm has a relatively fast convergence speed. The second stage uses a multi-objective optimization process to balance the accuracy and diversity of the CNN model set. Since the second stage is a combinatorial problem with a small-scale search space, the weight parameters of the candidate CNN models can be reused, and the multi-objective optimization does not consume too much time and computing resources.
[0208] DENE fixes the number of shared blocks, a trade-off between performance and algorithmic efficiency. Introducing more layer-sharing patterns would increase the dimensionality of the search space, leading to an exponential increase in search time. However, a fixed number of layers limits the diversity of the neural ensemble and the flexibility of the algorithm. Designing efficient methods with variable shared layers is a promising direction for future research.
[0209] This paper proposes DENE, a differential evolution algorithm for neural ensemble architecture search. DENE automatically constructs neural ensembles for image classification tasks using two phases. Experimental results show that the proposed algorithm achieves competitive performance with state-of-the-art evolutionary NAS and NEAS algorithms on CIFAR-10 and CIFAR-100. The adopted multi-head architecture with shared layers can save computational resources and reduce the number of parameters. The proposed performance ranking strategy based on surrogate models saves DENE's search time. In addition, a clustering-based initialization strategy can improve the classification accuracy of CNN ensembles. As for future research, improving the diversity of head structures and designing more flexible layer sharing strategies can further improve the classification performance of CNN ensembles.
[0210] FIG11 is a block diagram of a device for generating a deep integration model according to an exemplary embodiment of the present disclosure, wherein the device 1000 includes: a sampling module 1001, a data set construction module 1002, a training evaluation module 1003, a generation module 1004, and a model construction module 1005; the sampling module 1001 is used to randomly sample the neural architecture of a first group of CNN models from the search space of a classic CNN model for image classification; the data set construction module 1002 is used to collect representative image samples based on the scene of the image classification task, annotate the classification labels of the collected image samples, and construct a training image based on the annotated image samples. training dataset; image classification tasks include at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; a training evaluation module 1003, used to train and evaluate the image classification performance of a sampled first group of CNN models based on the training dataset; a generation module 1004, used to generate a second group of CNN models based on the proxy model and the trained first group of CNN models; a model construction module 1005, used to construct a structure of a deep integration model for image classification based on the second group of CNN models and a multi-objective optimization strategy, the structure of the deep integration model including a shared block.
[0211] Optionally, the generation module is specifically used to generate a second group of CNN models based on the single-objective differential evolution operation, the proxy model and the trained first group of CNN models.
[0212] Optionally, the generation module is specifically used to: encode and cluster the first group of CNN models after training, extract cluster centers to construct a first parent population; use the sampled first group of CNN models to train a performance comparator; when the termination condition of the training is not met, use a single-objective differential evolution operation to generate a first child population; merge the first parent population and the first child population into a new first parent population; determine whether to use a performance comparator; when the performance comparator is used, use the merge sort method of the proxy model to sort the new first parent population, and construct the next generation population through tournament selection; when the termination condition of the training is met, output the last generation population, and the last generation population is the second group of CNN models.
[0213] Optionally, the generation module is also used to, after determining whether to use the performance comparator, decode and train the CNN model in the first child population without using the performance comparator; sort the new first parent population according to accuracy; and use the first child population to train the performance comparator; wherein, every T iterations, the performance of the CNN model of the new first parent population is evaluated and ranked without using the trained performance comparator, the CNN model of the new first parent population is trained using the training data set, the true accuracy of the CNN model of the new first parent population is tested using the validation data set, and the candidate solutions are ranked according to the trained performance comparator and high-quality candidate solutions are selected; and the neural architecture encoding and corresponding accuracy ranking of the CNN model of the new first parent population are used as training data to perform incremental training on the trained performance comparator; or, in other cases, the performance of the CNN model of the new first parent population is evaluated and ranked using the trained performance comparator.
[0214] Optionally, the model building module is specifically used to: construct the structure of a deep integration model for image classification based on a dual-objective differential evolution operation, a second group of CNN models and a multi-objective optimization strategy.
[0215] Optionally, the model construction module is specifically used to: randomly construct a second parent population based on the second group of CNNs; if the termination condition is not met, use a dual-objective differential evolution operation on the second group of CNN models to generate a second child population; evaluate the second child population, and merge the second child population and the second parent population into a new second parent population; select a non-dominated solution to the current multi-objective optimization problem from the new second parent population; wherein, for any non-dominated solution, there is no solution in the new second parent population that is simultaneously better than the non-dominated solution in all optimization objectives; update the external archive to save the non-dominated solutions; construct the next generation population through tournament selection; if the termination condition is met, decode the solution with the highest accuracy in the external archive; output the constructed CNN integration model, which is a deep integration model for image classification.
[0216] Optionally, the generation module is specifically used to: encode each CNN model and the error rate corresponding to each CNN model in the first group of trained CNN models into a real vector, and cluster the encoded real vectors into multiple classes; wherein the error rate N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 if yes, otherwise returns 0.
[0217] Optionally, the search space of the CNN model includes 4 convolution blocks and 1 pooling layer; the structure of each convolution block is determined by the hyperparameters of the convolution block, and the hyperparameters of the convolution block include at least one of the following: convolution unit type, convolution layer channel expansion factor and number of convolution layer repetitions; the generation module is also used to encode each CNN model into a fixed-length integer array according to the hyperparameters of each convolution block.
[0218] Optionally, the generation module is specifically used to: change each overall CNN model structure in the first group of CNN models after training based on the first mutation operator; or, change a single convolution block in each CNN model in the first group of CNN models after training based on the second mutation operator.
[0219] Optionally, the first mutation operator and the second mutation operator are different values of the target mutation operator; the target mutation operator is: Among them, j1 represents the index of the j1th CNN model in the population, represents the mutated intermediate obtained by the j1th CNN model in the population after mutation in the tth generation, represents the j1th CNN model in the population, and for A random neighbor CNN model randomly selected from the neighborhood, F represents the factor that controls the range of the variant CNN model, and r represents a real number between [0,1] randomly generated by the random number generator rand(0,1). A vector indicating the direction of change of the optimal solution and The optimal solutions corresponding to the t-1 and t-2 generations, t is the current iteration round.
[0220] Optionally, the generation module is specifically configured to: exchange the original CNN model and the convolution block of the original CNN model based on a first cross operator; or exchange random bits of the original CNN model based on a second cross operator;
[0221] Among them, the first crossover operator is: r i2 =randI(0,2); r i2 is the i2th integer randomly generated by the random integer generator randI(0,2) in the range [0,2], CR represents the preset crossover probability factor, and j2 represents the index of the j2th variable dimension; represents the result of the j2th first crossover operator; represents the j2th dimension of the mutation solution of the target mutation operator, represents the j2th dimension of the original candidate solution of the target mutation operator; the second crossover operator is: represents the result of the j3-th second crossover operator, represents the j3rd dimension of the mutation solution of the target mutation operator, Represents the j3rd dimension of the original candidate solution of the target mutation operator
[0222] Optionally, the model construction module is specifically used to: evaluate the second group of CNN models based on a first objective function and a second objective function to construct a deep integration model for image classification; wherein the first objective function is used to determine the accuracy of the deep integration model, and the second objective function is used to determine the diversity of the deep integration model; wherein the first objective function is: Where N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 when , otherwise returns 0; the second objective function is: represents the output of the i4th head of the deep ensemble model, represents the architecture of the kth dimension of the i4th head, represents the architecture of the kth dimension of the j4th head, express On the pth test sample s p The output of , d represents the dimension of the encoding of the head, represents the Euclidean distance between the i4th head and the j4th head, Indicates the number of test samples with different classification results.
[0223] Optionally, the external archive stores non-dominated solutions that satisfy a preset objective function; if a first candidate solution among the candidate solutions of the preset objective function satisfies: no other candidate solution is superior to the first candidate solution in both optimization objectives of the first objective function and the second objective function, then the first candidate solution is a non-dominated solution, and the first candidate solution is any candidate solution; wherein the preset objective function is: represents the j5th candidate solution, W represents the set of candidate solutions, represents the i5th candidate solution The kth dimension of represents the j5th candidate solution The kth dimension of .
[0224] The external archive stores non-dominated solutions that satisfy the preset formula. The solutions stored in the external archive are copies of the solutions in the population that meet the requirements. The quality of the stored CNN model does not degrade due to the iteration of the population.
[0225] Optionally, the generation module is specifically used to encode the first group of CNN models after training based on the integer group; wherein each element of the integer group is the index of the first group of CNN models after training, the length of the array is m, the first element of the integer group represents the index of the CNN model contributing the shared layer, and the second element to the mth element of the integer group represent the index of the CNN model for constructing the head architecture.
[0226] It should be noted that the neural integration generation device provided in the embodiment of the present disclosure can implement each process in the above method embodiment and can achieve the corresponding technical effect, and the embodiment of the present disclosure will not be repeated here.
[0227] Corresponding to the aforementioned method embodiments, the present disclosure also provides embodiments of a device and a terminal to which the device is applied.
[0228] The embodiment of the generation device of the deep integration model disclosed in the present invention can be applied to a computer device, such as a server or a terminal device. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the file processing in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, as shown in Figure 12, it is a hardware structure diagram of the computer device where the file processing device of the embodiment of the present invention is located. In addition to the processor 1110, memory 1130, network interface 1120, and non-volatile memory 1140 shown in Figure 12, the server or electronic device where the neural integration generation device 1131 is located in the embodiment can generally include other hardware according to the actual function of the computer device, which will not be described in detail.
[0229] Accordingly, the present disclosure also provides a deep integration model generation device, which includes a processor; a memory for storing processor executable instructions; wherein the processor is configured to: randomly sample a first group of CNN models from the search space of classic CNN models for image classification; collect representative image samples based on the scene of the image classification task, annotate the classification labels of the image samples, and construct a training data set based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; train and evaluate the sampled first group of CNN models based on the training data set; generate a second group of CNN models based on the proxy model and the trained first group of CNN models; construct the structure of a deep integration model for image classification based on the second group of CNN models and a multi-objective optimization strategy, and the structure of the deep integration model includes a shared block.
[0230] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0231] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the disclosed solution. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0232] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0233] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the inventions claimed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0234] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
[0235] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A method for generating a deep integration model, characterized in that: The method comprises: Randomly sample the neural architectures of the first set of CNN models from the search space of classic convolutional neural network (CNN) models for image classification; Collect representative image samples based on the scene of the image classification task, annotate the classification labels of the collected image samples, and build a training data set based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; Training and evaluating the image classification performance of a sampled first set of CNN models based on the training dataset; Generate a second set of CNN models based on the proxy model and the trained first set of CNN models; A second parent population is randomly constructed based on the second group of CNN models; if the termination condition is not met, a second child population is generated by using a dual-objective differential evolution operation on the second group of CNN models; the second child population is evaluated, and the second child population and the second parent population are merged into a new second parent population; a non-dominated solution of the current multi-objective optimization problem is selected from the new second parent population; wherein, for any non-dominated solution, there is no solution in the new second parent population that is superior to the non-dominated solution in all optimization objectives; an external archive is updated to save the non-dominated solution; a next generation population is constructed through tournament selection; if the termination condition is met, the solution with the highest accuracy in the external archive is decoded; and the constructed CNN integration model is output, wherein the CNN integration model is a deep integration model for image classification, and the structure of the deep integration model includes a shared block.
2. The method according to claim 1, characterized in that The generating a second group of CNN models based on the proxy model and the trained first group of CNN models includes: Based on the single-objective differential evolution operation, the proxy model and the trained first group of CNN models, a second group of CNN models is generated.
3. The method according to claim 2, characterized in that Based on the single-objective differential evolution operation, the proxy model and the trained first group of CNN models, a second group of CNN models is generated, including: Encode and cluster the first group of trained CNN models, extract cluster centers to construct the first parent population; use the sampled first group of CNN models to train a performance comparator; When the termination condition of the training is not met, a first offspring population is generated using a single-objective differential evolution operation; the first parent population and the first offspring population are merged into a new first parent population; determining whether to use the performance comparator; Under the condition of using the performance comparator, sorting the new first parent population using the merge sort method of the agent model, and constructing the next generation population through tournament selection; When the termination condition of the training is met, the last generation of population is output, and the last generation of population is the second group of CNN models.
4. The method according to claim 3, characterized in that After determining whether to use the performance comparator, the method further includes: Decoding and training the CNN model in the first child population without using a performance comparator; sorting the new first parent population by accuracy; training the performance comparator using the first offspring population; Wherein, at every T iterations, the training data set is used to train the CNN model of the new first parent population, the verification data set is used to test the true accuracy of the CNN model of the new first parent population, the candidate solutions are sorted according to the trained performance comparator and high-quality candidate solutions are selected; and the neural architecture encoding and corresponding accuracy ranking of the CNN model of the new first parent population are used as training data to perform incremental training on the trained performance comparator; or, in other cases, the trained performance comparator is used to evaluate and rank the performance of the CNN model of the new first parent population.
5. The method according to claim 3, characterized in that: The encoding and clustering of the first group of trained CNN models includes: Encode each CNN model in the first group of trained CNN models and the error rate of image classification corresponding to each CNN model into a real vector, and cluster the encoded real vectors into multiple classes; Among them, the error rate N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when Returns 1 if true, otherwise returns 0.
6. The method according to claim 1, characterized in that The search space of the CNN model includes 4 convolution blocks and 1 pooling layer, and the structure of each convolution block is determined by the hyperparameters of the convolution block, and the hyperparameters of the convolution block include at least one of the following: convolution unit type, convolution layer channel expansion factor, and convolution layer repetition number; the method also includes: Each CNN model is encoded as a fixed-length integer array according to the hyperparameters of each convolutional block.
7. The method according to claim 2, characterized in that The generating a second group of CNN models based on the single-objective differential evolution operation, the proxy model and the trained first group of CNN models comprises: Changing each overall CNN model structure in the first group of trained CNN models based on the first mutation operator; or, A single convolutional block in each CNN model in the trained first group of CNN models is changed based on the second mutation operator.
8. The method according to claim 7, characterized in that The first mutation operator and the second mutation operator are different values of the target mutation operator; the target mutation operator is: Among them, j1 represents the index of the j1th CNN model in the population, represents the mutated intermediate obtained by the j1th CNN model in the population after mutation in the tth generation, represents the j1th CNN model in the population, and for A random neighbor CNN model randomly selected from the neighborhood, F represents the factor that controls the range of the CNN model mutation, r represents a real number between [0,1] randomly generated by the random number generator rand(0,1), A vector indicating the direction of change of the optimal solution and Corresponding to the optimal solution of the t-1th and t-2th generations, t is the current iteration round.
9. The method according to claim 8, characterized in that The generating a second group of CNN models based on the single-objective differential evolution operation, the proxy model and the trained first group of CNN models comprises: Exchanging the original CNN model and the convolution blocks of the original CNN model based on a first crossover operator; or exchanging random bits of the original CNN model based on a second crossover operator; Wherein, the first crossover operator is: r i2 =randI(0,2); is the i2th integer randomly generated by the random integer generator randI(0,2) in the range [0,2], CR represents the preset crossover probability factor, and j2 represents the index of the j2th variable dimension; represents the result of the j2th first crossover operator; represents the j2th dimension of the mutation solution of the target mutation operator, represents the j2th dimension of the original candidate solution of the target mutation operator; The second crossover operator is: represents the result of the j3rd second crossover operator, represents the j3th dimension of the mutation solution of the target mutation operator, Represents the j3th dimension of the original candidate solution of the target mutation operator.
10. The method according to claim 5, characterized in that The structure of constructing a deep integration model for image classification based on the second group of CNN models and the multi-objective optimization strategy includes: Evaluating the second group of CNNs based on a first objective function and a second objective function to construct a deep integration model for image classification; wherein the first objective function is used to determine the accuracy of the deep integration model, and the second objective function is used to determine the diversity of the deep integration model; Wherein, the first objective function is: Where N represents the number of test samples, |G| represents the number of test sample types, i1 represents the i1th test sample, x represents the solution of test sample s, and g x (s) is the classification result of the solution x for the test sample s, is the true classification of the test sample s, I(.) is the discriminant function, when 1 is returned when , otherwise 0 is returned; the second objective function is: represents the output of the i4th head of the deep ensemble model, represents the architecture of the kth dimension of the i4th head, represents the architecture of the kth dimension of the j4th head, express In the pth test sample s p The output of , d represents the dimension of the encoding of the head, represents the Euclidean distance between the i4th head and the j4th head, Indicates the number of test samples with different classification results.
11. The method according to claim 10, characterized in that The external archive stores non-dominated solutions that satisfy a preset objective function; if a first candidate solution among the candidate solutions of the preset objective function satisfies: no other candidate solution is better than the first candidate solution in both optimization objectives of the first objective function and the second objective function, then the first candidate solution is a non-dominated solution, and the first candidate solution is any candidate solution; Wherein, the preset objective function is: represents the j5th candidate solution, W represents the set of candidate solutions, represents the i5th candidate solution The kth dimension of represents the j5th candidate solution The kth dimension of .
12. The method according to claim 3, characterized in that The encoding and clustering of the first group of trained CNN models includes: The first set of CNN models trained based on integer group encoding; Among them, each element of the integer group is the index of the first group of CNN models after training, the length of the array is m, the first element of the integer group represents the index of the CNN model that contributes the shared layer, and the second element to the mth element of the integer group represent the index of the CNN model that constructs the head architecture.
13. A device for generating a deep integration model, the device comprising: Sampling module, dataset building module, training and evaluation module, generation module and model building module; The sampling module is used to randomly sample the neural architecture of the first group of CNN models from the search space of the classic convolutional neural network (CNN) model for image classification; The dataset construction module is used to collect representative image samples based on the scene of the image classification task, annotate the classification labels of the collected image samples, and construct a training dataset based on the annotated image samples; the image classification task includes at least one of the following: visual search, image labeling, content filtering, medical image analysis, security monitoring, agricultural monitoring, and environmental detection; The training evaluation module is used to train and evaluate the image classification performance of the sampled first group of CNN models based on the training data set; The generation module is used to generate a second group of CNN models based on the proxy model and the trained first group of CNN models; The model construction module is used to randomly construct a second parent population based on the second group of CNN models; if the termination condition is not met, the second group of CNN models is used to generate a second child population using a dual-objective differential evolution operation; the second child population is evaluated, and the second child population and the second parent population are merged into a new second parent population; a non-dominated solution of the current multi-objective optimization problem is selected from the new second parent population; wherein, for any non-dominated solution, there is no solution in the new second parent population that satisfies all optimization objectives and is better than the non-dominated solution at the same time; an external archive is updated to save the non-dominated solution; a next generation population is constructed through tournament selection; if the termination condition is met, a solution with the highest accuracy in the external archive is decoded; The constructed CNN integration model is output, where the CNN integration model is a deep integration model for image classification, and the structure of the deep integration model includes a shared block.
14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to claim 1 is implemented.
Citation Information
Patent Citations
Quick evolution method for optimized deep convolution neural network structure
CN108334949A
Deep learning parallel computing architecture method and hyper-parameter automatic configuration optimization thereof
CN111709519A
Image classification method based on neural network architecture search
CN111898689A
Evolutionary neural architecture search method and system based on Bayesian convolutional neural network
CN115908909A
Depth integration model generation method and device and computer equipment
CN117152568A
Cited By
Agent model assisted multi-target amplifier size optimization algorithm
CN120449797A
Planetary gearbox fault diagnosis algorithm automatic generation system and method based on large language model
CN121233973A
Lightweight millimeter wave radar two-dimensional feature map classification method and system oriented to FPGA hardware deployment
CN121637200A