Convolutional neural network automatic design method and system for remote sensing scene classification

Through automatic search of the convolutional neural network architecture and optimization model structure, the problem that the model is difficult to capture complex information and local optimal solutions in remote sensing scenario classification is solved, and higher classification accuracy and efficiency are achieved.

CN119942207APending Publication Date: 2025-05-06ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510047345.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing automatic design method of convolutional neural networks based on evolutionary computing has limitations in remote sensing scene classification tasks, making it difficult to effectively capture the complex diversity and heterogeneity in high-resolution remote sensing images, and random initialization of populations may lead to local optimal solutions.

Method used

A method of automatic design of convolutional neural networks for remote sensing scene classification is proposed. By automatically searching the convolutional neural network architecture, the new convolution operator and weight inheritance strategy are used to optimize the model structure to improve classification accuracy and efficiency.

Benefits of technology

The model's comprehensive understanding of global and local information of remote sensing images is improved, classification accuracy and generalization capabilities are improved, the dependence of manual design is reduced, and local optimal solutions are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942207A_ABST
    Figure CN119942207A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing scene classification-oriented convolutional neural network automatic design method and system, and the method comprises the steps: obtaining a remote sensing scene image needing to be classified, and dividing the remote sensing scene image into an evolution training set, an evolution test set and a test set; using new convolution operators to form a new search space to initialize the population; decoding the individuals of the current algebra into a network, inheriting the weight of the weight pool by using a weight inheriting strategy, carrying out training and fitness evaluation on the network, and recording the classification precision and codes of the network individuals after completing the training and fitness evaluation; constructing or updating a performance predictor using the evaluated individuals; generating a plurality of offspring individuals by using crossover and variation, inputting the plurality of offspring individuals corresponding to any parent individual into a performance predictor, and selecting high-quality offspring individuals as a next generation; judging a termination condition; and the best offspring individual is selected to be tested on the remote sensing scene test set. The method can adapt to flexible and changeable remote sensing scene images, and the algorithm has good stability and interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing scene classification methods, and in particular to an automatic design method and system for a convolutional neural network based on remote sensing scene classification. Background Art

[0002] With the rapid development of satellite sensor technology, high-resolution remote sensing images provide richer geospatial information. The goal of intelligent interpretation of remote sensing images has gradually shifted from pixel level to scene level. Remote sensing image scene classification has also received more and more attention in the fields of urban planning, environmental monitoring, natural disaster monitoring and land resource utilization. However, in the applications in the above fields, if you want to achieve real-time location detection, you need remote sensing scene classification technology to fully utilize and transform a large amount of remote sensing data. Remote sensing scene classification is the process of identifying and classifying surface features using satellite or aerial images. In the remote sensing scene classification task, the diversity and complexity of the objects make the scene classification task more challenging. There are large differences in the objects in the remote sensing images of the same scene, while there are similarities between some objects in the remote sensing images of different scenes. It is necessary to select and develop a new method to extract the information that best reflects the remote sensing scene image.

[0003] Early remote sensing image scene classification tasks mainly relied on manually extracting low-level visual features of images, such as statistics of color, shape, and texture, and then constructing feature vectors and using machine learning algorithms for modeling and classification. The core of traditional methods is that experienced scholars manually extract these features. Although these methods can effectively process the small number of remote sensing images collected at that time and the resolution is not high, their classification performance is poor. However, these methods are difficult to parse the high-level content in remote sensing images, especially when facing complex scenes, the problems of inter-class similarity and intra-class difference are still not effectively solved. In addition, manual feature extraction not only consumes a lot of time and energy, but also relies on professional knowledge and experience, which is difficult to meet the current needs of remote sensing image scene classification. This shows that traditional methods based on manual feature extraction are no longer suitable for classification tasks of high-resolution remote sensing images.

[0004] In recent years, deep learning methods used to solve remote sensing image scene classification problems have achieved certain classification results. Deep learning methods are data-driven, can automatically extract and understand features, and have shown better task processing capabilities in helping to solve many problems. In the field of remote sensing scene classification, convolutional neural network (CNN) methods have been widely used due to their excellent classification performance. However, although CNN has made significant progress, the traditional manually designed CNN structure is often fixed and difficult to flexibly adjust according to different task requirements. Remote sensing images often contain a large number of small targets and complex details, and standard CNNs are insufficient in capturing this information.

[0005] In order to further improve the performance of remote sensing scene classification, we proposed an automatic design method for convolutional neural networks for remote sensing scene classification. This method can automatically optimize the structure of the neural network and find the best model suitable for a specific task in the vast network architecture space through a search algorithm. This not only reduces the reliance on manual design, but also effectively improves the accuracy and efficiency of classification.

[0006] Existing automatic design methods for convolutional neural networks can be divided into three categories: algorithms based on reinforcement learning, optimization algorithms based on gradient descent, and optimization algorithms based on evolutionary computation. Optimization algorithms based on evolutionary computation are very popular and have been widely used in medical image segmentation, skin diagnosis, and traffic recognition, which lays the foundation for their application in remote sensing scene classification. Remote sensing scene classification tasks usually involve complex multi-source data and multi-scale features, which makes it challenging to design suitable deep learning models. In order to effectively extract and utilize these features, it is necessary to optimize the structure of the model to balance accuracy, computational complexity, and generalization ability. Evolutionary algorithms are good at gradually optimizing the model structure and finding a network architecture with excellent performance when the search space is large and the problem is complex. Therefore, applying automatic design methods based on evolutionary computation to remote sensing scene classification can automate the model design process, improve the performance of the model when processing high-dimensional and diverse remote sensing data, and promote remote sensing classification tasks to achieve higher accuracy and better generalization ability.

[0007] In the process of implementing the present disclosure, the inventors found that the prior art has the following technical problems:

[0008] Although the automatic design method of convolutional neural networks based on evolutionary computation has shown excellent performance in many fields, there are some limitations in reprocessing remote sensing scene images. In the remote sensing scene classification task, remote sensing images have high resolution, complex and diverse ground features, and significant spatial and spectral heterogeneity, which puts higher requirements on deep learning models. Remote sensing images usually contain rich global and local information, which requires the model to effectively capture and process these features.

[0009] In addition, the automatic design method of convolutional neural networks based on evolutionary computation usually initializes the population in a random way during the search, but this randomness may bring certain risks. If the distribution of individuals in the initial population is not diverse enough, or the quality of these individuals is low, the algorithm may fall into a local optimal solution, making it difficult to find the global optimal solution. Summary of the invention

[0010] The purpose of the present invention is to overcome the above defects and develop a new automatic design method of convolutional neural network for remote sensing scene classification. By automatically searching the convolutional neural network architecture, a convolutional neural network for remote sensing scene classification is obtained.

[0011] The technical solution adopted by the present invention is: a convolutional neural network automatic design method for remote sensing scene classification, comprising the following steps:

[0012] S1, obtain the remote sensing scene image that needs to be classified;

[0013] S2, divide the remote sensing scene data into evolutionary training set, evolutionary test set and test set;

[0014] S3, using a new convolution operator to form a new search space to initialize the population;

[0015] S4, decode the individuals of the current generation into a network, use the weight inheritance strategy to inherit the weights of the weight pool, and train and evaluate the fitness of the network. After the evaluation is completed, record the classification accuracy and encoding of the network individuals;

[0016] S5, update the weight pool;

[0017] S6, construct or update performance predictors using the evaluated individuals;

[0018] S7, using crossover and mutation to generate multiple offspring individuals, inputting multiple offspring individuals corresponding to any parent individual into the performance predictor, and selecting high-quality offspring individuals as the next generation;

[0019] S8, determine whether the termination condition is met; if the termination condition is not met, turn to S4; if the termination condition is met, turn to S9;

[0020] S9, select the best offspring individuals and test them on the remote sensing scene test set.

[0021] Further, the specific steps of S3 are:

[0022] S31, randomly initialize the population individual codes according to the coding strategy;

[0023] S32, decoding the corresponding population individual code, finding the corresponding convolution operator and pooling operator to form a basic search space;

[0024] S33, processing the remote sensing scene data into a new network according to the spatial decoding of the basic search.

[0025] Further, in S32, two convolution operators with different convolution kernels, two pooling operators and an identity mapping are set to form a new search space, wherein the convolution operators with two different convolution kernels are MB convolution and FR convolution, and the two pooling operators are maximum pooling and average pooling.

[0026] Further, the specific steps of S5 are:

[0027] S51, preferentially retaining the optimal individual's operation weight into the weight pool;

[0028] S52, secondly, retaining the current individual operation weight in the weight pool;

[0029] S53, forming a large supernet weight pool.

[0030] Furthermore, the steps for generating multiple offspring by crossover mutation in S7 are:

[0031] S711, randomly select two individuals from the population;

[0032] S712, traverse each bit code of the individual, perform crossover and mutation;

[0033] S713, judging whether the preset number of crossover and mutation rounds has been reached, if reached, turning to S711, continuing to select the remaining individuals, if not reached, jumping to S712;

[0034] Furthermore, the steps for selecting high-quality offspring in S7 are:

[0035] S721, inputting a plurality of offspring individuals corresponding to any one parent individual into a performance predictor;

[0036] S722, the performance predictor is scored. If the score is greater than 0.5, it proves that the offspring predicted by the performance predictor is better than the parent generation; if the score is less than 0.5, it will be directly discarded;

[0037] S723, select the offspring with the highest score among the offspring with a score greater than 0.5 as the offspring;

[0038] S724, repeat the whole process until all individuals in the parent population are traversed.

[0039] Further, the specific steps of S722 are:

[0040] S7221, obtain the code of each offspring individual;

[0041] S7222, use each decision tree to perform prediction scoring to determine whether it is better than the classification accuracy of the parent generation;

[0042] S7223, the score of the offspring is the sum of all decision trees that are better than the parent generation divided by the total number of decision trees.

[0043] Further, the specific steps of S9 are:

[0044] S91, reading and decoding the code of the final optimal individual required for the test;

[0045] S92, set the corresponding parameters and perform final training on the test set;

[0046] S93, obtaining the final classification accuracy.

[0047] Another technical solution of the present invention to solve the above technical problem is as follows: a convolutional neural network automatic design system for remote sensing scene classification, comprising:

[0048] Data acquisition module, which acquires images that need to be classified;

[0049] Preprocessing module, which preprocesses remote sensing scene data and divides data sets;

[0050] The population initialization module uses a new convolution operator to form a new search space to initialize the population;

[0051] The fitness evaluation module uses the weight inheritance strategy to decode individuals into networks and train them;

[0052] The offspring generation module generates multiple individuals using crossover and mutation, and selects high-quality individuals through performance predictors;

[0053] The testing module selects the best individuals to test on the test set and obtains the final classification results.

[0054] The beneficial effects of the present invention are:

[0055] 1. In remote sensing scene classification, this paper proposes a new search space to form a basic CNN; two new convolution operators are developed and applied to the new search module, which can take into account both the classification accuracy and model parameters, and improve the model's ability to comprehensively understand global and local information;

[0056] 2. This invention proposes a new weight inheritance strategy; by building an elite weight pool, the weight of the best individual is preferentially retained, which effectively improves the search efficiency and the convergence of the model;

[0057] 3. The present invention proposes a new offspring generation strategy, which generates multiple individuals through crossover and mutation, selects high-quality individuals through performance predictors, and improves the proportion of high-quality individuals in the offspring and the quality of the offspring individuals. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0059] The present invention will be further described below in conjunction with the accompanying drawings.

[0060] like Figure 1 As shown, the present invention is a convolutional neural network automatic design method for remote sensing scene classification, comprising the following steps:

[0061] S1, obtain the remote sensing scene image that needs to be classified.

[0062] S2, divide the remote sensing scene data into evolutionary training set, evolutionary test set and test set.

[0063] S3, use the new convolution operator to form a new search space to initialize the population; the specific steps are:

[0064] S31, randomly initialize the population individual codes according to the coding strategy;

[0065] S32, decode the corresponding population individual code, find the corresponding convolution operator and pooling operator to form a basic search space; set two convolution operators with different convolution kernels, two pooling operators and an identity mapping to form a new search space, which includes the seven operations in Table 1 below, which can take into account both the classification accuracy and model parameters, and improve the model's comprehensive understanding of global and local information.

[0066] Among them, MB convolution can not only effectively capture the multi-level information of input features, but also be suitable for resource-constrained deep learning applications at a low computational cost. The basic structure of MB convolution consists of three convolutional layers and their corresponding batch normalization layers. First, MB convolution performs channel expansion on the input features through a 1×1 convolutional layer. This process is controlled by the expansion factor parameter, which aims to increase the number of channels to enhance the feature representation capability while keeping the spatial resolution unchanged. Subsequently, depthwise separable convolution is used to perform spatial convolution on the expanded features. This convolutional layer uses convolution kernels and grouped convolution strategies to perform convolution operations only within each channel, thereby significantly reducing the computational complexity. The design of the padding strategy ensures that the spatial dimension of the output features is consistent with the input features, effectively maintaining the geometric structure of the features. Finally, after the second 1×1 convolutional layer, MB convolution compresses the expanded feature map back to the number of output channels so that it can be effectively connected with the subsequent network layers. This process also maintains the number of channels and spatial resolution of the features. Another FR operation consists of an expanded convolutional layer and an effective channel attention module. Specifically, given the input features, FR first uses a dilated convolutional layer to exponentially increase the receptive field, thereby incorporating more global information without losing resolution or coverage. Subsequently, a global average pooling layer is applied to each channel. Next, a pooling layer of size A one-dimensional convolutional layer and a The sigmoid activation layer generates the weights for each channel. This one-dimensional convolutional layer is designed to capture the relationship between each channel and its The nonlinear interaction between adjacent channels, where the kernel size Indicates the number of neighboring channels that participate in channel attention prediction.

[0067] Table 1 Operation Type

[0068]

[0069] S33, processing the remote sensing scene data into a new network according to the spatial decoding of the basic search.

[0070] S4, decode the individuals of the current generation into a network, use the weight inheritance strategy to inherit the weights of the weight pool, and train and evaluate the fitness of the network. After the evaluation is completed, record the classification accuracy and encoding of the network individuals;

[0071] S5, update the weight pool; the specific steps are:

[0072] S51, preferentially retaining the optimal individual's operation weight into the weight pool;

[0073] S52, secondly, retaining the current individual operation weight in the weight pool;

[0074] S53, forming a large supernet weight pool.

[0075] S6, construct or update a performance predictor using the evaluated individuals.

[0076] S7, using crossover and mutation to generate multiple offspring individuals, multiple offspring individuals corresponding to any parent individual are input into the performance predictor, and high-quality offspring individuals are selected as the next generation.

[0077] Because the initialization of the population is often completed randomly, this randomness brings certain risks. If the individual distribution in the initial population is not diverse enough or the quality of the initial individuals is low, the algorithm may fall into a local optimal solution and it is difficult to find a global optimal solution. In this case, the quality of the offspring population may decrease from generation to generation, causing the algorithm to converge more slowly or even have poor results. In order to improve the efficiency of the algorithm, the core strategy adopted by the present invention is to perform fine screening through a performance predictor during offspring generation and population update, that is, only retain those offspring individuals that are better than the parent generation.

[0078] Among them, the steps for crossover mutation in S7 to produce multiple offspring are:

[0079] S711, randomly select two individuals from the population;

[0080] S712, traverse each bit code of the individual, perform crossover and mutation;

[0081] S713, determine whether the preset number of crossover and mutation rounds has been reached. If it has been reached, go to S711 and continue to select the remaining individuals. If it has not been reached, jump to S712.

[0082] The steps for selecting high-quality offspring in S7 are as follows:

[0083] S721, inputting a plurality of offspring individuals corresponding to any one parent individual into a performance predictor;

[0084] S722, the performance predictor is scored. If the score is greater than 0.5, it proves that the offspring predicted by the performance predictor is better than the parent. If the score is less than 0.5, it will be directly discarded. The specific steps are:

[0085] S7221, obtain the code of each offspring individual;

[0086] S7222, use each decision tree to perform prediction scoring to determine whether it is better than the classification accuracy of the parent generation;

[0087] S7223, the score of the offspring is the number of decision trees that are better than the parent generation divided by the total number of decision trees;

[0088] S723, select the offspring with the highest score among the offspring with a score greater than 0.5 as the offspring;

[0089] S724, repeat the whole process until all individuals in the parent population are traversed.

[0090] S8, determine whether the termination condition is met; if the termination condition is not met, turn to S4; if the termination condition is met, turn to S9.

[0091] S9, select the best offspring individuals and test them on the remote sensing scene test set; the specific steps are:

[0092] S91, reading and decoding the code of the final optimal individual required for the test;

[0093] S92, set the corresponding parameters and perform final training on the test set;

[0094] S93, obtaining the final classification accuracy.

[0095] In order to further illustrate the superiority of the present invention in the remote sensing scene classification problem, Table 2 shows the classification accuracy of the present invention method and some typical manually designed networks ResNet34, GoogleNet and the classic RelativeNAS based on the search space in processing the UCM typical remote sensing scene classification data set.

[0096] The experimental results on a typical remote sensing scene classification dataset are given in the embodiment. The UCM dataset has 21 categories, each category has 100 images, and the image size is 256×256. The experimental results are obtained by running on four V100 GPUs, with the population size set to 20 and the number of generations set to 50. Without any pre-training, the artificially designed network runs independently 10 times, and our method runs independently 5 times. Table 1 shows the maximum accuracy, average accuracy, model parameter quantity and running time. Each block in the table shows the experimental results on the UCM dataset. By comparison, it can be seen that in the remote sensing scene classification, the method proposed in the present invention shows superior classification performance. Specifically, whether it is the highest accuracy or the average accuracy, the proposed method has excellent performance than other methods. These results also show that the proposed method can find a good architecture for remote sensing scene classification, proving its effectiveness in classification tasks.

[0097] Table 2 Comparison of experimental results on UCM dataset

[0098]

[0099] In summary, the automatic design method of convolutional neural network for remote sensing scene classification proposed in the present invention can effectively process remote sensing data sets and find a good architecture.

[0100] The present invention also provides a convolutional neural network automatic design system for remote sensing scene classification, comprising:

[0101] Data acquisition module, which acquires images that need to be classified;

[0102] Preprocessing module, which preprocesses remote sensing scene data and divides data sets;

[0103] The population initialization module uses a new convolution operator to form a new search space to initialize the population;

[0104] The fitness evaluation module uses the weight inheritance strategy to decode individuals into networks and train them;

[0105] The offspring generation module generates multiple individuals using crossover and mutation, and selects high-quality individuals through performance predictors;

[0106] The testing module selects the best individuals to test on the test set and obtains the final classification results.

[0107] Although the present invention has been described in detail above with general description and specific embodiments, some modifications or improvements can be made on the basis of the present invention. The above description is only a preferred embodiment of the present invention and is not limited to the scope of the present invention. Other changes and modifications made by those skilled in the art without departing from the spirit and scope of protection of the present invention are still included in the scope of protection of the present invention.

Claims

1. An automatic design method of convolutional neural network for remote sensing scene classification, characterized by: The following steps are involved: S1, obtain the remote sensing scene image that needs to be classified; S2, divide the remote sensing scene data into evolutionary training set, evolutionary test set and test set; S3, using a new convolution operator to form a new search space to initialize the population; S4, decode the individuals of the current generation into a network, use the weight inheritance strategy to inherit the weights of the weight pool, and train and evaluate the fitness of the network. After the evaluation is completed, record the classification accuracy and encoding of the network individuals; S5, update the weight pool; S6, construct or update performance predictors using the evaluated individuals; S7, using crossover and mutation to generate multiple offspring individuals, inputting multiple offspring individuals corresponding to any parent individual into the performance predictor, and selecting high-quality offspring individuals as the next generation; S8, judging whether the termination condition is met; If the termination condition is not met, go to S4; If the termination condition is reached, go to S9; S9, select the best offspring individuals and test them on the remote sensing scene test set.

2. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 1, characterized in that: The specific steps of S3 are: S31, randomly initialize the population individual codes according to the coding strategy; S32, decoding the corresponding population individual codes, and finding the corresponding convolution operator and pooling operator to form a basic search space; S33, processing the remote sensing scene data into a new network according to the spatial decoding of the basic search.

3. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 2, characterized in that: In S32, two convolution operators with different convolution kernels, two pooling operators and an identity mapping are set to form a new search space, wherein the convolution operators with two different convolution kernels are MB convolution and FR convolution, and the two pooling operators are maximum pooling and average pooling.

4. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 1, characterized in that: The specific steps of S5 are: S51, preferentially retaining the optimal individual's operation weight into the weight pool; S52, secondly, retaining the current individual operation weight in the weight pool; S53, forming a large supernet weight pool.

5. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 1, characterized in that: The steps for crossover mutation to produce multiple offspring in S7 are: S711, randomly select two individuals from the population; S712, traverse each bit code of the individual, perform crossover and mutation; S713, determine whether the preset number of crossover and mutation rounds has been reached. If it has been reached, go to S711 and continue to select the remaining individuals. If it has not been reached, jump to S712.

6. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 1, characterized in that: The steps for selecting high-quality offspring in S7 are: S721, inputting a plurality of offspring individuals corresponding to any one parent individual into a performance predictor; S722, the performance predictor is scored. If the score is greater than 0.5, it proves that the offspring predicted by the performance predictor is better than the parent generation; If the score is less than 0.5, it will be discarded directly; S723, select the offspring with the highest score among the offspring with a score greater than 0.5 as the offspring; S724, repeat the whole process until all individuals in the parent population are traversed.

7. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 6, characterized in that: The specific steps of S722 are: S7221, obtain the code of each offspring individual; S7222, use each decision tree to perform prediction scoring to determine whether it is better than the classification accuracy of the parent generation; S7223, the score of the offspring is the sum of all decision trees that are better than the parent generation divided by the total number of decision trees.

8. The method for automatically designing a convolutional neural network for remote sensing scene classification according to claim 1, characterized in that: The specific steps of S9 are: S91, reading and decoding the code of the final optimal individual required for the test; S92, set the corresponding parameters and perform final training on the test set; S93, obtaining the final classification accuracy.

9. Automatic design system of convolutional neural network for remote sensing scene classification, characterized by: include: Data acquisition module, which acquires images that need to be classified; Preprocessing module, which preprocesses remote sensing scene data and divides data sets; The population initialization module uses a new convolution operator to form a new search space to initialize the population; The fitness evaluation module uses the weight inheritance strategy to decode individuals into networks and train them; The offspring generation module generates multiple individuals using crossover and mutation, and selects high-quality individuals through performance predictors; The testing module selects the best individuals to test on the test set and obtains the final classification results.