A method for crop recognition in remote sensing images based on selective learning and playback
By introducing an incremental learning method SeRe, a selective learning and playback strategy, in the remote sensing semantic segmentation model, the problems of model update delay and catastrophic forgetting are solved, the rapid update of the model and the retention of old knowledge are achieved, and the overall reasoning performance of the model is improved.
Patent Information
- Application Number
- CN202210454684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-24
AI Technical Summary
In the field of remote sensing, semantic segmentation models are difficult to update quickly when facing massive new data, resulting in delay in model updates and prone to catastrophic forgetting problems in the incremental learning process, affecting the model's inference performance.
A incremental learning method based on selective learning and replay is proposed, called SeRe (Selective Learning with Replay). Through selective learning, obtain high-value samples from new data for training and quickly update the model; at the same time, obtain difficult samples from old data through playback strategies to consolidate and alleviate catastrophic forgetting.
It realizes the rapid iterative update of the model in the context of remote sensing big data, while retaining old knowledge, improving the model's inference ability on new and old knowledge, and is better than other existing incremental learning methods.
Smart Images

Figure CN114821324B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of remote sensing technology, and in particular relates to a remote sensing image crop recognition method based on selective learning and playback. Background Art
[0002] In the field of remote sensing, semantic segmentation is an important image interpretation task. Deep learning has made great progress in the field of semantic segmentation in recent years. With the continuous generation of massive remote sensing data, deep learning models need to continuously learn new data and iteratively generate new models to maintain good generalization capabilities. As we all know, model training relies on a lot of annotation costs and training time, and faces the problem of catastrophic forgetting in the incremental learning process.
[0003] The semantic segmentation method originated from FCN. FCN uses a skip-layer structure to fuse shallow features with deep features to obtain more comprehensive features and achieve better segmentation results. Later, U-Net used this encoder-decoder idea to effectively extract contextual information and spatial features of images. With the emergence of dilated convolution, more spatial information is retained by increasing the receptive field, thereby improving the model effect. Recently, research work combining semantic segmentation tasks with attention mechanisms, multi-scale problems, and pooling structures has also emerged. Semantic segmentation has important research and application value in intelligent understanding of remote sensing images. Benefiting from the rapid development of deep learning, the performance of remote sensing semantic segmentation models has been significantly improved. As we all know, deep learning relies on a large amount of labeled data and requires a long training cycle. However, with the development of remote sensing technology, a large amount of remote sensing image data is continuously collected, and the model needs to learn labeled data to improve generalization ability. In addition, especially in some fields, such as crop remote sensing identification and crop remote sensing yield estimation, models need to be updated quickly to play a real-time role. Based on biological principles, these fields collect and analyze the different spectral characteristics of various crops at different growth stages, and use the surface information recorded by sensors on the platform to identify crop types, monitor crop growth, and predict crop yields before harvesting. Therefore, the rapid update of semantic segmentation models in the context of remote sensing big data is of great significance.
[0004] Model updating refers to the process of building a new model by learning new data from the old model. Due to the rapid accumulation of massive remote sensing data, the model needs to learn new data in a short period of time to complete the update, otherwise there will be a problem of model update delay. To solve the problem of update delay, the following two aspects need to be considered. First, the accuracy of the model is improved by learning new knowledge, and the new knowledge is contained in the growing massive data. It is almost impossible to learn all the new data; second, the model update process is centered on model training, and the model training speed is limited by the volume of new data and the cost of data annotation. This requires us to make a compromise between model accuracy and update speed. Catastrophic forgetting refers to the model forgetting old knowledge in the process of learning new knowledge. Catastrophic forgetting will seriously affect the reasoning performance of the model and is an issue that cannot be ignored in model updating. At present, the field of incremental learning pays the most attention to the problem of catastrophic forgetting. The corresponding solutions are mainly divided into two categories: replay and regularization. The replay method mainly retains some old samples and consolidates the old knowledge when the model is trained on new tasks. The regularization method is to add some constraints when the model is trained on a new task. The most typical method is knowledge distillation, which overcomes forgetting by narrowing the difference between the output results of the new model and the old model on the old data.
[0005] Regarding the above problems of model update delay and catastrophic forgetting: As for the former, although the active learning-based method uses the query mechanism to select more uncertain data from the newly added unlabeled data and annotates it for model training, it reduces the time cost of model training to a certain extent, but does not consider the problem of catastrophic forgetting in the incremental segmentation scenario, and the model performs poorly on old data. As for the latter, catastrophic forgetting is the most concerned issue in the current incremental learning field. A lot of effort is spent on maintaining the stability of the model on old data while ignoring the plasticity of new data. There is a lack of control over the amount of new data and the cost of annotation, and the model updates slowly. It is worth mentioning that the above methods are mainly used for image classification tasks, while methods for semantic segmentation tasks are currently rarely studied.
[0006] Active learning refers to querying the data that needs to be labeled from unlabeled data. Its query strategies can be mainly divided into three types: uncertainty-based methods, diversity-based methods, and expected model changes. The uncertainty-based method is to find the most uncertain samples of the model for labeling, which is a commonly used method in active learning. For the case of two classifications, samples with a posterior probability less than 0.5 are selected. For multi-classification problems, cross entropy is used as a measure of uncertainty. In general, samples with smaller posterior probabilities are samples with higher uncertainty. Most uncertainty-based methods can be directly applied to neural networks because the Softmax of neural networks can output information about uncertainty such as confidence and cross entropy. Active learning also has relevant applications in the field of semantic segmentation. The application of active learning in probabilistic models is explored, and probability distribution encoding is performed through large structured output spaces (such as semantic segmentation). The network uncertainty based on MCdropout is used for region selection to reduce the labeling cost.
[0007] According to different incremental scenarios, semantic segmentation tasks can be divided into category increment and domain increment. At present, incremental learning applied to semantic segmentation tasks focuses on the scenario of category increment, such as the background shift problem when the category is incremented. A loss function including distillation loss and cross entropy loss for the old category is designed to solve this problem. This work is considered to be the beginning of incremental learning for semantic segmentation tasks in the remote sensing field. Some studies have made improvements on this basis, overcoming the problem of forgetting by using a multi-scale spatial distillation loss, and solving the problem of background semantic expertise by using a confidence-based pseudo-labeling strategy. At present, there is little work on domain increment for semantic segmentation tasks, but its necessity cannot be ignored. Summary of the invention
[0008] In view of this, the present invention proposes an incremental learning method based on selective learning and replay. Aiming at the problems of semantic segmentation model update delay and catastrophic forgetting faced in the incremental scenario of remote sensing images, a selective learning and replay framework is designed, called SeRe (Selective Learning with Replay). SeRe takes high-value sample sampling as its core idea, makes up for the shortcomings of existing methods, and reflects its unique advantages. For the problem of model update delay, SeRe uses an uncertainty-based selective learning method to obtain high-value new data for model training, thereby speeding up the update of the model. For the problem of catastrophic forgetting, SeRe uses a method of uniform sampling of difficult and easy samples to obtain samples from old data for playback, which improves the problem of catastrophic forgetting. The advantage of SeRe is that it can achieve rapid model updates while retaining old knowledge, so that the model exhibits strong reasoning ability on both new and old knowledge.
[0009] The present invention discloses a remote sensing image crop recognition method based on selective learning and playback, comprising the following steps:
[0010] The selective learning selects data with more learning value from the new data as the training set D of this stage. t , and D t Marking, D t We have paired images and annotation files (I t ,G t ), where I t Represents the W×H, G of an image t Represents the corresponding true label of the same size; in the learning stage t, the set of all parameters of the model after learning is recorded as θ t , the segmentation result of the model for the image in the test set is recorded as Where W is width, H is height, and C is the number of categories;
[0011] The playback uses the model to infer and evaluate the old data, calculates the segmentation accuracy IoU of a single image, sorts the difficulty of the old samples by IoU, and obtains this part of the old samples by uniform sampling.
[0012] Furthermore, the data with more learning value uses the uncertainty measurement of the sample, and the measurement method of the uncertainty of the sample is as follows:
[0013] represents the inference result of the model in stage t for unlabeled data, and the old model in stage t-1 for new data D t The inference result is expressed as The inferred pseudo-label of the old model for a new image is:
[0014]
[0015] Since each pixel’s pseudo-label has its confidence, the uncertainty of each pixel is calculated based on it:
[0016]
[0017] Where w is the width, h is the height, c is the number of categories, and u is the value of the uncertainty of the image pixel;
[0018] Then the uncertainty of the sample is:
[0019] ∑U t-1→t (w,h,u).
[0020] Furthermore, the calculation process of the IoU of the single image is as follows:
[0021] The old data D t-1 As the input of the t-stage model, the output after model processing is Through one-hot encoding, the category with the highest inference confidence is taken as the pixel category to obtain the pseudo label:
[0022]
[0023] Old data D t-1 It is labeled data, denoted by G t-1 According to the pseudo-label and the true label G t-1 Calculate the IoU of this image. First calculate the IoU of each category in this image, and then find the average of the IoU of all categories as the IoU of the image:
[0024]
[0025]
[0026] Furthermore, the higher the IoU value, the better the reasoning effect of the model, and the better the model's memory of this part of the data. Conversely, the lower the IoU value, the worse the reasoning effect of the model, and the model has forgotten this part of the data.
[0027] Furthermore, all images are sorted according to the IoU value, and then 20% of the samples are sampled for playback at a certain interval, and the playback ratio is 25%.
[0028] Furthermore, the model is a Deeplab-v3+ model, and uses ResNet-50 as a backbone network, and the backbone network uses a pre-trained model of ImageNet.
[0029] The beneficial effects of the present invention are as follows:
[0030] 1) Apply the incremental learning method to the task of remote sensing semantic segmentation, so that the model can be quickly iterated and updated in the context of remote sensing big data.
[0031] 2) We propose an incremental learning framework, SeRe, for semantic segmentation tasks. The model learns new knowledge and retains old knowledge through selective learning and replay. The selective learning and replay approach can not only allow the model to be updated quickly, but also alleviate catastrophic forgetting. Experiments show that this invention is superior to other current incremental learning methods.
[0032] 3) The present invention was evaluated on multiple data sets, and experiments proved that with 20% of new data, the present invention can achieve 80% of the joint training accuracy. Joint training is considered to be the upper limit of the accuracy of incremental learning. At the same time, the training speed is improved with the same accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 Model framework diagram of the present invention;
[0034] Figure 2 This is a diagram of experimental results, which shows from left to right the original image, the result of processing by the annotation method, the result of processing by the static learning method, the result of processing by the fine-tuning method, the result of processing by the combined method and the result of processing by the present invention. DETAILED DESCRIPTION
[0035] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.
[0036] Deep learning model training relies on a large amount of manually annotated data, which is also very difficult to solve in the field of semantic segmentation of remote sensing. Figure 1 As shown, the present invention provides a selective learning strategy based on uncertainty, which is mainly used for learning value assessment and sampling of unlabeled samples. By determining a quantitative ratio, part of the data with more learning value in the unlabeled samples can be obtained, and then it is labeled to obtain labeled data for model training. The advantage of selective learning based on uncertainty is that on the one hand, it reduces the time cost and material cost of labeling, and on the other hand, it reduces the amount of data that the model needs to learn through a selection mechanism, indirectly shortening the training time of the model.
[0037] The semantic segmentation domain incremental model can continue to learn in t=1…T stages, and the model is updated at the end of each stage. In each learning stage, a new data arrives, and we assume that this new data is temporarily unlabeled. Through our selective learning method, we select data with more learning value from the new data as the training set D for this stage. t , and D t Marking, D t We have paired images and annotation files (I t ,G t ). t Represents the W×H, G of an image t represents the corresponding true label of the same size, namely W×H. Since we are dealing with the problem of domain increment, the category C in the dataset at each learning stage is fixed, and the number of categories is constant, namely |C| = |C 1 |=|Ct |=|C t+1 |. Generally speaking, a neural network model can be represented as a combination of a feature extractor f(·) and a classifier g(·). In the learning stage t, the features of the image are extracted by the feature extractor f t (·) is obtained by classifier g t (·) Output the classification result. We record the segmentation result of the model for the image of the test set in stage t as The set of all parameters of the model after learning is recorded as θ t .
[0038] The content of the selective learning module includes: uneven distribution of learning value - high learning value samples - uncertainty measurement. The present invention pays attention to the phenomenon of uneven distribution of learning value of samples in new data. When the model is inferring, different samples show different results. On the one hand, there may be some samples in the new data that are highly similar to the old samples, so the model will have better reasoning effect on this type of samples; on the other hand, there are some samples in the new data that are relatively unfamiliar to the model. Although the model's reasoning effect on the above is not good, these samples contain richer new knowledge. Therefore. Samples with high similarity contain lower learning value. If the model learns too much of this type of data, it will hinder the model from learning new knowledge, and the robustness of the model will also deteriorate. To achieve iterative updates in a short period of time, the model needs to learn some samples that it is not good at reasoning. These samples are considered to be more valuable uncertainty samples. The following will explain how to measure the uncertainty of samples.
[0039] The last activation function layer of the semantic segmentation network, the softmax layer, outputs the model's predicted category probability for each pixel in the image. This probability represents the confidence that the pixel belongs to each category. The greater the confidence, the greater the possibility that the pixel belongs to that category, indicating that the model has a higher degree of certainty in reasoning about this pixel. Usually, when we reason about unknown images, we directly take the category with the highest confidence as the category to which the pixel belongs. We use represents the inference result of the model in stage t for unlabeled data, then the old model in stage t-1 is used to infer the new data D t The inference result can be expressed as The inference pseudo-label of the old model for a new image is
[0040]
[0041] Since each pixel's pseudo-label has its confidence, the uncertainty of each pixel is calculated based on it. Usually the inverse of the confidence is taken as the uncertainty. The lower the confidence, the higher the uncertainty, and the higher its learning value. The uncertainty of the old model for a pixel in a new image is:
[0042]
[0043] Then the uncertainty of an image is:
[0044] ∑U t-1→t (w,h,u)
[0045] Playback
[0046] To address the catastrophic forgetting problem during the incremental learning process of the model, we designed a playback strategy to sort the difficulty of old samples by IoU and obtain these old samples by uniform sampling. The samples learned by the model are all labeled data, so the model is used to infer and evaluate the old data to calculate the segmentation accuracy IoU of a single image. The higher the IoU value, the better the reasoning effect of the model, and the better the model's memory ability for this part of the data. Conversely, the lower the IoU value, the worse the reasoning effect of the model, and the model has forgotten this part of the data. The calculation process of IoU for a single image is as follows.
[0047] The old data D t-1 As the input of the t-stage model, the output after model processing is Through one-hot encoding, the category with the highest inference confidence is taken as the pixel category to obtain the pseudo label:
[0048]
[0049] Old data D t-1 It is labeled data, denoted by G t-1 According to the pseudo-label and the true label G t-1 Calculate the IoU of this image. We first calculate the IoU of each category in this image, and then find the average of the IoU of all categories as the IoU of the image:
[0050]
[0051]
[0052] In order to better simulate the distribution of the difficulty and ease features of all old data, we sorted all images according to the IoU value, and then sampled 20% of the samples at a certain interval for playback. When replaying old data, due to time cost constraints, it is impossible to replay a large amount of old data to help the model overcome forgetting. Therefore, in this experiment, the method adopted is to obtain 20% of the old data from the old data for playback.
[0053] The Xiangtan dataset used in this experiment is a remote sensing image taken by the Gaofen-2 satellite, covering the Xiangtan City of China. The image has 3 bands, RGB, and a resolution of 2 meters. There are 9 categories in the image: background, farm, urban area, suburbs, water, woods, grassland, road, and others. The dataset contains 106 images with an original size of 4096×4096, which is too large for model training. Therefore, we cropped it into 19866 images of 256×256 size and randomly divided it into training set A, training set B, and test set C according to 4:4:2, with 8026, 8025, and 3815 images respectively. Training set A is regarded as the first batch of data, and training set B is regarded as the second batch of data.
[0054] The ISPRS Potsdam dataset contains 38 high-resolution remote sensing images with a size of 6000×6000, a spatial resolution of 5cm, and a total of 4 bands: R, G, B, and NIR. There are 6 categories in the dataset, namely impervious surfaces, buildings, low plants, trees, cars, and others. In order to facilitate model training, we cropped the images to obtain 20,102 256×256 images. Since a large dataset would consume too much training cost, we extracted 10,000 images as the experimental dataset while ensuring that the number of images was sufficient, and randomly divided them into training set A, training set B, and test set C according to 4:4:2, with 4,000, 4,000, and 2,000 images respectively.
[0055] Since the task of the present invention is essentially a semantic segmentation task, the general indicator mIoU in the field of semantic segmentation is used as the evaluation indicator in terms of model accuracy, which represents the intersection-over-union ratio of the real area and the predicted area. In addition, aAcc and OA are also used as accuracy indicators, representing the average accuracy and overall accuracy, respectively. Since our experimental scenario focuses on the iterative update speed of the model, we use GPU / hour to evaluate the training speed of the model. In summary, we comprehensively evaluate the effect of the model from the two aspects of accuracy and speed.
[0056] The present invention uses the following baseline:
[0057] Static learning: This is a traditional non-incremental learning method. It assumes that all training data arrives at once and is labeled. The model is obtained by training this data, and the model is no longer updated.
[0058] Joint: This is a classic joint training method that has the ability of incremental learning and is considered to be the upper limit of incremental learning accuracy. The training data can arrive in batches, but when new data arrives, all data, including new data and old data, will be trained.
[0059] Fine-tuning: After the model learns the old data, it will get a model with reasoning ability, which we call the base model. When new data arrives, the base model is trained with the new data. Although the model can effectively learn the new data, it will have catastrophic forgetting problems on the old data.
[0060] For all methods, the present invention adopts the Deeplab-v3+ model and uses ResNet-50 as the backbone network. The backbone network is initialized using the pre-trained model of ImageNet. In the Xiangtan dataset, the batch-size of each training stage is 2, and the iteration is 20 cycles. SGD is used to accelerate the convergence of the model during training, and the learning rate is 0.0001. In the Potsdam dataset, the batch-size of each training stage is 32, and the iteration is 200 cycles. Adam is used as the optimizer during training, and the learning rate is 0.001. For each dataset, we use 40% of the data as historical data, 40% of the data as new data, and the remaining 20% of the data as test data.
[0061] Comparison of experimental results under fast model updating
[0062]
[0063] The performance of each method on the Xiangtan dataset is shown in the table above, and the method with the best performance in each evaluation index is highlighted. The joint method has the best performance in terms of accuracy, because in the training stage, compared with other methods, the model learns all the data and there is no catastrophic forgetting problem. Although Joint is better than other methods in terms of accuracy, it takes more storage space to save historical data, and the training speed is much slower than other methods, which is not desirable in real-world applications. Static learning is a non-incremental method. After the training is completed, the model is no longer updated, so the model cannot improve the generalization ability by learning new data, and the performance effect is the worst. The fine-tuning method uses the model initialization obtained by training based on historical data, which retains the model's memory of old knowledge to a certain extent, but with the addition of new data in the new training stage, the model forgets the old data very seriously. The SeRe of the present invention not only achieves the result closest to the highest accuracy, but also the model training speed is the fastest under the condition of equal accuracy. In summary, the present invention has shown faster and better advantages in model evaluation in the experiment.
[0064] Ablation Experiment Results
[0065]
[0066] The effectiveness of the selective learning and playback mechanism in the present invention is verified by ablation experiments. When incremental learning is performed without playback, the present invention without playback is significantly better than the random learning of the present invention, with mIoU increased by 10.94 and mAc increased by 11.65. This is because selective learning helps the model acquire richer knowledge under the condition of the same number of pictures. In addition, the playback results of the present invention have increased mIoU by 1.64 and 8.97 respectively compared with the no playback results and random playback results. It can be noted that the sampling strategy of playback has a great impact on the results. Improper sampling will cause the model accuracy to decrease instead of improve. The exploration of the playback sampling method will be described in the following analysis experiment.
[0067] On the same basic model, the playback ratio is fixed at 25%, and different sample sampling methods are used to obtain samples from old samples for playback. Random playback is to randomly select 25% of the samples for playback. Weighted playback based on pixel frequency (.ref) is to calculate the weight of the category according to the frequency of occurrence of different categories, and score each image and take the top 25% of the samples with the highest score for playback. Since the basic model has a certain reasoning ability and has old data and its annotations, the basic model is used to reason and evaluate the old data, and mIoU is used as an evaluation indicator to judge the quality of the model reasoning results. mIoU also reflects the difficulty of sample reasoning. Sort by the mIoU value of each image in descending order. Take 25% of the samples from the head, middle, and end of the sample sequence for playback, respectively, corresponding to the methods of replaying difficult samples first, replaying medium samples first, and replaying simple samples first. SeRe obtains 25% of the old samples from the sequence at a certain step length, and the sample range covers samples from difficult to simple. As can be seen from the figure, the effect of random replay is the worst, because it does not take into account the value of old samples in consolidating the learning of the model. The method of prioritizing the replay of medium samples and the method of prioritizing the replay of simple samples reduces the accuracy of the basic model because it is difficult for the model to learn old knowledge that has not been mastered or forgotten from these data, and ignores the importance of difficult samples. Although the priority replay of difficult samples performs better, it takes into account that difficult samples are easier for the model to forget, but it also ignores the usefulness of simple samples. The replay based on pixel frequency weighting performs better because it alleviates the problem of class imbalance. The uniform replay of difficult and easy samples of the present invention best simulates the difficulty distribution of old data and achieved the best performance in the experiment.
[0068] Comparison of the performance of selective learning on rare and common categories:
[0069] The performance of selective learning on rare categories is very outstanding. Take the new data ratio of 20% as an example: the mIoU accuracy of the background class is improved from 11.71 to 82.27, the accuracy of the town class is improved from 0 to 18.93, and the accuracy of the grass class is improved from 2.98 to 4.91. After pixel statistical analysis, we found that the proportion of pixels in these categories is small, which is why the basic model has poor inference accuracy in these categories. By sampling new data based on the method of uncertainty-based difficult samples, we tend to choose data with poor model learning effects when selecting data, and these data happen to contain more pixels of rare categories. Therefore, after selective learning, the performance of the model in rare categories is significantly improved.
[0070] Hyperparameter Analysis:
[0071] 1) The impact of selective learning ratio
[0072]
[0073]
[0074] After training the basic model using data A, we use the uncertainty-based difficult sample sampling method to select the new data that needs to be learned, and set the ratio of different new data to 20%, 40%, 60%, 80%, and 100%. It can be found that when the proportion of new data is between 20% and 60%, the accuracy of the model is significantly and stably improved. This shows that selective learning can use less data and quickly learn the rich knowledge in it, which not only saves the annotation cost, but also achieves rapid updates.
[0075] 2) Impact of playback ratio
[0076] After completing the selection of 20% new sample ratio, the old samples are sampled using the uniform sampling method of easy and difficult samples based on mIoU. Since storing old samples requires a lot of storage space, we limit the playback ratio to less than 25%, specifically divided into 5%, 10%, 15%, 20%, and 25%. When the amount of playback data is too small, such as 5%, 10%, and 15%, the accuracy of the model cannot be effectively improved and is unstable. When the model playback data reaches 25%, the accuracy of the model is significantly improved.
[0077] The beneficial effects of the present invention are as follows:
[0078] 1) Apply the incremental learning method to the task of remote sensing semantic segmentation, so that the model can be quickly iterated and updated in the context of remote sensing big data.
[0079] 2) We propose an incremental learning framework, SeRe, for semantic segmentation tasks. The model learns new knowledge and retains old knowledge at the same time through selective learning and replay. The selective learning and replay approach can not only allow the model to be updated quickly, but also alleviate catastrophic forgetting. Experiments show that this invention is superior to other current incremental learning methods.
[0080] 3) The present invention was evaluated on multiple data sets, and experiments proved that with 20% of new data, the present invention can achieve 80% of the joint training accuracy. Joint training is considered to be the upper limit of the accuracy of incremental learning. At the same time, the training speed is improved with the same accuracy.
[0081] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.
[0082] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".
[0083] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.
[0084] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A remote sensing image crop recognition method based on selective learning and playback, characterized in that: The following steps are involved: Obtain multiple crop remote sensing images and input them into the neural network model; The selective learning selects data with more learning value from the new data as the training set Dt of this stage, and annotates Dt, wherein Dt has paired images and annotation files (I t ,G t ), where I t Represents the W×H, G of an image t Represents the corresponding true label of the same size, i.e., W×H; In the learning stage t, the features of the image are extracted by the feature extractor f t (·) is obtained by classifier g t (·) Output the classification result; the segmentation result of the model for the image of the test set in stage t is recorded as The set of all parameters of the model after learning is recorded as θ t ; The playback performs reasoning and evaluation on the old data, calculates the segmentation accuracy IoU of a single image, sorts the difficulty of the old samples by IoU, and obtains these old samples by uniform sampling; Use the labeled samples to train the neural network model. After the loss between the prediction of the neural network model and the real label of the data is generated, back propagation is performed to update the parameters of the neural network model. Input the newly collected crop remote sensing images into the trained neural network model and output the crop recognition results; The data with more learning value uses the uncertainty measurement of the sample, and the measurement method of the uncertainty of the sample is as follows: represents the inference result of the model in stage t for unlabeled data, W, H, C represent the width, height and category of the image respectively, and the old model in stage t-1 for the new data D t The inference result is expressed as The old model's inferred pseudo-label for a pixel in a new image is: Since each pixel’s pseudo-label has its confidence, the uncertainty of each pixel is calculated based on it; the inverse of the confidence is taken as the uncertainty. The lower the confidence, the higher the uncertainty, and the higher its learning value. The uncertainty of the old model for a pixel in a new image is: Where w is the width, h is the height, c is the number of categories, and u is the value of the pixel uncertainty of this image. After calculating the uncertainty of all pixels in an image, the uncertainty of all pixels is summed to obtain the uncertainty of the sample: ∑U t-1→t (w,h,u)。 2. The method for crop recognition based on remote sensing images based on selective learning and playback according to claim 1, characterized in that: The calculation process of the IoU of the single image is as follows: The old data D t-1 As the input of the t-stage model, the output after model processing is Through one-hot encoding, the category with the highest inference confidence is taken as the pixel category to obtain the pseudo label: Old data D t-1 It is labeled data, denoted by G t-1 ; According to the pseudo label and the true label G t-1 Calculate the IoU of this image; calculate the IoU of each category in this image, and then find the mean of the IoU of all categories as the IoU of the image:
3. The method for crop recognition based on remote sensing images based on selective learning and playback according to claim 2, characterized in that: The higher the IoU value, the better the model's reasoning effect is, and the model has a better memory for this part of the data. Conversely, the lower the IoU value, the worse the model's reasoning effect is, and the model has forgotten this part of the data.
4. The method for crop identification based on remote sensing images based on selective learning and playback according to claim 3 is characterized in that: All images are sorted according to the IoU value, and then 20% of the samples are sampled for playback at a certain interval, and the playback ratio is 25%.
5. The method for crop identification based on remote sensing images based on selective learning and playback according to claim 1, characterized in that: The neural network model is a Deeplab-v3+ model, and ResNet-50 is used as the backbone network, and the backbone network uses a pre-trained model of ImageNet.
Citation Information
Patent Citations
All-weather video measurement method for visibility
CN101957309A
Multi-object tracking using correlation filters in video analytics applications
CN113950702A