Sparse sample enhanced sampling method for sample imbalance crop fine classification

Through the Sparse Sample Enhanced Sampling Method (MES), by intensively resampling sparse category samples and combining data augmentation operations, the problem of sample imbalance in crop fine classification is solved, and the classification performance of the model and the recognition ability of sparse categories are improved.

CN119942200APending Publication Date: 2025-05-06KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510017623.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In crop fine classification, the problem of sample imbalance causes deep learning models to tend to learn head category characteristics, weakening the overall performance of the model in actual scenarios.

Method used

A sparse sample enhancement sampling method (MES) is proposed. By intensively resampling sparse category samples and combining data enhancement operations, the training data distribution is balanced, the model biases over most classes, and the feature representation ability of sparse class samples is improved.

Benefits of technology

It effectively alleviates the limitations of sample imbalance on network performance, improves the classification performance of the model on unbalanced data sets, especially in identifying sparse categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942200A_ABST
    Figure CN119942200A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crop classification, in particular to a rare sample enhanced sampling method for sample imbalance crop fine classification, which comprises the following steps of: counting sample information, including the number of pixels of each category, sample distribution and category pixel frequency; calculating the resampling probability of each category by taking a softmax function as a prototype formula; selecting a category according to the category resampling probability obtained by calculation, and randomly and uniformly sampling from an image subset containing the category to obtain an image; checking whether the number of pixels of the target category contained in the obtained image meets the minimum pixel threshold value or not, and if not, re-obtaining the image meeting the condition from the image subset; data enhancement is carried out on the selected image, the diversity and feature space of samples are increased, and the images are input into the network for training and learning. According to the method, the limitation of sample imbalance on network performance can be effectively relieved, and the practical application of a deep learning technology in crop fine classification is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of crop classification, and in particular to a sparse sample enhancement sampling method for fine classification of sample imbalanced crops. Background Art

[0002] The widespread application of deep learning technology has promoted the rapid development of high-resolution remote sensing image coverage mapping and achieved remarkable results. However, different land cover types, such as forests, farmlands, urban land, etc., have significant differences in the distribution area in geographic space, resulting in the widespread problem of class imbalance in land cover mapping. The training effect of deep learning networks depends largely on massive and high-quality labeled data. The problem of data sample imbalance causes the network to be more inclined to learn head category features, weakening the overall performance of the network in actual scenarios. In agricultural landscapes, the problem of sample imbalance is equally thorny. Taking the fine classification of crops in plateau mountainous areas as an example, the cultivated land in plateau mountainous areas is fragmented and of various shapes, the planting types are rich and diverse, the crop planting methods are flexible, and the problem of crop category sample imbalance is serious. Therefore, it is urgent to explore effective solutions to deal with the challenges brought by sample imbalance to fine classification of crops.

[0003] The distribution of data samples with imbalanced categories is called long-tail distribution. A large number of studies have focused on the solution strategies for the long-tail distribution problem, which has promoted the development of deep learning technology. Deep long-tail learning methods can be divided into three categories according to their main technical characteristics: model improvement methods, category rebalancing methods, and information enhancement methods. Model improvement methods are mainly based on the idea of ​​category rebalancing, which improves performance by optimizing network modules and has comprehensive performance. However, in practical applications, specific module design and model training are required, which has high design costs and is difficult to avoid increasing the complexity of the model. Category rebalancing methods include resampling methods and category-sensitive learning methods, which achieve the effect of rebalancing category distribution by adjusting the number of samples at the sample level or adjusting the category weights in the loss function. Traditional deep learning network training ignores the problem of sample imbalance based on random sampling. For this reason, researchers have explored various sampling methods and designed specific sampling strategies based on sample distribution. Shen et al. proposed the "Class-Aware Sampling" sampling method to ensure that the probability of each class appearing in each batch is the same as much as possible. However, this method will cause the head class samples to be insufficiently sampled in multiple small batches. Recent studies have proposed a variety of adaptive sampling strategies, such as dynamic sampling rate. Wang et al. proposed a dynamic curriculum learning method (Dynamic CurriculumLearning, DCL) to achieve category rebalancing through dynamic sampling data. Although the above method alleviates the problem of category imbalance to a certain extent, due to the limited sample data of the minority class, it is still difficult for the model to learn sufficient feature information from the minority class. Information enhancement is to introduce additional information into model training to improve the performance of the model in long-tail learning. It mainly includes two methods: transfer learning and data augmentation (DA). Among them, the transfer learning method transfers knowledge from the source domain to the target domain to enhance the training of the image model. First, all long-tail samples are used for model pre-training, and then the model is fine-tuned on a more balanced training subset. In this way, the learned features are gradually transferred to the tail category, so as to obtain a more balanced performance in all categories. However, the limitation of transfer learning is that enough samples are needed for effective transfer. Data augmentation is a common technique in deep learning model training. It increases the diversity of data samples by performing random flipping, random cropping, random scaling, color transformation and other operations on images at the data and feature levels, thereby improving the generalization ability of the model. Ahn et al. first proposed CUDA (CUrriculum of Data Augmentation) method based on category-related data augmentation, which dynamically adjusts the data augmentation intensity according to the degree of category learning, provides a new paradigm for processing long-tail distribution data, and combines data augmentation with dynamic adjustment of category learning, providing an effective way to solve the problem of category imbalance.

[0004] The long-tail distribution problem stems from the serious imbalance in the distribution of training set data. Improving data distribution is one of the most intuitive ideas to solve this problem. Resampling, as the mainstream method for dealing with the long-tail problem, balances the category distribution by increasing or decreasing data samples, which will not interfere with the training of subsequent classifiers, and most resampling methods can flexibly adapt to the existing long-tail image classification model. Data enhancement technology performs well in improving the generalization ability of the model, but when it is applied alone to long-tail classification, the head class samples are dominant in number, and their features will cover the tail class features, resulting in the model being unable to learn accurate classification boundaries. Existing studies have attempted to combine resampling with data enhancement to alleviate the long-tail problem. However, these methods usually rely on specific algorithm support, have low adaptability, and are difficult to flexibly apply in different task scenarios. To this end, the present invention proposes a minority sample enhanced sampling method (Minority sample Enhanced Sampling, MES) for the sample imbalance problem in crop fine classification, combining the advantages of category resampling and data enhancement to provide an efficient and universal solution to solve the actual long-tail problem. Specifically, MES balances the distribution of training data by densely resampling rare class samples and combining data augmentation operations, alleviates the model's bias towards the majority class, and improves the feature representation ability of rare class samples, thereby improving the classification performance of deep learning models under unbalanced sample conditions. Compared with existing methods, the modular design of MES can be plug-and-play for semantic segmentation networks, improving recognition performance without increasing the number of parameters. Summary of the invention

[0005] The purpose of the present invention is to provide a sparse sample enhancement sampling method for fine classification of crops with sample imbalance. MES balances the distribution of training data by densely resampling rare category samples and combining data enhancement processing, alleviates the network's cognitive bias for the majority class, and improves the generalization performance of the model. MES is easy to implement and has strong adaptability. It can be integrated into the network model as a general sampler module for semantic segmentation tasks. To verify the effectiveness of MES, experiments were carried out on the Dali dataset and the barley remote sensing detection dataset. The results show that the performance of MES on the four benchmark networks of CNN and Transformer architectures is effectively improved, and finally its stability and reliability are verified by hyperparameter sensitivity analysis. The proposed method can effectively alleviate the limitations of sample imbalance on network performance and promote the practical application of deep learning technology in the fine classification of crops.

[0006] In order to achieve the above technical objectives and the above technical effects, the present invention is implemented through the following technical solutions:

[0007] A sparse sample enhancement sampling method for fine classification of sample imbalanced crops includes the following steps:

[0008] S1: Count the sample information of the training set, calculate the number of sample pixels of each category in the training set, and calculate the frequency of pixels of each category.

[0009] S2: The resampling probability of each category is calculated by using the softmax function as the prototype formula through the frequency of each category pixel, so that the lower the category pixel frequency, the higher the resampling probability. Through this mechanism, rare categories can obtain higher resampling probabilities, and the network's learning of rare categories can be strengthened.

[0010] S3: Select a category based on the calculated category resampling probability, and obtain a sample image from the image subset containing the category. Check whether the number of pixels of the target category contained in the acquired image meets the minimum pixel threshold, otherwise reacquire the image that meets the condition from the image subset. This ensures that the resampled image contains enough pixels of the target category, avoids invalid or inefficient resampling, and improves the efficiency of resampling of rare categories.

[0011] S4: Perform data enhancement on the selected images to increase sample diversity and feature space, and input them into the network for training and learning.

[0012] Through the above steps, the MES algorithm ensures that rare category samples are resampled more frequently, and at the same time improves sample diversity through data enhancement, thereby enhancing the network's learning ability for rare categories and ultimately improving the performance of the model on unbalanced data sets.

[0013] Furthermore, the step S1 specifically includes:

[0014] MES first analyzes the frequency of occurrence of each category to assign a sampling probability to each category, and the tail category will be assigned a higher sampling probability. In the pixel-level classification task, the instance object is a pixel. MES first counts the number of pixels in each category when calculating the sampling probability distribution. For any category C in the data set, its frequency f c It can be calculated based on the number of pixels of this category in the sample label:

[0015]

[0016] Where: (i, j) – pixel coordinates;

[0017] ——Indicator function, when the pixel coordinate (i, j) is of category c, the value is 1, otherwise it is 0;

[0018] H×W——image size;

[0019] N S - number of images;

[0020] C – Category;

[0021] (i,j)——pixel coordinates;

[0022] Furthermore, the step S2 specifically includes:

[0023] Calculate the resampling probability of each category, the sampling probability P of a certain category c c is defined as its frequency f c Function:

[0024]

[0025] Where: T——smoothness parameter;

[0026] C – total number of categories;

[0027] c′——Category c′

[0028] Therefore, the less frequent classes will have a higher sampling probability, and the sampling distribution is adjusted by the softmax function and the specified smoothness parameter. Parameter T controls the smoothness of the distribution: the higher the value of T, the higher the sampling probability P. c The more uniform the distribution, the lower the T value and the higher the probability of sampling rare categories.

[0029] Furthermore, the step S3 includes:

[0030] For each sample image, from the probability distribution c~P c Select the corresponding category and then sample an image from the data subset containing the category. The mathematical expression can be expressed as:

[0031] I c ~uniform(X S,c )(3)

[0032] Where: I c ——sampling image;

[0033] X S,c ——The dataset contains a subset of samples of category c;

[0034] After the tail category is selected, sample images are obtained from the image subset containing the category by uniform sampling. The purpose of uniform sampling is to ensure that the selection of samples is not biased towards certain specific samples in the subset, thereby reducing the bias in sample selection. Specifically, random selection is performed from all images containing the category, rather than repeatedly selecting certain specific images, so as to ensure the diversity of samples. After the image is selected, check whether the number of pixels of the target category contained in the image meets the preset minimum pixel threshold. The minimum pixel threshold parameter P is used to ensure that the resampled image contains enough target category pixels. This avoids too few target category pixels in the selected image, resulting in invalid or inefficient resampling. If the number of target category pixels in the selected image is insufficient, continue to reselect qualified samples from the image subset. Ensure that each resampling is effective, improve the efficiency and effect of resampling of rare categories, and enable the network to better learn the characteristics of these categories.

[0035] Furthermore, the step S4 specifically includes:

[0036] A series of data augmentation operations are performed on the target image and label data, such as rotation, adjusting brightness, contrast and sharpness, to increase the diversity and feature space of the samples. These enhanced sample data are then packaged in a data container and input into the network for training, thereby improving the generalization ability of the model and the learning effect on rare categories.

[0037] The MES method thus achieves resampling of images containing rare categories according to the scarcity of the sample categories, including sampling of rare category pixels and common categories, enriching the contextual information in the network model learning data; by increasing the number of samples for training rare categories, the model's recognition ability for rare categories is improved during the training process.

[0038] Beneficial effects of the present invention:

[0039] The design of MES of the present invention combines the functions of data loader and sampler, and as a component of deep learning framework, provides a flexible way to deal with data imbalance problem. First, MES obtains the pixel frequency of each category of samples by calculating the pixel quantity statistics of each category of samples in the training data set. Then, using the softmax function as the prototype formula, the resampling probability of each category is calculated using the category sample pixel frequency, so that the resampling probability of each category is inversely proportional to the pixel frequency. The sampler selects rare categories according to the resampling probability distribution, obtains image samples that meet the conditions, performs data enhancement on the samples, and finally inputs them into the model for training. The MES method solves the sample imbalance problem by dynamically adjusting the sample distribution during the network training process, improves the model performance, and combines data enhancement processing to prevent the risk of overfitting caused by resampling and enhance the generalization performance of the model.

[0040] The MES of the present invention uses a special sampling design to ensure that the network model resamples the rare category samples after traversing the training samples. In the resampling stage, by increasing the sampling probability of rare categories, these analogies can receive more attention during the training process, alleviating the training bias problem caused by data imbalance and improving the model's recognition ability for these categories.

[0041] After selecting the sample images, the MES of the present invention performs random data enhancement processing (such as image flipping, scaling, rotation, etc.) on the sample images before inputting them into the network model to expand the feature space of the training data, so that the model can see more diverse data during the training process, thereby enhancing the robustness and generalization ability of the model.

[0042] The MES of the present invention is flexible in design and includes three adjustable hyperparameters (sampling probability smoothness parameter T; resampling ratio α; minimum pixel threshold P) and adjustable data enhancement methods. In practical applications, it can be adjusted according to the specific data set conditions to adapt to diverse data sets and maintain its effectiveness and versatility in a wide range of application scenarios.

[0043] The MES of the present invention automates the resampling and data enhancement logic, making the network's learning of rare samples more efficient and systematic. MES does not require additional data processing or specific network design. As a network sampler, it solves the sample imbalance problem at the data level and can be plug-and-play for semantic segmentation networks, alleviating the limitations of data imbalance on network performance.

[0044] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0046] Figure 1 Schematic diagram of manual labeling for the Dali dataset;

[0047] Figure 2 This is a schematic diagram of pixel statistics of the Dali dataset training sample categories;

[0048] Figure 3 Schematic diagram of manual labeling for the barley remote sensing detection dataset;

[0049] Figure 4 This is a statistical diagram of pixel categories of training samples of the barley remote sensing detection dataset;

[0050] Figure 5 This is a schematic diagram of the experimental results of the Dali dataset;

[0051] Figure 6 Schematic diagram of the curve of training loss of each method changing with the iteration cycle;

[0052] Figure 7 This is a diagram showing the curve of IoU changing with iteration period during training of the vegetable class (tail class);

[0053] Figure 8 This is a schematic diagram of the experimental results of the barley remote sensing dataset;

[0054] Fig. 9 This is a schematic diagram of the category resampling frequency corresponding to different T values ​​of the Dali dataset;

[0055] Fig.10 Schematic diagram of the number of sample pixels (T = 0.05) corresponding to different a values ​​in the Dali dataset;

[0056] Fig.11 Schematic diagram of the number of sample pixels (a=1) corresponding to different T values ​​of the Dali dataset. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] Example 1

[0059] In order to solve the problem of relatively scarce tail class samples in training data, a resampling scheme for rare samples is designed. MES first analyzes the frequency of occurrence of each category to assign a sampling probability to each category, and the tail category will be assigned a higher sampling probability. In the pixel-level classification task, the instance object is a pixel. When calculating the sampling probability distribution, MES first counts the number of pixels in each category. For any category C in the data set, its frequency f c It can be calculated based on the number of pixels of this category in the sample label:

[0060]

[0061] Where: (i, j) – pixel coordinates;

[0062] ——Indicator function, when the pixel coordinate (i, j) is of category c, the value is 1, otherwise it is 0;

[0063] H×W——image size;

[0064] N S - number of images;

[0065] C – Category;

[0066] (i,j)——pixel coordinates;

[0067] The sampling probability P of a certain category c c is defined as its frequency f c Function:

[0068]

[0069] Where: T——smoothness parameter;

[0070] C – total number of categories;

[0071] c′——Category c′

[0072] Therefore, the less frequent classes will have a higher sampling probability, and the sampling distribution is adjusted by the softmax function and the specified smoothness parameter. Parameter T controls the smoothness of the distribution: the higher the value of T, the higher the sampling probability P. c The more uniform the distribution, the lower the T value, and the higher the probability of sampling rare categories. For each sample image, select the corresponding category from the probability distribution c~P, and then sample an image from the data subset containing this category. It can be expressed as:

[0073] I c ~uniform(X S,c ) (3)

[0074] Where: I c ——sampling image;

[0075] X S,c ——The dataset contains a subset of samples of category c;

[0076] After selecting the tail category, uniform sampling is used to ensure that the sample selection is not biased towards certain specific samples in the subset, which helps to reduce the bias in sample selection. Formula (2) allows the model to resample images containing rare categories according to the scarcity of the sample category, achieving a balanced sampling effect for each category. Rare category samples usually coexist with multiple common category samples in a single image. Therefore, when resampling, not only rare category pixels are sampled, but also common categories are sampled, enriching the context information in the network model learning data.

[0077] The core idea of ​​MES is to enable the model to pay more attention to the rare categories in the training set by resampling according to the frequency of rare categories. The main purpose of resampling is to increase the number of samples of rare categories and improve the model's recognition ability of rare categories during training. In order to avoid too few pixels of rare categories in the resampled sample image, the minimum sampling pixel parameter P is introduced to ensure that the tail category features of the resampled image are sufficient and to avoid invalid or inefficient resampling. The designed threshold P is adjustable to ensure that it can be flexibly applied to different data sets.

[0078] To avoid overfitting, data augmentation is often used in the network training process. Common data augmentation methods include random image flipping, scaling, rotation, cropping, etc. These methods can expand the feature space, increase intra-class diversity, and improve the generalization ability of the model. However, the strategy of applying the same intensity of data augmentation to all categories is not applicable to long-tail datasets, because the imbalance of category distribution in long-tail datasets will cause the effect of feature space expansion to be inconsistent between different categories. To this end, MES applies data augmentation operations to resampled rare category samples to expand the feature space of rare categories, thereby improving the balance of the feature space.

[0079] The detailed process of executing the rare sample enhancement sampling method is as follows:

[0080]

[0081] MES algorithm process:

[0082] 1. Initialization: MES first reads the data set and related parameters from the configuration file to build the basic data set object. Then it reads the statistical information of the training set samples and calculates the pixel frequency f of each category according to formula (1): c . Using f c And the given smoothness parameter T, the sampling probability P of each category is calculated by formula (2): c .

[0083] 2. Obtain (rare category) samples: according to the probability distribution P c Select a category and randomly select a sample image from the samples containing the category using the previously obtained sample statistics, and check whether the number of pixels of the target category in the sample meets the minimum pixel threshold P. If the number of pixels meets the minimum pixel threshold P, perform random data augmentation on the sample image and return it; otherwise, try to obtain a better sample.

[0084] 3. Data enhancement: Perform a series of random data enhancement operations on the selected sample images, including rotation, adjustment of brightness, contrast, and sharpness. Data enhancement can avoid the risk of overfitting caused by resampling and improve the generalization ability of the model.

[0085] 4. Input the processed sample images into the network model for training.

[0086] Example 2

[0087] This study uses the Dali dataset and the Barley Remote Sensing Dataset (BRSD) published on the Ali Tianchi platform as data sources. Both datasets have the characteristics of complex planting structures in typical plateau mountainous areas, and crop categories show a long-tail distribution. The crops are finely classified and compared through four benchmark networks, a benchmark network combined with data enhancement (+DA), and a benchmark network combined with rare sample enhanced sampling (+MES), which fully verifies the effectiveness of the method of the present invention. Finally, through the hyperparameter sensitivity study of MES, the influence of MRS hyperparameter settings on the classification results is analyzed.

[0088] Data Source

[0089] Dali Dataset

[0090] The Dali dataset was collected from Longkan Village, Dali Town, Dali Bai Autonomous Prefecture, Yunnan Province. The cultivated land plots are fragmented, and the crop planting types are diverse and scattered. The drone image data was collected on August 3, 2022, using a DJIPHANTOM 4RTK drone platform equipped with a 1-inch CMOS sensor and 8.8 mm and 24 mm focal length lenses, with an effective pixel of 20 million. The spatial resolution of the orthophoto of the study area is 0.0285m. The study area is divided into a training set and a test set. The image size of the training set is 1850×9655 pixels, and the image size of the test set is 1001×1001 pixels. According to the field survey data, the two images were manually labeled using Labelme software. The labels are divided into 8 categories of samples: corn, rice, celery, onion, green vegetables, lettuce, coriander and background. The label image is a single-channel grayscale image with a bit depth of 8bit. The training set contains 1432 images and the test set contains 676 images after the cropping process. The manual labels of the Dali dataset and the statistical distribution of pixels in the training sample categories are shown in the figure below: Figure 1 , Figure 2 shown.

[0091] Barley remote sensing detection dataset

[0092] The barley remote sensing detection dataset comes from the 2019 Ali Tianchi County Agricultural Brain AI Challenge. The data was collected from a farmland area in Xingren City, Qianxinan Buyi and Miao Autonomous Prefecture, Guizhou Province. The dataset includes five types of sample elements: background, flue-cured tobacco, corn, coix seed rice, and buildings. The corresponding labels are as follows: Figure 3As shown. In the experiment, images 1 and 2 are used for training, and image 3 is used for testing. Image 1 and image 2 are slidably cropped with a size of 512×512 pixels, and the overlap is not set (0 pixel overlap). During the cropping process, images without features are screened out, and 5926 images are finally obtained, of which 80% are used for training and 20% are used for verification. An image without features refers to a cropped image that does not contain any valid features in the original image and its label image. It usually appears in the edge area of ​​the original image in a rectangular crop. The use of such images for training will result in data redundancy. Image 3 of the test set is slidably cropped with a size of 512×512 pixels and an overlap of 128 pixels, and 6517 images are obtained by cropping. The pixel distribution of the image and training sample categories of the barley remote sensing detection dataset are shown as follows: Figure 3 , Figure 4 shown.

[0093] Evaluation indicators

[0094] The crop classification results are evaluated using pixel-level accuracy. The present invention sets up two main evaluation indicators of semantic segmentation on crop classification accuracy, namely mean Intersection over Union (mIoU) and F1 score, which are calculated as follows:

[0095]

[0096] Among them, TP represents true positive examples, TN represents true negative examples, FP represents false positive examples, FN represents false negative examples, and mIoU and mF1 are the average values ​​of IoU and F1 scores of all categories, respectively.

[0097]

[0098] Among them, C is the total number of categories, F1 i F1 score for class i.

[0099] In an imbalanced dataset, the mF1 score (Macro F1) can better reflect the classification performance of the model. The mF1 score is an equal weighting of the precision and recall of a class. The mF1 score gives the same weight to each class, regardless of the number of samples in the class. Therefore, the mF1 score can better reflect the overall performance of the model in each class, especially the class with fewer samples.

[0100] Experimental results analysis

[0101] In order to verify the applicability of the proposed method in deep learning semantic segmentation networks, the crop classification results of MES under the same parameter conditions in different benchmark models were analyzed through experiments. The benchmark models include typical CNN architecture networks (Deeplablv3+ and SegNeXt) and typical Transformer architecture networks (SegFormer and SwinTansformer).

[0102] Experimental software environment: Windows operating system, PyTorch 1.10.1 deep learning framework and Python3.8 development environment. Experimental hardware environment: processor Intel Core i7-13700KF, graphics card NVIDIA GeFrce RTX 3090, video memory 24G, running memory 64G. Experimental parameter setting: The training area data is randomly divided into training set and validation set according to 8:2. The loss function is cross entropy loss (CE). The network is trained using the AdamW optimizer with a momentum of 0.9, an initial learning rate of 0.00006, and a weight decay of 0.01. The training batch size is 16, the number of threads is 4, the total training iteration cycle is 100 epochs, and the optimal mIoU weight is always saved during training. The final results of the experiment are obtained by accuracy evaluation of the test set.

[0103] Experimental analysis of the Dali dataset

[0104] In view of the long tail degree, data size and image input size of the Dali dataset, the MES hyperparameters are set as follows in the experiment: T = 0.05; α = 1; P = 100000. The accuracy evaluation results are shown in Table 1.

[0105] Table 1. Accuracy evaluation of experimental results of Dali dataset

[0106]

[0107]

[0108] DA represents the data augmentation experimental results.

[0109] As can be seen from Table 1, data enhancement can improve the impact of sample imbalance on classification results to a certain extent. Data enhancement processing can deepen the network's learning of crop category characteristics and enhance the generalization performance of the model. The methods of the present invention all improve the classification performance of the benchmark models. Compared with the simple data enhancement method, the method of the present invention obtains +1.10%, +0.68%, +1.30%, +1.47% mIoU gains and +1.09%, +0.27%, +1.08%, +0.76% mF1 gains in the four benchmark models, verifying the effectiveness and versatility of the method of the present invention. It can be seen from the IoU performance of each category that the gain of the tail category is significantly improved, and the overall improvement does not only come from the tail category samples. Figure 5 The experimental results of the method of the present invention on the Dali dataset are shown. Figure 5 It can be seen that the proposed method effectively improves the recognition effect of each benchmark model. The red rectangle highlighted part in the figure shows the processing of the tail class details by each method. The proposed method effectively reduces the category confusion of each benchmark method, which verifies the effectiveness of the proposed method.

[0110] The training loss change curve can reflect the training dynamics of the model. Figure 6 The training loss curves of each method are displayed as they change with the iteration cycle. The loss curves can be used to observe the convergence of the network model. Figure 6 It can be seen that the method of the present invention accelerates the model convergence in each benchmark network, and the loss changes are smoother and more stable, which verifies the effectiveness and versatility of the method of the present invention. Minority class samples in the long-tail distribution are difficult to train effectively and are usually the bottleneck of model convergence. The rapidly decreasing loss indicates that the model can better adapt to long-tail data and optimize the learning weights of long-tail categories. By introducing a category-balanced learning mechanism, the contribution of minority class features to the loss can be captured faster. In order to observe the effect of MES on tail categories during network training, take the green vegetable category as an example, Figure 7 The curve of IoU change of green vegetable training with iteration cycle is shown. The results show that MES can enable the network model to learn rare class features earlier, preventing the network from accumulating head class prediction bias too early.

[0111] Experimental analysis of barley remote sensing detection dataset

[0112] The present invention verifies the generalization performance of the method of the present invention through the barley remote sensing detection dataset. Considering the long tail layer of the dataset, the amount of data, and the image input size, the MES hyperparameters in the experiment are set as follows: T = 0.1; α = 0.2; P = 100000. The accuracy evaluation results are shown in Table 2.

[0113] Table 2 Accuracy evaluation of experimental results of barley remote sensing detection dataset

[0114]

[0115] DA represents the data augmentation experimental results.

[0116] It can be observed from Table 2 that the method of the present invention obtains mIoU gains of +2.84%, +1.07%, +4.16%, and +2.69% and mF1 gains of +1.83%, +0.62%, +2.84%, and +1.8% on the original benchmark model. It is worth noting that simple random data augmentation causes a decrease in classification accuracy on some networks, and all of them occur on Transformer-based network models. The Transformer model relies on a global attention mechanism to capture the long-distance dependency characteristics of the data. However, when the data augmentation strategy does not fully consider the long-tail distribution, the generated enhanced samples may be biased towards the majority class, further expanding the dominance of the majority class features, resulting in a decrease in the model's discriminative ability. In comparison, the barley remote sensing detection dataset has more sufficient samples and a lighter long-tail problem. Further experiments on this dataset have demonstrated the effectiveness and good generalization ability of the method of the present invention in this scenario.

[0117] Figure 8 An example of experimental results in the barley remote sensing dataset is shown. As can be seen from the figure, the MES method effectively improves the internal integrity and boundary contour effects of each benchmark network for land parcel recognition. For the recognition of buildings (tail category), the recognition results of each method showed different levels of "salt and pepper" phenomenon, among which the MES method had the least hole phenomenon, which further verified that the introduction of the MES mechanism can improve the classification ability of rare samples.

[0118] Hyperparameter sensitivity analysis

[0119] To be applicable to data sets with various distributions, MES sets three adjustable hyperparameters. The hyperparameter T is used to adjust the smoothness of the sampling probability distribution; the hyperparameter α is used to adjust the resampling ratio of MES; and the hyperparameter P is used to control the minimum number of pixels of the rare class contained in the MES sampled image. The present invention refers to the parameters used to control the minimum number of retained sample pixels in the Online Hard Example Mining (OHEM) method, and sets the default value P = 100000. In order to analyze the sensitivity of parameters T and α in MES, the present invention conducts experiments on the Dali dataset using SegFormer as the benchmark network, and analyzes the effects of different parameter settings on crop classification results by controlling variables. The experimental results are shown in Tables 3 and 4.

[0120] Table 3 Experimental results of MES hyperparameter analysis of SegFormer in Dali dataset (α=1)

[0121]

[0122]

[0123] Table 4 Experimental results of MES hyperparameter analysis of SegFormer in Dali dataset (T=0.05)

[0124]

[0125] As can be seen from Tables 3 and 4, when the resampling ratio α of MES is 1, the sampling probability smoothness parameter T can maintain a significant overall performance gain in the interval [0.01, 0.5]; when the sampling probability smoothness parameter T is 0.05, the resampling ratio α also shows a stable performance improvement in the interval [0.2, 2]. The selection of the sampling probability smoothness parameter T and the resampling ratio α can be reasonably set based on an intuitive strategy: for example, T can be adjusted according to the degree of long tail by maximizing the number of resampled pixels, while α is determined according to the size of the data set. On the basis of the default strategy, the parameters are allowed to be adjusted within a certain range while still maintaining the effectiveness of MES. Hyperparameter sensitivity analysis further verifies the stability of the method of the present invention, indicating that it has reliability and scalability in practical applications.

[0126] Fig. 9 and Fig.10 The resampling frequency and number of sampled pixels of the categories under different T values ​​are shown respectively. As can be seen from the figure, the smaller the T value, the steeper the sampling probability distribution. The MES mechanism makes the sampling probability distribution of the tail category samples approximately inversely proportional to the sample frequency distribution, so that the rare categories occupy a larger proportion in the resampling. Fig.11 The amount of resampled pixels for different alpha values ​​during network training is shown. Fig.11 It can be found that under the action of MES, as the number of resampled samples (resampling ratio) increases, the distribution of training samples gradually tends to be balanced. Fig.10 It can be seen from (red column) that when T = 0.05, the distribution of the sampled samples is closest to the equilibrium state, and the classification accuracy also reaches the highest.

[0127] The original intention of setting MES hyperparameters is to adapt to a variety of data sets. However, when actually selecting hyperparameters, it may be necessary to conduct a preliminary analysis in combination with the sample distribution. Based on this research, further research can be conducted to integrate the analysis strategy with the MES mechanism to explore adaptive MES methods suitable for diverse data sets and task scenarios.

[0128] In summary, in order to solve the problem of sample imbalance in crop classification, the present invention proposes a rare sample enhancement sampling method (MES). As a universal sampler, MES can be plug-and-play for semantic segmentation networks. This method solves the problem of classification pixel deviation caused by sample imbalance in the data set by designing the resampling probability to be approximately inversely proportional to the frequency of the category pixels, and combines the data enhancement method to improve the generalization ability of the model. The present invention takes the fine classification of crops in plateau mountainous drone remote sensing images as an example and conducts experiments on two data sets. The results show that the method of the present invention shows good performance improvement in four semantic segmentation networks (including CNN and Transformer architectures) without increasing the number of network parameters. Compared with the use of data enhancement methods alone, MES does not require additional data processing, and the implementation process is simple and efficient. The hyperparameter design of MES is flexible and can adapt to a variety of data sets and application environments. The present invention verifies its stability and reliability through hyperparameter sensitivity analysis. In the future, more rigorous hyperparameter analysis experiments will be further carried out, and the MES method of adaptive data sets will be implemented in combination with a variety of data sets (such as satellite remote sensing images). The implementation of the method of the present invention does not depend on a specific network architecture and can be widely used in the field of semantic segmentation, especially in application scenarios such as crop classification and remote sensing image analysis, which helps to promote the application of deep learning technology in agricultural fine classification tasks.

[0129] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A sparse sample enhancement sampling method for fine classification of sample imbalanced crops, characterized in that: The following steps are involved: S1: Count the sample information of the training set, calculate the number of sample pixels of each category in the training set, and calculate the frequency of pixels of each category; S2: The resampling probability of each category is calculated by using the softmax function as the prototype formula through the pixel frequency of each category; S3: Select a category based on the calculated category resampling probability, and obtain a sample image from the image subset containing the category; check whether the number of pixels of the target category contained in the obtained image meets the minimum pixel threshold, if so, proceed to S4, otherwise re-acquire an image that meets the conditions from the image subset; S4: Perform data enhancement on the selected images to increase sample diversity and feature space, and input them into the network for training and learning.

2. The method for enhanced sampling of rare samples for fine classification of sample-imbalanced crops according to claim 1, characterized in that: The step S1 specifically includes: MES first analyzes the frequency of occurrence of each category to assign a sampling probability to each category, and the tail category will be assigned a higher sampling probability; in the pixel-level classification task, the instance object is a pixel, and MES first counts the number of pixels of each category when calculating the sampling probability distribution; for any category C in the data set, its frequency f c It can be calculated based on the number of pixels of this category in the sample label: Where: (i, j) – pixel coordinates; ——Indicator function, when the pixel coordinate (i, j) is of category c, the value is 1, otherwise it is 0; H×W——image size; N S - number of images; C – Category; (i,j)——Pixel coordinates.

3. The sparse sample enhancement sampling method for fine classification of sample imbalanced crops according to claim 1, characterized in that: The step S2 specifically includes: Calculate the resampling probability of each category, the sampling probability P of a certain category c c is defined as its frequency f c Function: Where: T——smoothness parameter; C – total number of categories; c′——Category c′ The sampling distribution is adjusted by the softmax function and the specified smoothness parameter; the parameter T controls the smoothness of the distribution: the higher the T value, the higher the sampling probability P. c The more uniform the distribution, the lower the T value and the higher the probability of sampling rare categories.

4. The method for enhanced sampling of rare samples for fine classification of sample-imbalanced crops according to claim 1, characterized in that: The step S3 specifically includes: For each sample image, from the probability distribution c~P c Select the corresponding category and then sample an image from the data subset containing the category; it can be expressed as follows through mathematical expressions: I c ~uniform(X S,c ) (3) Where: I c ——sampling image; X S,c ——The dataset contains a subset of samples of category c; After the tail category is selected, sample images are obtained from the image subset containing the category through uniform sampling; uniform sampling ensures that the selection of samples is not biased towards certain specific samples in the subset, thereby reducing the bias in sample selection; Uniform sampling includes: randomly selecting from all images containing the category, after selecting the image, checking whether the number of pixels of the target category contained in the image meets the preset minimum pixel threshold; the minimum pixel threshold parameter P is used to ensure that the resampled image contains enough target category pixels; to avoid too few target category pixels in the selected image, resulting in invalid or inefficient resampling; if the number of target category pixels in the selected image is insufficient, continue to reselect qualified samples from the image subset; until the number of pixels meets the requirement, perform S4; ensure that each resampling is effective, improve the efficiency and effect of resampling of rare categories, and enable the network to better learn the characteristics of these categories.

5. The method for enhanced sampling of rare samples for fine classification of sample-imbalanced crops according to claim 1, characterized in that: The step S4 specifically includes: performing a series of data enhancement operations on the target image and label data, including but not limited to rotation, adjusting brightness, contrast and sharpness, to increase the diversity and feature space of the samples; these enhanced sample data are then packaged in a data container and input into the network for training, thereby improving the generalization ability of the model and the learning effect on rare categories.