High-resolution oil palm tree identification method and system based on noise label learning

By using noise label learning, high-resolution mapping, weak-supervised learning and Bayesian inference methods in high-resolution remote sensing image processing, the problem of mismatch between label data and image resolution is solved, and the accuracy of oil palm tree recognition and the generalization ability of model are significantly improved.

CN120126017APending Publication Date: 2025-06-10SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510091770.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has problems in the processing of high-resolution remote sensing image, which leads to insufficient classification accuracy and generalization capabilities, especially in complex terrain and variable terrain areas that are difficult to accurately identify.

Method used

Using a noise label learning method, 10-meter resolution images are learned through 100-meter resolution oil palm tree labels, combining high-resolution mapping, weak-supervised learning and Bayesian reasoning to improve the accuracy of recognition and the generalization ability of the model.

Benefits of technology

It significantly improves the accuracy of oil palm tree recognition and the adaptability and generalization capabilities of the model, and can more accurately identify and classify oil palm trees, maintain high accuracy even in complex terrain and variable terrain areas, reducing the dependence on high-resolution labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126017A_ABST
    Figure CN120126017A_ABST
Patent Text Reader

Abstract

The invention relates to a remote sensing image processing and ground object identification technology, in particular to a high-resolution oil palm tree identification method and system based on noise label learning. The method comprises the following steps: acquiring remote sensing image data and label data, and performing mask mapping and preliminary cutting processing on the remote sensing image data and the label data to obtain image data after preliminary cutting; pre-processing the image data after the preliminary clipping to obtain a plurality of image blocks and a plurality of corresponding mask images which are clipped again and subjected to label screening, and taking the image blocks and the corresponding mask images as a training set; using the training set to train a U-shaped convolutional neural network model; outputting a prediction mask graph through the trained model, and identifying whether each pixel in the input image belongs to the region of interest; and carrying out post-processing on a model prediction result, and splicing to generate a full-size image map. According to the method, the generalization ability of the oil palm tree identification model is enhanced, and the adaptability and generalization ability of the model in image identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to remote sensing image processing and ground object recognition technologies, in particular to a high-resolution oil palm tree recognition method and system based on noisy label learning. Background Art

[0002] In the prior art, large-scale high-resolution land cover mapping has become an important technical means for surface monitoring, urban planning, and ecological protection. However, current methods face significant challenges in the resolution matching problem between labeled data and remote sensing images, resulting in insufficient classification accuracy and generalization ability.

[0003] The closest prior arts include the Paraformer framework proposed by Li et al., the SinoLC-1 project, and the L2HNet network. The Paraformer framework proposed by Li et al. combines a convolutional neural network (CNN) and a Transformer model, and effectively improves the performance of land cover mapping through the synergistic effect of local and global context information. This technology has achieved remarkable results in the fusion of low-resolution labels and high-resolution images. However, when processing extremely high-resolution remote sensing images, the computational complexity of Paraformer increases significantly, especially in areas with complex terrain and variable ground objects, where there are performance bottlenecks and it is difficult to accurately capture detailed features. SinoLC-1 is a project for constructing a 1-meter resolution land cover map. It generates training labels by integrating multi-source data (such as OpenStreetMap and high-resolution remote sensing images), and uses a weakly supervised module and a self-supervised loss function to optimize the model. However, the adaptability of this project is limited, and the classification effects for special terrains (such as mountains and wetlands) and special ground object categories (such as sparse vegetation) are not good. L2HNet is a network for low-resolution label to high-resolution mapping. By a confidence region selection module and a low-to-high loss function, it significantly improves the problem of resolution mismatch between labels and images. The test results on a small-scale dataset are excellent, but its adaptability and stability under different climate and geographical conditions still need to be verified. In addition, this method has deficiencies in the fusion of multi-source data and the generalization ability of complex scenarios. Summary of the Invention

[0004] To solve the problems existing in the prior art, the present invention proposes a high-resolution oil palm tree recognition method and system based on noisy label learning. It learns the noisy labels for 10-meter resolution oil palm tree recognition through 100-meter resolution oil palm tree labels, and introduces noisy label learning (LNL), high-resolution mapping (HRC), and weakly supervised learning (WSL).

[0005] And Bayesian Inference (BI) improves the accuracy of recognition, reduces confirmation bias, and significantly enhances the ability of the model to perform image recognition, thereby enhancing the generalization ability of the oil palm tree recognition model and improving the adaptability and generalization ability of the model in image recognition.

[0006] The recognition method of the embodiment of the present invention is implemented by the following technical solutions: A high-resolution oil palm tree recognition method based on noise label learning, including the following steps:

[0007] S1. Obtain remote sensing image data and label data, and perform mask mapping and preliminary cropping processing on them to obtain the preliminarily cropped image data;

[0008] S2. Preprocess the preliminarily cropped image data to obtain a number of image patches and corresponding mask maps after re-cropping and label screening as the training set;

[0009] S3. Use the training set obtained in step S2 to train the U-shaped convolutional neural network model; the trained model outputs a predicted mask map to identify whether each pixel in the input image belongs to the region of interest;

[0010] S4. Post-process the results predicted by the model and splice them to generate a full-size image map.

[0011] The recognition system in the embodiment of the present invention is implemented by the following technical solutions: A high-resolution oil palm tree recognition system based on noise label learning, including:

[0012] A data acquisition module for obtaining remote sensing image data and label data, and performing mask mapping and preliminary cropping processing on them to obtain the preliminarily cropped image data;

[0013] A preprocessing module for preprocessing the preliminarily cropped image data to obtain a number of image patches and corresponding mask maps after re-cropping and label screening as the training set;

[0014] A model training and prediction module for training the U-shaped convolutional neural network model using the training set; the trained model outputs a predicted mask map to identify whether each pixel in the input image belongs to the region of interest;

[0015] A splicing module for post-processing the results predicted by the model and splicing them to generate a full-size image map.

[0016] The present invention has the following advantages and positive effects compared with the prior art:

[0017] 1. Due to the adoption of the Label Noise Learning (LNL) method, combined with the processed labels, it brings better prediction results and model stability. Existing oil palm tree recognition models have low accuracy because they can only learn based on existing labels. In contrast, the present invention introduces a label noise learning module, which uses the preprocessed dataset as label noise for learning, significantly improving the prediction accuracy from 100-meter recognition unit to 10 meters.

[0018] 2. Due to the adoption of the High-Resolution Cartography (HRC) method, the existing data is preprocessed, greatly improving the resolution of the image, making it correspond to the pixels of satellite images, and bringing better model training effects. The gdal library is used to cut the image, and the image is processed by the interpolation method. Without manually relabeling the labels, batch preprocessing of the data also reduces the manual intervention link, making the data preprocessing link more convenient.

[0019] 3. Due to the adoption of the Weak Supervised Learning (WSL) method, it brings stronger feature expression ability and generalization ability to the target domain. By adopting the generalized labels, the feature representation ability of the target model is improved, the error caused by the processed labels is reduced, and the adaptability of the model in the target domain is improved.

[0020] 4. Due to the adoption of the Bayesian Inference (BI) method, non-target areas are effectively extracted from the land cover data, and the interference of these areas is accurately excluded, so as to focus on the areas where oil palm trees are planted. This innovation not only improves the accuracy of the oil palm tree recognition model, but also ensures that the model focuses on the target area, avoiding misclassification that may be caused by non-target areas, and significantly improving the effectiveness and reliability of the model in practical applications. In addition, the Bayesian inference method also optimizes the adaptability of the model to complex terrain and variable ground object areas. Especially in high-resolution remote sensing images, it can more accurately identify and classify oil palm trees, and can maintain high accuracy even near non-target areas, thus greatly improving the robustness and practicality of the model.

[0021] In summary, the present invention combines label noise learning, high-resolution cartography, weak supervised learning, and Bayesian inference. Compared with existing oil palm tree recognition models, it significantly improves the detection accuracy and generalization ability of the model, and can better handle the situation where the resolution of label data is lower than that of remote sensing images. Especially in the scenarios of complex terrain and variable ground object areas, it improves the fine-grained processing ability of the classification boundary and the recognition accuracy of oil palm trees under complex geographical conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flowchart of the high-resolution oil palm tree recognition method in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0023] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0024] Embodiment

[0025] This embodiment proposes a high-resolution oil palm tree recognition method based on noise label learning, aiming to achieve higher-precision recognition of cash crops in the target area, improve the production efficiency of cash crop planting, increase economic benefits, and promote local development. This method is achieved through noise label learning (LNL), high-resolution cartography (HRC), U-Net semantic segmentation model, weakly supervised learning (WSL), and Bayesian inference.

[0026] Based on the existing 100m-resolution oil palm labels, this embodiment proposes a high-resolution cartography method, which realizes the 10m high-resolution recognition of oil palm trees in the target area through interpolation to generate noisy labels, noise label learning, weakly supervised learning, and Bayesian inference.

[0027] Figure 1 The overall process of the high-resolution oil palm tree recognition method is shown. Starting from obtaining multi-source remote sensing image data of the Malaysia and Indonesia regions, through data preprocessing and feature extraction, then using the U-Net segmentation model for training, and finally performing data post-processing steps and stitching generation, high-precision recognition results are obtained, covering four major parts: data acquisition, data preprocessing, model training, post-processing, and stitching generation.

[0028] S1. Obtain remote sensing image data and label data, and perform mask mapping and preliminary cropping processing on them to obtain the preliminarily cropped image data.

[0029] In this step, multi-source remote sensing image data of the Malaysia and Indonesia regions are downloaded from Google Earth Engine, specifically including Sentinel-2 image data and annual oil palm area dataset data; oil palm labels of Indonesia and Malaysia are downloaded from AOPD Data.

[0030] Sentinel-2 image (Sentinel-2) data: sourced from the Sentinel-2 satellite, including a total of 6 bands: Red (R), Green (G), Blue (B), Near Infrared (NIR), Short Wave Infrared 1 (SWIR1), and Short Wave Infrared 2 (SWIR2), with a resolution of 10 meters; the size of each image data is 13568×13568, and the file format is tif file.

[0031] Annual Oil Palm Area Dataset Data (AOPD data): On the annual oil palm area dataset data, the regions of interest (i.e., the regions where oil palm is planted, with a value range of 1) and non - interest regions (value range of 0) are labeled with the corresponding classification mask. The spatial resolution is 100 meters, and the size is 283929×14747 pixels.

[0032] The resolution of the oil palm label is 100m, the size is 283929×14747 pixels, the value is 0 or 1, and the file format is tif file.

[0033] Mask Mapping & Initial Cropping: Through georegistration technology, align the mask data with the original Sentinel - 2 image; the data after initial cropping can be used in conjunction with high - resolution images to provide pixel - level supervision for subsequent models.

[0034] To ensure that each label map can correspond to each satellite map, crop the oil palm label map to make it the same size as the satellite map. Generate a shp file from the satellite image using the gdal library, and use the shp file to crop the label image, making the coordinates of each map the same as those of the corresponding satellite map, and save the new label as a png file.

[0035] S2. Pre - process the image data after initial cropping to obtain several image patches and corresponding mask maps that have been cropped again and label - screened as the training set.

[0036] The size of the image data after the initial cropping process in step S1 is 13568×13568. However, this resolution is still too high and is not conducive to subsequent calculation processing and model training. Therefore, it still needs to be pre - processed.

[0037] Re - cropping & Label Screening: To avoid memory shortage or excessive computational overhead when processing large - size images, continue to crop the initially cropped image patches into smaller sub - images, each with a size of 512×512 pixels, to obtain the cropped image patches. Perform mask screening on the cropped image patches, and only retain the image samples where the proportion of the region of interest (mask value of 1) exceeds a preset proportion value (in this embodiment, the preset proportion value is 30%) to avoid interference of invalid data on the model.

[0038] Use the interpolation method to change the 100m - resolution label map to 10m - resolution, magnify the pixels, so that the image samples after mask screening and the pixels of the satellite image correspond one by one. Since the resolution of the target result is 10m, and the input label is a 10m - resolution label obtained from the 100m - resolution label by the interpolation method, it can be considered that this label is a noisy label, that is, a noisy label.

[0039] Finally, the number of effective samples after re - cropping and label screening is 8,409 image patches and the corresponding 8,409 mask images, which are used as the training set. Each image patch has rich multi - spectral features and corresponds to a high - resolution mask, ensuring that the model training has weak - supervision ability.

[0040] S3. Use the training set obtained in step S2 to train the U - shaped convolutional neural network model (i.e., the U - Net model); through the trained model, output a predicted mask image to identify whether each pixel in the input image belongs to the region of interest.

[0041] Due to the lack of a high - quality 10 - meter resolution label dataset, manually annotating a large - scale dataset is both time - consuming and laborious. Therefore, in this embodiment, the 100 - meter resolution label obtained by the interpolation method is regarded as a noisy label. Directly training on the noisy label may lead to a decline in model performance. Therefore, a Noise - robust Learning method is introduced to improve the robustness of the model. For this purpose, the U - Net model is selected for image segmentation in this embodiment, and the DMI Loss is used as the loss function; specifically, it includes the following steps:

[0042] S31. Use U - Net as the image segmentation model and modify the U - Net model parameters.

[0043] Select U - Net, which performs well in image segmentation tasks, as the segmentation model. The number of input channels in_channels = 6, the number of output channels out_channels = 1, input size = 512 * 512, and output size = 512 * 512.

[0044] The U - Net model is a deep - learning model widely used in image segmentation tasks and can perform semantic segmentation accurately at the pixel level. Input the 8,409 image patches after re - cropping and label screening and the corresponding 8,409 mask images into the U - Net model to train the model.

[0045] S32. Select DMI Loss as the loss function to process the noisy labels.

[0046] During the model training process, a customized Difference Maximization Information Loss function (DMI Loss) is adopted to make full use of the potential information in the noisy labels and efficiently learn from the noisy data. The DMI (Determinant-based Mutual Information) loss is a loss function specifically designed to improve the performance of image segmentation. By measuring the distance between the predicted values and the ground truth labels, it optimizes the classification accuracy and boundary details of the model. The DMI Loss is based on the determinant-based mutual information, which not only satisfies theoretical properties such as information monotonicity and relative invariance, but also is the first loss function that can be directly applied to any classification neural network without prior information. The design of the DMI Loss enables it to effectively handle the problem of noisy labels and avoid the interference caused by noisy labels to model training. Compared with traditional cross-entropy loss or Dice loss, the DMI Loss can better handle imbalanced data and fine-grained boundaries.

[0047] The DMI Loss function is specifically designed to improve the performance of image segmentation. By measuring the distance between the predicted values and the ground truth labels, it optimizes the classification accuracy and boundary details of the model. The formula for the DMI Loss is:

[0048] L DMI (data;classifier):=-log[DMI(classifier's output;labels)].

[0049] where DMI is defined as follows:

[0050] DMI(X;Y)=|det(U X,Y )|

[0051] U X,Y (x,y)=Pr[X=x,Y=y]

[0052] X and Y are two discrete random variables; Pr[X=x,Y=y] is the matrix representation of the joint distribution of the discrete random variables X and Y; |det(U X,Y )| is the absolute value of the determinant of the matrix U X,Y .

[0053] By combining noisy label learning and DMI Loss, the accuracy of training with noisy labels is significantly improved, and the confirmation bias caused by noise during the target domain adaptation process of the model is reduced. In addition, this method has a low computational cost and a fast processing speed, effectively improving the prediction accuracy of the model and enhancing its stability and robustness.

[0054] S33. Crop the high-resolution satellite image into tiles of size 512×512 pixels, and set an overlapping area with a preset number of pixels (20 pixels in this embodiment) between each tile. Then input the cropped tiles (i.e., the input image) into the trained U-shaped convolutional neural network model to output a predicted mask image, which identifies whether each pixel in the input image belongs to the region of interest, so as to predict the probability that each pixel point belongs to an oil palm tree. The output mask has the same resolution and size as the input image and can be used for subsequent analysis and visualization.

[0055] This step retains cropping with keep overlap. First, crop the large-size image into small-size image tiles of size 512×512, and retain an overlapping area of 20×20 pixels between each tile. By retaining the overlapping part, the continuity of the prediction results at the boundary is ensured during subsequent stitching.

[0056] This step makes predictions on small-size image tiles (Prediction on Small Images). Input each cropped small-size image tile (i.e., each sub-image) into the trained U-Net model to generate a predicted mask (probability map) for each pixel; each sub-image will obtain a predicted probability mask.

[0057] S4. Post-process the results predicted by the model and stitch them to generate a full-size image.

[0058] After the model prediction, through a series of post-processing steps, stitch the prediction results from the small-size image tiles back to the large-size image (i.e., the full-size image), and at the same time perform optimization processing to further improve the overall prediction quality. Specifically, it includes:

[0059] S41. Calculate the mean probability of the overlapping area and stitch (Calculating Mean Probability&Merging).

[0060] For the overlapping areas of all the cropped small-size image tiles in step S33, calculate the average value of their predicted probabilities to improve the robustness of the prediction. Then, stitch these cropped small-size image tiles back to the original full-size image to generate a complete high-resolution prediction map, forming a high-resolution oil palm tree distribution probability map covering the entire study area, with the size restored to 13568×13568.

[0061] S42. Using Bayesian inference: To improve the prediction accuracy of oil palm planting areas, in this embodiment, a 10-meter resolution mask is used as the prior, and a 100-meter resolution mask is used as the likelihood. Bayesian inference is used to calculate the posterior probability of each pixel in the stitched full-size image. The specific steps are as follows: Define a 21×21 pixel window around each 10-meter pixel marked as 1, calculate the combination of the prior and the likelihood. If the posterior probability is greater than 0.5, the pixel is marked as 1 (oil palm area); otherwise, it is marked as 0.

[0062] In this step, Bayesian inference is used to refine the prediction results of each pixel. Bayesian inference combines prior knowledge and observed data, enabling more accurate posterior probabilities to be obtained. The specific operations are as follows:

[0063] (1) Definition of the prior and the likelihood:

[0064] In this embodiment, two key elements are first defined: the prior and the likelihood.

[0065] Prior: A 10-meter resolution mask is used as the prior information (Prior), where the area marked as 1 represents the known oil palm planting area. Since the 10-meter resolution mask is more detailed and has a higher spatial accuracy, it is considered more reliable prior information.

[0066] Likelihood: A 100-meter resolution mask is used as the likelihood information (Likelihood). The prediction results of the 100-meter resolution mask may have some ambiguity, but it can provide more extensive regional information for subsequent calculations.

[0067] (2) Bayesian inference calculation process:

[0068] For each pixel point with a 10-meter resolution, its predicted value (i.e., whether the pixel belongs to the oil palm planting area) will be updated through Bayesian inference based on the overlapping information of the 10-meter and 100-meter resolution masks.

[0069] First, a 21×21 pixel window is defined for each pixel point, where the central pixel of the window is the current target pixel. This window contains the adjacent pixel data around it, which helps to smooth the prediction of the target pixel.

[0070] Within this window, calculate the prior probability in the 10-meter mask (i.e., the probability that the area is an oil palm planting area) and the likelihood probability in the 100-meter mask (i.e., the predicted probability that the area belongs to the oil palm planting area).

[0071] Based on Bayes' formula, calculate the posterior probability:

[0072]

[0073] Among them, P(palm oil = 1) is the prior (i.e., the probability of the 10-meter mask), P(observed data | palm oil = 1) is the likelihood (i.e., the predicted probability of the 100-meter mask), and P(observed data) is the normalization factor.

[0074] (3) Posterior probability determination:

[0075] By calculating the posterior probability, if the posterior probability of a certain pixel is greater than 0.5, then this pixel is considered to belong to the palm oil planting area, and the label value is set to 1; otherwise, the label value is set to 0. This processing method helps to improve the segmentation accuracy of the palm oil planting area, especially in the boundary area, and can effectively reduce the errors caused by over-labeling or mis-labeling.

[0076] (4) Pixel-by-pixel processing:

[0077] For each pixel with a 10-meter resolution, through the above Bayesian inference processing, more accurate prediction results can be obtained. During each processing, the same operation is performed on all pixels to ensure the global consistency of the final result.

[0078] Combining to Big Image: After completing the prediction of all small-sized image tiles, the ultimate goal is to stitch the prediction results back to the original large image. After Bayesian inference, small-sized image tiles are stitched by retaining the overlapping area (20×20 pixels), and the probability mean of the overlapping area is calculated to ensure smooth and consistent boundaries. Due to the boundary differences between the 10-meter and 100-meter masks, that is, there are differences in spatial resolution, an overlapping area processing strategy is adopted during the stitching process to eliminate non-overlapping areas, and finally a high-resolution palm oil planting distribution map (with a size of 279489×147839) is stitched. Specifically, a 20×20 pixel overlapping area is retained between every two adjacent small-sized image tiles to avoid discontinuous or inconsistent boundaries during stitching. Within these overlapping areas, first, the predicted probability mean of each pixel is calculated to smooth the boundaries and reduce the impact of stitching errors on the image quality. Due to the boundary differences between the 10-meter and 100-meter masks, the 10-meter mask in some areas may be empty or all 0, and these areas need to be eliminated during stitching to ensure that the final image only contains valid palm oil planting areas.

[0079] After stitching and generation, after the predicted results of all valid sub-tiles are subjected to mean calculation and smoothing processing, they are merged back into the original large image, and finally a high-resolution palm oil planting distribution map with a size of 279489×147839 pixels is generated. This image not only maintains a high spatial resolution but also provides accurate spatial data support for subsequent analysis and decision-making.

[0080] Through the above steps, from data acquisition, preprocessing, model training to post-processing and synthesis, this embodiment has successfully achieved high-precision automated prediction for large-scale oil palm planting areas. Especially in the post-processing stage, by combining the methods of Bayesian inference and probability mean calculation, the accuracy and continuity of the stitching results are effectively improved, greatly enhancing the efficiency and reliability of remote sensing image analysis. This method provides a new technical route for remote sensing image classification and recognition and has broad application prospects.

[0081] Based on the same inventive concept, this embodiment also proposes a high-resolution oil palm tree recognition system based on noisy label learning, including:

[0082] A data acquisition module, configured to acquire remote sensing image data and label data, and perform mask mapping and preliminary cropping processing on them to obtain the preliminarily cropped image data;

[0083] A preprocessing module, configured to preprocess the preliminarily cropped image data to obtain several image patches and corresponding several mask maps after re-cropping and label screening as the training set;

[0084] A model training and prediction module, which trains the U-shaped convolutional neural network model using the training set; and outputs a predicted mask map through the trained model to identify whether each pixel in the input image belongs to the region of interest;

[0085] A stitching module, configured to post-process the results predicted by the model and stitch them to generate a full-size image.

[0086] The above modules are respectively used to implement the steps S1-S4 of the above recognition method, and for the detailed implementation process, refer to the description of each step.

[0087] In this embodiment, by adopting technical means combining noise label learning, high-resolution mapping, weakly supervised learning, and Bayesian inference, the cost of manual annotation is reduced, and the accuracy and generalization ability of oil palm tree target recognition in multi-source remote sensing image datasets for Indonesia and Malaysia regions are significantly improved. Compared with the existing Paraformer framework, SinoLC-1 project, and L2HNet method, the present invention breaks through the adaptability bottleneck of complex terrains and diverse land cover types, significantly reduces the dependence on high-resolution labels, and provides an efficient and accurate solution for large-scale high-resolution land cover mapping and species recognition; not only achieves higher computational efficiency in high-resolution image data processing, but also maintains better recognition effects under complex terrains (such as mountains and wetlands), with the generalization ability improved by 2.1%-3.7%. Through the Bayesian inference method, in the recognition of oil palm tree planting areas, the present invention can effectively eliminate the interference of non-oil palm areas, further improving the accuracy and reliability of the recognition results. Through the innovative regional attention mechanism, the prediction accuracy of the model in data imbalance scenarios is significantly improved, and the target consistency of weakly supervised training is increased by more than 15%, thus reducing the risk of model overfitting. On the datasets of Malaysia and Indonesia, the correct rate is increased by 3.5% through validation set learning compared with traditional methods. The solution of the present invention shows efficient and stable characteristics in the fine-grained segmentation of remote sensing images and high-resolution species recognition tasks.

[0088] As can be seen from the technical solutions described in this embodiment, the implementation of this embodiment mainly relies on the following technical means:

[0089] 1. Learning with Noise Label (LNL):

[0090] By introducing noise label learning LNL, the 100-meter resolution label is used as a noisy label for learning to improve the accuracy of pseudo-labels and reduce the confirmation bias of the model when adapting to the target domain. Specifically, the DMI loss obtained through LNL is used for noisy label learning to calculate the difference between the model prediction and the label, improving the reliability and accuracy of the prediction results, and also enhancing the training effect and stability of the target model.

[0091] The noise label learning module makes full use of the potential information in the noise labels through a customized divergence maximization information loss function (DMI Loss) to efficiently learn the noisy data. This module optimizes the distance between the predicted value and the true label, enhancing the adaptability of the model in imbalanced data and complex geographical conditions. The Bayesian inference module, aiming at the interference of uncertainty, significantly improves the generalization ability of the model in diverse geographical environments by accurately inferring the probability of non-target areas.

[0092] 2. High-Resolution Cartography (HRC)

[0093] Using high-resolution cartography techniques, batch preprocessing of label data is carried out by methods such as interpolation to improve the resolution of label images. This method does not require additional manual annotation, reduces labor costs, improves the efficiency of process execution, enhances model adaptability, creates a pipeline-style data processing standard, and strengthens the reliability of data.

[0094] 3. Weakly Supervised Learning (WSL)

[0095] This invention adopts weakly supervised learning. Instead of re-annotating the images, it uses the existing labels for preprocessing to replace the per-pixel annotation method, reduces the dependence on image annotation, enhances model adaptability, and improves the performance of the model. It further reduces errors and decreases the requirements for data.

[0096] 4. Bayesian Inference (BI)

[0097] The Bayesian inference method is introduced to eliminate the interference of non-target areas in the land cover data through a probability model, thus focusing on the area of interest - the area for planting oil palm trees. Bayesian inference utilizes prior knowledge and observed data to accurately identify the target area by calculating the posterior probability, optimizing the model's adaptability to complex terrains and areas with diverse features. Especially in high-resolution remote sensing images, it can more accurately identify and classify oil palm trees, maintaining high accuracy even near non-target areas, thereby significantly enhancing the robustness and practicality of the model.

[0098] In summary, this invention adopts methods of noisy label learning, high-resolution cartography, weakly supervised learning, and Bayesian inference, significantly improving the recognition accuracy of the model, enhancing the model's adaptability and generalization ability to the dataset, making up for the deficiency of the existing oil palm tree recognition model's high requirements for the dataset, enhancing the usability of the model, and filling the gap in the existing field of oil palm tree recognition.

[0099] The above embodiments are one of the implementation manners of the present invention. The implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.

Claims

1. A high-resolution oil palm tree recognition method based on noise label learning, characterized in that: The following steps are involved: S1, obtaining remote sensing image data and label data, and performing mask mapping and preliminary cropping processing on them to obtain preliminary cropped image data; S2, preprocessing the image data after preliminary cropping to obtain a number of image blocks that have been cropped again and label-screened and a number of corresponding mask images as a training set; S3, using the training set obtained in step S2 to train the U-shaped convolutional neural network model; outputting a predicted mask map through the trained model to identify whether each pixel in the input image belongs to the region of interest; S4. Post-process the results predicted by the model and stitch them together to generate a full-size image.

2. The high-resolution oil palm identification method according to claim 1, wherein The process of re-cropping and label screening in step S2 includes: The image block after the initial cropping is further cropped into a sub-image of a smaller size to obtain a cropped image block; the cropped image block is subjected to mask screening to retain only image samples whose proportion of the area of ​​interest exceeds a preset ratio value; The interpolation method is used to change the 100m resolution label image to 10m resolution, and the pixels are enlarged so that the pixels of the image samples after mask screening and the satellite image correspond one to one.

3. The high-resolution oil palm identification method according to claim 1, wherein Step S3 includes: S31, using a U-shaped convolutional neural network model as an image segmentation model, and training the image segmentation model using the impact block that has been cropped again and label-screened and the corresponding mask map; S32. During the model training process, select DMI Loss as the loss function to handle noisy labels; S33, cropping the high-resolution satellite image into tiles of size 512×512 pixels, setting an overlapping area of ​​a preset pixel amount between each tile; inputting the cropped tiles into the trained image segmentation model to output a predicted mask map, marking whether each pixel in the input image belongs to the area of ​​interest, so as to predict the probability of each pixel belonging to an oil palm tree.

4. The high-resolution oil palm identification method according to claim 3, wherein Step S4 includes: S41, for the overlapping areas of the image blocks cropped in step S33, the average value of the prediction probabilities is calculated; the cropped image blocks are spliced ​​back to the original full-size image to generate a complete high-resolution prediction map, thereby forming a high-resolution oil palm tree distribution probability map covering the entire study area; S42. Using the 10-meter resolution mask as the prior and the 100-meter resolution mask as the likelihood, Bayesian inference is used to calculate the posterior probability of each pixel of the spliced ​​full-size image to obtain an accurate prediction result.

5. The high-resolution oil palm identification method according to claim 4, wherein The calculation method of the posterior probability based on the Bayesian formula is: Among them, P(Oil Palm = 1) is the prior, P(Observed Data|Oil Palm = 1) is the likelihood, and P(Observed Data) is the normalization factor.

6. The high-resolution oil palm identification method according to claim 2, wherein: The preset ratio value is 30%.

7. A high-resolution oil palm tree recognition system based on noisy label learning, characterized in that: include: The data acquisition module is used to acquire remote sensing image data and label data, and perform mask mapping and preliminary cropping processing on them to obtain the image data after preliminary cropping; A preprocessing module is used to preprocess the image data after preliminary cropping, and obtain a number of image blocks that have been cropped again and label-screened and a number of corresponding mask images as a training set; The model training and prediction module uses the training set to train the U-shaped convolutional neural network model; the trained model outputs a prediction mask map to identify whether each pixel in the input image belongs to the region of interest; The stitching module is used to post-process the results predicted by the model and stitch them together to generate a full-size image.

8. The high-resolution oil palm tree identification system according to claim 7, characterized in that: In the preprocessing module, the process of re-cropping and label screening includes: The image block after the initial cropping is further cropped into a sub-image of a smaller size to obtain a cropped image block; the cropped image block is subjected to mask screening to retain only image samples whose proportion of the area of ​​interest exceeds a preset ratio value; The interpolation method is used to change the 100m resolution label image to 10m resolution, and the pixels are enlarged so that the pixels of the image samples after mask screening and the satellite image correspond one to one.

9. The high-resolution oil palm tree identification system according to claim 7, characterized in that: The processing of the model training and prediction module includes: A U-shaped convolutional neural network model is used as an image segmentation model, and the image segmentation model is trained using the impact block that has been cropped again and label-screened and the corresponding mask map; During model training, DMI Loss is selected as the loss function to handle noisy labels; The high-resolution satellite image is cropped into tiles of size 512×512 pixels, and an overlapping area with a preset number of pixels is set between each tile. The cropped tiles are input into the trained image segmentation model to output a prediction mask map, which identifies whether each pixel in the input image belongs to the area of ​​interest, so as to predict the probability of each pixel belonging to an oil palm tree.

10. The high-resolution oil palm tree identification system according to claim 9, characterized in that: The processing of the splicing module includes: For the overlapping areas of the cropped tiles, the average value of their predicted probabilities is calculated; the cropped tiles are stitched back to the original full-size image to generate a complete high-resolution prediction map, forming a high-resolution oil palm tree distribution probability map covering the entire study area; The 10-meter resolution mask is used as the prior and the 100-meter resolution mask is used as the likelihood. Bayesian inference is used to calculate the posterior probability of each pixel to obtain accurate prediction results.