Unmanned aerial vehicle sugarcane planting area image recognition method
Through the improved U-Net segmentation model and rule engine optimization technology, the problem of large errors in measurement accuracy and low efficiency of the drone sugarcane planting area is solved, and efficient and accurate identification of sugarcane planting area in resource-constrained environments is achieved.
Patent Information
- Application Number
- CN202510090294.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The measurement of existing drone sugarcane planting area has problems of large accuracy error and low efficiency, and traditional U-Net segmentation models are difficult to maintain stability and accuracy in resource-constrained environments.
Using an improved U-Net segmentation model, we generate geo-referenced stitching images by identifying feature points in the orthogonal image, cropping, chunking and enhancement processing, and building a rule engine to optimize model performance to ensure stability and accuracy in resource-constrained environments.
It realizes efficient processing of large-area images, improves the accuracy and efficiency of sugarcane planting area recognition, ensures the stability and accuracy of the model in resource-constrained environments, and is suitable for a variety of segmentation tasks.
Smart Images

Figure CN120014493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles (UAVs), and in particular to a method for unmanned aerial vehicle (UAV) image recognition of sugarcane planting areas. Background Art
[0002] A drone is an aircraft that can fly autonomously or remotely without human control. In sugarcane planting, drones are widely used for filming and monitoring. The high-definition cameras on drones can obtain bird's-eye views of sugarcane fields to monitor sugarcane growth, pests and diseases, and soil moisture. After processing and analysis, these image data can provide accurate agricultural management information for sugarcane planting, helping farmers improve planting efficiency and optimize planting strategies.
[0003] Sugarcane is an important sugar crop in the world. It has a wide range of application values in agriculture, industry, medicine and ecological protection. Therefore, it is of great significance to accurately count the sugarcane planting area. The existing sugarcane planting area measurement is basically carried out by using UAV remote sensing images. However, the photographed image is generally divided into multiple tiles. The planting area is obtained by calculating the area of the corresponding tile by combining the center point. This has the problems of large accuracy error and low efficiency.
[0004] In addition, for orthophotos of sugarcane fields collected by drones, the traditional U-Net segmentation model is usually used to deal with orthophotos. There are challenges in adaptability and optimization in this process. Specific problems include: it is difficult for the model to maintain stability and accuracy in a resource-constrained environment, and it is often unable to fully exert its performance due to memory limitations. At the same time, the key performance indicators of the model, such as accuracy, precision, and recall, vary on different images, making it difficult to achieve a unified high standard, which will affect the effectiveness of image recognition of sugarcane planting areas to a certain extent. Summary of the invention
[0005] 1. Technical issues to be resolved
[0006] In view of the shortcomings of the prior art, the present invention provides a method for unmanned aerial vehicle (UAV) sugarcane planting area image recognition, which solves the problems raised in the background technology.
[0007] (II) Technical solution
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0009] A method for unmanned aerial vehicle sugarcane planting area image recognition comprises the following steps:
[0010] S1, obtaining the orthophoto image corresponding to the target sugarcane planting area;
[0011] S2, identifying feature points in the orthophoto image, completing image consistency operations, and generating a mosaic image with geographic reference;
[0012] S3, cutting, dividing and enhancing the stitched image in sequence;
[0013] S4, configuring an improved U-Net segmentation model and optimizing it until a standard is set, and then building a rule engine to determine a minimum performance directional value according to the amount of memory occupied by the current orthophoto image, and when the performance state of the improved U-Net segmentation model does not reach the minimum performance directional value, executing a search optimization strategy until the performance state of the improved U-Net segmentation model reaches the minimum performance directional value;
[0014] S5, using the improved U-Net segmentation model to recognize the processed spliced image and generate a semantic mask;
[0015] S6, converting and adjusting the semantic mask;
[0016] S7, extracting the mask image of the sugarcane field category and obtaining its boundary pixels, and calculating the planting area of sugarcane according to the pixel size and number occupied by the mask image of the sugarcane field category.
[0017] Furthermore, the orthophoto image is obtained by a drone equipped with a multispectral camera, and the multispectral camera has a GPS positioning module, and the heading overlap of the drone in the working state shall not be less than 80%, and the lateral overlap shall not be less than 70%.
[0018] Furthermore, the image consistency operation includes at least: alignment, stitching and geometric correction of images, and matching of overlapping areas.
[0019] Furthermore, the enhancement processing at least includes: denoising, contrast adjustment and rotation adjustment.
[0020] Furthermore, the improved U-Net segmentation model consists of an encoder and a decoder;
[0021] The encoder gradually extracts the features of the image through several convolutional layers and pooling layers, while reducing the spatial size of the feature map. Each convolutional layer contains two 3*3 convolution operations and one 2*2 maximum pooling operation;
[0022] The decoder gradually restores the spatial size of the feature map through upsampling operations and fuses it with the feature map of the corresponding encoder layer. Each upsampling layer contains a 2*2 deconvolution operation, followed by two 3*3 convolution operations, connected by several nested dense convolution blocks.
[0023] Furthermore, the set standard indicates that: the average intersection-over-union ratio corresponding to the improved U-Net segmentation model reaches a preset value;
[0024] The content of the constructed rule engine is as follows:
[0025] First, summarize and generate the performance indicator value Iot based on the performance indicator data set. The formula is as follows:
[0026]
[0027] In the formula, Ac, Pr, Re, Fs, Ma and Io are expressed as follows:
[0028] Performance indicators Accuracy, precision, recall, F1 score, average accuracy, and intersection over union ratio in the dataset;
[0029] Secondly, based on the amount of memory occupied by the orthophoto image, the sigmoid function is used to calculate the minimum performance orientation value Iot_min. The formula is as follows:
[0030]
[0031] Wherein, P0 is expressed as: the adjusted basic performance index value parameter, and P0 < 1;
[0032] k, M, B and n are represented as:
[0033] Adjustment coefficients, memory amounts, baseline memory amounts, and power exponents;
[0034] Among them, k>0, n=1.
[0035] Furthermore, the search optimization strategy implemented is as follows:
[0036] Compare any type of data in the performance indicator data set with the corresponding threshold. If any type or types of data do not reach the corresponding threshold, perform optimization operations. The optimization operations at least include:
[0037] Adjust model parameters, change model architecture, and hardware acceleration.
[0038] Furthermore, the spliced image is input into the trained improved U-Net segmentation model, and the model predicts its category pixel by pixel to generate a binary or multi-class mask map of the plot segmentation, namely the semantic mask. The semantic mask assigns each pixel to a different category, including at least 8 categories, namely: sugarcane field, pond, river, house, tractor, car, road and rice field; among them, the background is represented by 0, and the 8 categories are represented by numbers 1 to 8 respectively.
[0039] Furthermore, the content of the conversion adjustment process is as follows:
[0040] Use OpenCV and Shapely libraries to convert the semantic masks generated by the improved U-Net segmentation model into polygon outlines. Use the GDAL library to convert polygon outlines into GeoJSON format. During the conversion process, refer to the geographic information of the cropped image to convert pixel coordinates into the geographic coordinate system. In the generated GeoJSON format file, each polygon contains at least the following attributes: category and boundary.
[0041] (III) Beneficial effects
[0042] The present invention provides a method for unmanned aerial vehicle (UAV) sugarcane planting area image recognition, which has the following beneficial effects:
[0043] (1) This solution can efficiently process images of large areas. The drone automatically performs flight shooting tasks and can process images of large sugarcane planting areas after stitching. At the same time, there are fewer data processing links in the intermediate process. The improved U-Net segmentation model can also achieve good segmentation effects on a very small training set, without the need for a large number of labeled data sets for training. After completing semantic segmentation, the sugarcane planting area can be quickly obtained.
[0044] (2) This scheme constructs a rule engine to dynamically determine the minimum performance benchmark value based on the memory capacity of the orthophoto image, thereby ensuring the stability and accuracy of the model in a resource-constrained environment. At the same time, by implementing a search optimization strategy, targeted optimization is performed on indicators that do not reach the performance threshold, effectively improving the model's key performance indicators such as accuracy, precision, and recall. In addition, the introduced skip connections and nested dense convolution blocks enhance the network feature extraction and transmission efficiency, and the optimized loss function and dynamic learning rate decay strategy further improve the model's convergence speed and segmentation accuracy. Finally, this scheme achieves a balance between high accuracy and robustness in a variety of segmentation tasks, providing an efficient and flexible segmentation model optimization method for the image processing field.
[0045] (3) This scheme uses the improved U-Net segmentation model to classify the spliced images captured by the drone pixel by pixel and generate semantic masks containing categories such as sugarcane fields. Subsequently, the mask edges are optimized through morphological operations or boundary smoothing techniques and converted into GeoJSON format for easy reading and analysis by the GIS system. Finally, the mask map of the sugarcane field category is extracted and combined with the latitude and longitude mapping relationship of the pixels to calculate the planting area of sugarcane.
[0046] This process not only improves the accuracy and efficiency of plot segmentation, but also achieves seamless conversion from images to geographic information, providing accurate spatial data support for agricultural management and decision-making; in addition, by merging multiple images into one monitoring area image and calculating the area and boundaries, the applicability and flexibility of the technology in practical applications are further enhanced, effectively meeting the needs of agricultural remote sensing monitoring and precision agriculture. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of a flow chart of a method for image recognition of sugarcane planting area using a drone according to the present invention;
[0048] Figure 2 This is the architecture diagram of the improved U-Net segmentation model;
[0049] in, Figure 2 The four lines of English in the middle right subscript are expressed from top to bottom as follows:
[0050] Downsampling, upsampling, skip connections, and X 4,0 convolution. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] See also Figure 1 to Figure 2 This embodiment provides a method for unmanned aerial vehicle sugarcane planting area image recognition, the method comprising the following steps:
[0053] S1, obtaining the orthophoto image corresponding to the target sugarcane planting area;
[0054] Among them, this step is completed by using a drone equipped with a multispectral camera to execute the flight plan;
[0055] The drone takes aerial photos over the sugarcane planting area according to the preset flight path, and the multispectral camera is responsible for capturing orthophoto images of the ground. These images carry GPS longitude and latitude information to ensure that the geographic spatial position can be accurately located during subsequent processing. The requirements for heading overlap and lateral overlap ensure that there is enough overlapping area between the images to facilitate subsequent image stitching and geometric correction.
[0056] Specifically, the multispectral camera has a GPS positioning module to obtain an orthophoto image with GPS latitude and longitude information, requiring that the heading overlap should not be less than 80% and the lateral overlap should not be less than 70%;
[0057] Note: This step is the basis of the entire process. Obtaining high-quality orthophotos is crucial for subsequent image processing and segmentation.
[0058] S2, identifying feature points in the orthophoto image to complete the image consistency operation and generate a mosaic image with geographic reference;
[0059] The image consistency operation includes at least: alignment, splicing and geometric correction of images, and matching overlapping areas;
[0060] Explanation: Use software such as Pi*4D to identify feature points in orthophoto images, and use these feature points to complete image alignment, stitching, and geometric correction; at the same time, apply algorithms such as Bundle Adjustment to ensure the spatial consistency of each image, and finally generate a stitched image with geographic reference;
[0061] Note: This step ensures that multiple orthophotos can be accurately stitched together to form a complete, georeferenced mosaic image, which provides the basis for subsequent segmentation and cropping;
[0062] The image recognition accuracy of this solution is high. After the original image taken by the drone multispectral is spliced, cropped, and semantically segmented, the pixels have a specific mapping relationship with the longitude and latitude. Compared with the method of calculating the coordinates of the center point of the picture which is affected by the ground terrain, the flight altitude of the drone, and the camera equipment, the recognition accuracy of the present invention is higher.
[0063] S3, cutting, dividing and enhancing the stitched image in sequence;
[0064] The enhancement processing at least includes: denoising, contrast adjustment and rotation adjustment;
[0065] The GeoTIFF format stitching results containing geographic coordinate information are obtained in S3; the stitched GeoTIFF format images are cropped and divided into image blocks with a resolution of 1024*1024 using the GDAL library; at the same time, the geographic information of the image is retained for subsequent analysis; cropping and dividing are to facilitate subsequent image processing and segmentation, while retaining the geographic information ensures the accuracy and traceability of the processing results;
[0066] Image enhancement can effectively remove noise from the image, thereby improving image quality and enhancing the performance and robustness of the model; contrast adjustment helps enhance image details, making it easier for the model to identify targets in segmentation tasks; rotation adjustment introduces diversity to the dataset, simulating image inputs at different angles, and further improving the generalization ability of the model.
[0067] S4, configuring an improved U-Net segmentation model and optimizing it until a standard is set, and then building a rule engine to determine a minimum performance directional value according to the amount of memory occupied by the current orthophoto image, and when the performance state of the improved U-Net segmentation model does not reach the minimum performance directional value, executing a search optimization strategy until the performance state of the improved U-Net segmentation model reaches the minimum performance directional value;
[0068] Among them, the basic network architecture of the improved U-Net segmentation model is the U-Net model;
[0069] This application adopts an image segmentation model based on an improved U-Net basic network architecture, and continuously optimizes the model through optimization processing; wherein the optimization processing at least includes: adjusting hyperparameters and modifying the model structure;
[0070] The improved U-Net segmentation model consists of an encoder and a decoder;
[0071] The encoder and decoder gradually extract high-level features of the image through a series of convolutional layers and pooling layers, while reducing the spatial size of the feature map. Each convolutional layer contains two 3*3 convolution operations, followed by a 2*2 maximum pooling operation; the decoder gradually restores the spatial size of the feature map through upsampling operations and fuses it with the feature map of the corresponding encoder layer. Each upsampling layer contains a 2*2 deconvolution operation, followed by two 3*3 convolution operations, which are connected through a series of nested dense convolution blocks;
[0072] The core idea of the improved U-Net is to narrow the semantic gap between the feature maps of the encoder and decoder before fusion;
[0073] The improved U-Net introduces skip connections between the encoder and decoder at each level and within each sub-U-Net structure, which enhances the efficiency of feature transmission and utilization;
[0074] Specifically, the convolution block X is introduced 0,1 , X 0,2 and X 1,1 The encoder and decoder are connected by nested dense convolution blocks. These blocks perform feature transformation and extraction at different levels, which effectively enhances the network's expressive power. For the model structure, see Figure 2 You can know;
[0075] In the first layer of skip connection of the encoder, add two convolution modules X 0,1 and X 0,2 , by using convolution kernels of different sizes to extract multi-scale features; in the second layer of skip connection, add a convolution module X 1,1 , to enhance the local feature extraction capability; these newly added convolution modules use ReLU activation function and BatchNormalization to ensure the effectiveness and stability of feature extraction; in addition, in order to optimize the computational efficiency of the model, the number of channels of the convolution module is appropriately trimmed to match the subsequent decoding part;
[0076] In the decoding part of the segmentation model, the output feature layer of the encoding part is used to perform feature fusion with the corresponding layer of the decoder through skip connections. In the feature fusion process, for each preliminary effective feature layer, an upsampling operation is first performed, in which the upsampling adopts bilinear interpolation to restore the resolution of the image.
[0077] The upsampled feature layer is concatenated with the feature layer transmitted by the skip connection, and the features are refined through a series of convolution operations to ensure the effective combination of semantic information and spatial detail information.
[0078] Through the above improvements, the model can more accurately restore the segmentation results of the input image;
[0079] During the model training process, a weighted combination of the cross entropy loss function and the Dice loss function is used to improve the segmentation performance of small target areas. To further optimize the convergence speed and performance of the model, the Adam optimizer is used and a dynamic learning rate decay strategy is set. Finally, the model shows high accuracy and robustness in a variety of segmentation tasks.
[0080] After feature extraction is completed, the extracted feature map is sent to the decoder part, and the original resolution of the image is gradually restored through multiple upsampling operations; in each upsampling process, the corresponding feature map from the encoder part is fused to retain important spatial information; finally, the model fuses the features of each layer in turn and outputs a feature map of the same size as the original image;
[0081] The feature map is then sent to the prediction module for classification prediction, and the segmentation result is output using the full convolution layer and Sigmoid activation function. Combined with the post-processing steps, the complete segmented image is finally obtained;
[0082] The improved U-Net model is implemented based on the Pytorch2.3.0 framework, using the Python3.11.5 programming language, OpenCV4.9.0.80 for image processing, and the operating environment is the PyCharm integrated development environment (IDE). During the model training process, the batch size is set to 16, and 300 epochs are trained, of which the first 50 epochs are in a frozen state and the model parameters are not updated.
[0083] The set standard means: the average intersection-over-union ratio corresponding to the improved U-Net segmentation model reaches the preset value;
[0084] For example, after training and optimization, the average intersection-over-union (MIOU) of the final improved U-Net segmentation model reached 94.31%, where 94.31% is the preset value;
[0085] The content of the constructed rule engine is as follows:
[0086] First, summarize and generate performance indicator values based on the performance indicator data set. The formula is as follows:
[0087]
[0088] Where, Iot represents the performance index value, and Ac, Pr, Re, Fs, Ma and Io are expressed as follows:
[0089] Accuracy, precision, recall, F1 score, average accuracy, and intersection-over-union ratio; Therefore, the performance indicator data set consists of accuracy, precision, recall, F1 score, average accuracy, and intersection-over-union ratio;
[0090] It should be noted that the molecular part:
[0091] The product of accuracy, precision, recall, F1 score, average accuracy and IoU: This reflects the comprehensive performance of the model in multiple dimensions; the average accuracy is processed using the square root to avoid excessive impact of its value on the overall result while maintaining its importance (because accuracy is usually a ratio, which may be between 0 and 1, and its square root will change more smoothly within this range);
[0092] Denominator:
[0093] The sum of accuracy and precision: Both are important indicators for measuring the accuracy of model predictions. Their sum, as part of the denominator, can balance the model's performance in accurate and precise predictions; The product of recall and (1-F1 score): The recall reflects the model's ability to identify the positive class, while the (1-F1 score) reflects the model's lack of balance between precision and recall. Multiplying the two can penalize models that perform well in recall but poorly in balancing precision and recall; The product of square root average accuracy and (1-IoU): This is also to balance the model's performance in overall accuracy and spatial overlap (for tasks such as image segmentation);
[0094] Overall logic:
[0095] Through the combination of numerator and denominator, the formula aims to comprehensively evaluate the performance of the model in multiple dimensions, while penalizing those models that perform well in some aspects but have obvious deficiencies in others;
[0096] Here are some examples:
[0097] Suppose there is a model that performs as follows on the following metrics:
[0098] Accuracy = 0.9;
[0099] Precision = 0.85;
[0100] Recall = 0.8;
[0101] F1 score = 0.825 (harmonic mean of precision and recall);
[0102] Average accuracy = 0.85 (assumed to be the average of all categories’ accuracy);
[0103] Intersection-over-union ratio = 0.75;
[0104] Substituting these values into the formula:
[0105] Then, the performance index value is ≈0.597,
[0106] This performance index value reflects the comprehensive performance of the model in multiple dimensions, and due to the design of the formula, it penalizes the relatively low performance of the model in recall and IoU;
[0107] Secondly, based on the amount of memory occupied by the orthophoto image, the sigmoid function is used to calculate the minimum performance orientation value Iot_min. The formula is as follows:
[0108]
[0109] In the formula, P0 is expressed as:
[0110] The adjusted basic performance index value parameter, and P0<1;
[0111] k, M, B and n are represented as:
[0112] Adjustment coefficients, memory amounts, baseline memory amounts, and power exponents;
[0113] Wherein, k>0, n=1;
[0114] It should be noted that the adjusted basic performance index value parameter P0 is used to provide a basic offset in the sigmoid function; k, B and n are used to adjust the degree of influence, standardize the amount of memory and control nonlinear influence respectively;
[0115] Assume we set the following parameters:
[0116] P0 = -1 (the adjusted basic performance index value parameter makes Iot_min close to 0.5 under the baseline memory amount);
[0117] k = 1 (adjustment coefficient);
[0118] B = 100MB (baseline memory);
[0119] M = 50MB (memory);
[0120] n = 1 (power exponent, used for linear influence, can be adjusted to nonlinear as needed);
[0121] Then, the minimum performance index value ≈ 0.377;
[0122] This means that for an orthophoto occupying 50MB of memory, the minimum performance heading value is about 0.377;
[0123] Combined with the previous example, 0.597>0.377, so the performance of the improved U-Net segmentation model has reached the minimum performance benchmark value;
[0124] Performance indicator dataset:
[0125] Accuracy, Precision, Recall, F1 Score, Mean Accuracy, Frequency Weighted Accuracy, and Intersection over Union (IoU);
[0126] Accuracy is the proportion of samples predicted correctly by the model to the total samples. The formula is: TP, FP, TN, and FN respectively represent TP is a true positive, FP is a false positive, TN is a true negative, and FN is a false negative.
[0127] Precision refers to the proportion of samples predicted by the model to be positive that are actually positive. The formula is:
[0128]
[0129] Recall is the ratio of samples that are actually positive that are correctly predicted as positive by the model. The formula is:
[0130] The F1 score (F1Score) represents the harmonic mean of precision and recall, and is used to comprehensively evaluate the performance of the model. The formula is:
[0131] Mean Accuracy averages the accuracy of each category, which is useful when dealing with multi-category problems and can reflect the performance of each category. You can also use Weighted Accuracy as an alternative, which is the average accuracy weighted by the sample frequency of each category, which can better reflect the performance of the model on an unbalanced dataset.
[0132] The intersection over union (IoU) measures the degree of overlap between the predicted area and the true area. The formula is:
[0133] In addition, the mean intersection over union (Mean IoU) represents the average value of the IoU of all categories, which is usually used to comprehensively evaluate the performance of the model in each category; after testing the images of the validation set to verify the performance of the segmentation method, the above-mentioned index data are calculated as follows: Accuracy 97.39%, Precision 97.49%, Recall 96.68%, F1 Score 97.06%, Mean Accuracy 96.68%, Frequency Weighted Accuracy 94.92%, IoU 96.18%, Mean IoU 94.31%;
[0134] When the performance status of the improved U-Net segmentation model does not reach the minimum performance benchmark value, a search optimization strategy is executed, and the content of the search optimization strategy is as follows:
[0135] Compare any type of data in the performance indicator data set with the corresponding threshold (for example, the corresponding threshold of accuracy is 90%, the corresponding threshold of precision is 90%...). If any type or types of data do not reach the corresponding threshold, then perform optimization operations. The optimization operations include at least:
[0136] Adjust model parameters, change model architecture, and hardware acceleration;
[0137] Among them, adjust the model parameters:
[0138] Hyperparameter tuning: adjust the model's hyperparameters (such as learning rate, regularization coefficient, batch size, etc.) to improve the model's performance; Network structure adjustment: adjust the number of neural network layers, the number of neurons in each layer, activation functions, etc. according to the characteristics of the data and the requirements of the task;
[0139] Changing the model architecture:
[0140] Change the model architecture: try to use more advanced neural network architectures (such as convolutional neural network CNN, recurrent neural network RNN, Transformer, etc.) to adapt to complex data and tasks; ensemble learning: combine the prediction results of multiple models to improve the overall prediction performance;
[0141] Hardware Acceleration:
[0142] Use more powerful computing resources: Use GPU, TPU and other hardware to accelerate computing, shorten training time and improve model convergence speed;
[0143] By adopting the above technical solution, the technical effect of significantly improving the performance of the improved U-Net segmentation model is achieved;
[0144] The adaptability and optimization problems of the model on orthophotos with different memory sizes are solved. The solution builds a rule engine to dynamically determine the minimum performance benchmark value according to the memory size of the orthophoto image, ensuring the stability and accuracy of the model in a resource-constrained environment. At the same time, by implementing a search optimization strategy, targeted optimization is performed on indicators that do not reach the performance threshold, including adjusting model parameters, changing model architecture, and hardware acceleration, effectively improving the model's key performance indicators such as accuracy, precision, and recall. In addition, the introduced jump connections and nested dense convolution blocks enhance the network feature extraction and transmission efficiency, and the optimized loss function and dynamic learning rate decay strategy further improve the model's convergence speed and segmentation accuracy. Finally, the solution achieves a balance between high accuracy and robustness in a variety of segmentation tasks, providing an efficient and flexible segmentation model optimization method for the image processing field.
[0145] S5, using the improved U-Net segmentation model to recognize the processed spliced image and generate a semantic mask;
[0146] Among them, explanation: the cropped 1024*1024 resolution image is input into the trained improved U-Net segmentation model, and the model predicts its category pixel by pixel to generate a binary or multi-class mask map (i.e., semantic mask) for land parcel segmentation; this mask assigns each pixel to a different category, such as background, land parcel, road, etc.; description: this step is the core part of image segmentation, and the generated semantic mask provides the basis for subsequent analysis and calculation;
[0147] In order to facilitate reading and analysis by applications such as geographic information systems (GIS), it is usually necessary to convert image mask data into geographic information formats (such as GeoJSON);
[0148] It should be noted that the dataset used in the present invention for training and testing images includes images captured using a commercial drone at an altitude range of 60 to 120 meters; the dataset provides 5766 images of 1024*1024 pixels and 8 categories (excluding background) of ground truth masks; the 8 categories include: sugarcane field, pond, river, house, tractor, car, road and rice field (sugarcane field corresponds to 1, pond corresponds to 2, river corresponds to 3, house corresponds to 4, tractor corresponds to 5, car corresponds to 6, road corresponds to 7 and rice field corresponds to 8); in the labeled image, the background is represented by 0, and the category is represented by numbers 1 to 8; the dataset is divided into a training set of 4324 pictures and a validation set of 1442 pictures;
[0149] S6, converting and adjusting the semantic mask (through morphological operations or boundary smoothing technology to improve the accuracy of segmentation edges);
[0150] Explanation: Use OpenCV and Shapely libraries to convert the generated binary or multi-class masks into polygon outlines, and further convert them into GeoJSON format; in this process, it is necessary to refer to the geographic information of the cropped image and convert the pixel coordinates into the geographic coordinate system; Description: This step converts the semantic mask into a format that is easier to read and analyze in the GIS system, which is convenient for subsequent application and display;
[0151] The content of conversion adjustment processing is as follows:
[0152] Use OpenCV and Shapely libraries to convert the binary or multi-class masks generated by the improved U-Net segmentation model into polygonal contours; first use OpenCV's cv.findContours function to find the contours in the binary mask, and use cv.drawContours function to draw and extract the contours; then use Shapely's polygon construction function to convert these contours into polygons in geographic coordinate space; use the GDAL library to convert polygons into GeoJSON format; during the conversion process, it is necessary to refer to the geographic information of the cropped image to convert the pixel coordinates into the geographic coordinate system; in the generated GeoJSON file, each polygon (plot) contains its category, boundary and other attributes, which is convenient for further GIS application analysis and display;
[0153] S7, extracting the mask image of the sugarcane field category and obtaining its boundary pixels, and calculating the planting area of sugarcane according to the pixel size and number occupied by the mask image of the sugarcane field category;
[0154] Explanation: First, extract the mask map of the sugarcane field category, and then obtain its boundary pixels; calculate the sugarcane planting area based on the pixel size and the number of sugarcane field pixels and boundary pixels; Description: This step is the ultimate goal of the entire process. By calculating the sugarcane planting area, it provides important data support for agricultural management and decision-making;
[0155] Among them, according to the latitude and longitude mapping relationship of the image pixels, multiple images are merged into a monitoring area image according to the latitude and longitude coordinates of each pixel; finally, the area and boundary are calculated; first, according to the boundary extraction algorithm in the attached figure, the boundary pixels of each type of crop are extracted; then according to the pixel size, the number of crop patch pixels and the number of boundary pixels, the sugarcane planting area is calculated.
[0156] By adopting the above technical solution, the technical effect of efficiently and accurately extracting the sugarcane planting area is achieved;
[0157] The difficult problems of plot information extraction and area calculation in agricultural remote sensing images are solved. The solution first uses the improved U-Net segmentation model to classify the stitched images captured by drones pixel by pixel to generate semantic masks containing categories such as sugarcane fields. Subsequently, the mask edges are optimized through morphological operations or boundary smoothing techniques and converted into GeoJSON format for easy reading and analysis by GIS systems. Finally, the mask map of the sugarcane field category is extracted, and the planting area of sugarcane is calculated by combining the latitude and longitude mapping relationship of the pixels. This process not only improves the accuracy and efficiency of plot segmentation, but also realizes the seamless conversion from images to geographic information, providing accurate spatial data support for agricultural management and decision-making. In addition, by merging multiple images into one monitoring area image and calculating the area and boundaries, the applicability and flexibility of the technology in practical applications are further enhanced, effectively meeting the needs of agricultural remote sensing monitoring and precision agriculture.
[0158] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. A person of ordinary skill in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.
[0159] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0160] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. A method for unmanned aerial vehicle sugarcane planting area image recognition, characterized in that: The steps include: S1, obtaining the orthophoto image corresponding to the target sugarcane planting area; S2, identifying feature points in the orthophoto image, completing image consistency operations, and generating a mosaic image with geographic reference; S3, cutting, dividing and enhancing the stitched image in sequence; S4, configuring an improved U-Net segmentation model and optimizing it until a standard is set, and then building a rule engine to determine a minimum performance directional value according to the amount of memory occupied by the current orthophoto image, and when the performance state of the improved U-Net segmentation model does not reach the minimum performance directional value, executing a search optimization strategy until the performance state of the improved U-Net segmentation model reaches the minimum performance directional value; S5, using the improved U-Net segmentation model to recognize the processed spliced image and generate a semantic mask; S6, converting and adjusting the semantic mask; S7, extracting the mask image of the sugarcane field category and obtaining its boundary pixels, and calculating the planting area of sugarcane according to the pixel size and number occupied by the mask image of the sugarcane field category.
2. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The orthophoto image is obtained by a drone equipped with a multispectral camera, and the multispectral camera has a GPS positioning module. In the working state, the heading overlap of the drone shall not be less than 80%, and the lateral overlap shall not be less than 70%.
3. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: Image consistency operations include at least: image alignment, stitching and geometric correction, and matching overlapping areas.
4. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The enhancement processing at least includes: denoising, contrast adjustment and rotation adjustment.
5. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The improved U-Net segmentation model consists of an encoder and a decoder; The encoder gradually extracts the features of the image through several convolutional layers and pooling layers, while reducing the spatial size of the feature map. Each convolutional layer contains two 3*3 convolution operations and one 2*2 maximum pooling operation; The decoder gradually restores the spatial size of the feature map through upsampling operations and fuses it with the feature map of the corresponding encoder layer. Each upsampling layer contains a 2*2 deconvolution operation, followed by two 3*3 convolution operations, connected by several nested dense convolution blocks.
6. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The set standard means: the average intersection-over-union ratio corresponding to the improved U-Net segmentation model reaches the preset value; The content of the constructed rule engine is as follows: First, summarize and generate the performance indicator value Iot based on the performance indicator data set. The formula is as follows: In the formula, Ac, Pr, Re, Fs, Ma and Io are expressed as follows: Performance indicators Accuracy, precision, recall, F1 score, average accuracy, and intersection over union ratio in the dataset; Secondly, based on the amount of memory occupied by the orthophoto image, the sigmoid function is used to calculate the minimum performance orientation value Iot_min. The formula is as follows: Wherein, P0 is expressed as: the adjusted basic performance index value parameter, and P0 < 1; k, M, B and n are represented as: Adjustment coefficients, memory amounts, baseline memory amounts, and power exponents; Among them, k>0, n=1.
7. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The search optimization strategy implemented is as follows: Compare any type of data in the performance indicator data set with the corresponding threshold. If any type or types of data do not reach the corresponding threshold, perform optimization operations. The optimization operations at least include: Adjust model parameters, change model architecture, and hardware acceleration.
8. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 1, characterized in that: The spliced image is input into the trained improved U-Net segmentation model. The model predicts its category pixel by pixel and generates a binary or multi-class mask map of the plot segmentation, namely the semantic mask. The semantic mask assigns each pixel to a different category, including at least 8 categories, namely: sugarcane field, pond, river, house, tractor, car, road and rice field; among them, the background is represented by 0, and the 8 categories are represented by numbers 1 to 8 respectively.
9. A method for unmanned aerial vehicle sugarcane planting area image recognition according to claim 8, characterized in that: Conversion adjustments process the following: Use OpenCV and Shapely libraries to convert the semantic masks generated by the improved U-Net segmentation model into polygon outlines. Use the GDAL library to convert polygon outlines into GeoJSON format. During the conversion process, refer to the geographic information of the cropped image to convert pixel coordinates into geographic coordinate system. In the generated GeoJSON file, each polygon contains at least the following attributes: category and boundary.