A large-scale greenhouse interpretation method based on drone imagery

Through the multi-task greenhouse extraction network and Pix4D stitching technology, the problems of limited interpretation range of drone images and deterioration of legal quality of segmentation and grouping were solved, and high-precision large-scale land cover classification and greenhouse interpretation were achieved.

CN116958834BActive Publication Date: 2025-10-03INST OF ELECTRONICS & INFORMATION ENG OF UESTC IN GUANGDONG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310421509.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-10-03
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

The existing technology has a limited interpretation range for drone images, and the segmentation and grouping process has problems such as splicing traces and image quality degradation, which cannot meet the needs of large-scale land cover classification.

Method used

A multi-task greenhouse extraction network is used to perform feature extraction and multi-task prediction on drone small imagery, combined with Pix4D software for image stitching, and the registration and point cloud information of multispectral and visible light images are used to achieve large-scale land cover classification.

Benefits of technology

It improves the interpretation accuracy and image quality of large-scale land cover classification, reduces splicing traces, and improves the accuracy and consistency of greenhouse interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958834B_ABST
    Figure CN116958834B_ABST
Patent Text Reader

Abstract

The present invention discloses a large-scale greenhouse interpretation method based on drone images, which belongs to the field of image processing technology. The present invention first constructs and trains a multi-task greenhouse extraction network, which includes an image feature extraction network and a multi-task prediction network, and obtains a greenhouse interpreter based on the trained image feature extraction network and the greenhouse segmentation task prediction branch. A first multispectral image and a first visible light image are obtained by aligning multispectral images of different channels; and the first multispectral image is interpreted by the greenhouse interpreter to obtain a small image prediction result; the first visible light image is spliced ​​to obtain a large visible light image, and the small image prediction result is used as the spliced ​​input image, and the corresponding camera posture information and point cloud information of the first visible light image are used as the camera posture of the spliced ​​input image and the point cloud information between images, and image splicing processing is performed to obtain a greenhouse interpretation result of the large visible light image. The present invention can have high-precision registration performance in a variety of scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a large-scale greenhouse interpretation method based on drone images. Background Art

[0002] Land cover classification is one of the important applications of drone imagery. In applications such as surveying and mapping, agriculture, etc., it is necessary to interpret large-scale images at the pixel level to obtain the category of each pixel in the large-scale image in order to obtain accurate distribution of land objects, etc. One of the important categories of land cover is greenhouses. Greenhouses create a small environment to grow food, vegetables, fruits, etc. Greenhouses enable agricultural production to overcome adverse natural conditions and greatly increase crop yields. As of the end of 2020, my country's greenhouse area was 1.873 million hectares, which is of great significance to ensuring my country's food security. Greenhouses have greatly improved agricultural productivity and have become an important part of agriculture. Existing agricultural production is inseparable from the management of greenhouses. Therefore, in land cover classification, accurate and large-scale greenhouse classification is very important, which is related to national food security, effective use of cultivated land, ecological environmental protection, rural economic development, etc.

[0003] Most current research interprets small images captured by drones, interpreting only a small area within the sensor's field of view during a single drone shot. It's important to note that the field of view of a drone's sensor during a single shot is limited, and the coverage of a small drone image is very limited. Although increasing the drone's flight altitude can achieve a larger field of view, the image resolution decreases accordingly. Furthermore, greater sensor parallax results in greater image distortion at the edges. Given that land cover classification typically requires interpretation of a wide range of features, methods based on small image interpretation can only produce results for small areas, failing to meet practical requirements. Therefore, research on large-scale land cover classification methods is crucial.

[0004] Currently, the most common method for interpreting large images is the segmentation and grouping method. This involves splitting the large image into multiple smaller images, then predicting each of these images separately. For example, a 6000×2000 image can be split into 6×2 1000×1000 smaller images. Each image is interpreted separately using MTGEN, and the interpretation results of the smaller images are then combined based on their position within the larger image to form the interpretation of the larger image.

[0005] However, the slicing and grouping method produces noticeable stitching artifacts, resulting in poor image quality. This is because MTGEN doesn't have enough information at the edges of the small images when interpreting them, resulting in poor prediction results at these edges. Consequently, when the predictions are combined into a larger image, the predictions at the junctions between the small images are poor, resulting in noticeable stitching artifacts. Furthermore, the image quality of the large image after Pix4D stitching is lower than that of the original image, and this lower image quality also reduces MTGEN's prediction performance. In summary, the slicing and grouping method suffers from two issues: noticeable stitching artifacts and reduced image quality. Summary of the Invention

[0006] The purpose of the present invention is to achieve large-scale land cover classification based on drone images, so as to address the problems of limited coverage and the existence of splicing gaps in the existing segmentation groups when interpreting drone thumbnails.

[0007] The technical solution adopted in the present invention is:

[0008] A method for interpreting a large-scale greenhouse image based on drone imagery includes the following steps:

[0009] Step 1: Greenhouse interpretation based on drone thumbnail images, i.e. building and training a multi-task greenhouse extraction network;

[0010] A multi-task greenhouse extraction network is used to perform image feature extraction and multi-task prediction on UAV images with image size smaller than a specified value;

[0011] The multi-task greenhouse extraction network includes an image feature extraction network and a multi-task prediction network, wherein the image feature extraction network is used to extract image features of the input image, and the multi-task prediction network includes three task prediction branches: a greenhouse segmentation task prediction branch, a greenhouse edge segmentation task prediction branch, and a greenhouse edge direction segmentation task prediction branch. The three task prediction branches respectively perform greenhouse segmentation task prediction, greenhouse edge segmentation task prediction, and greenhouse edge direction segmentation task prediction based on the extracted image features;

[0012] Based on the set training data set, the multi-task greenhouse extraction network is trained and learned by using cross entropy loss on the network parameters of the multi-task greenhouse extraction network. When the preset training end condition is met, the training ends, and a greenhouse interpreter is obtained based on the trained image feature extraction network and the greenhouse segmentation task prediction branch in the multi-task prediction network;

[0013] That is, in this invention, the greenhouse edge and direction prediction tasks only provide more supervision during training, resulting in better greenhouse segmentation results. After network training, during the inference phase, only the greenhouse segmentation prediction task is used to determine the category of each pixel (greenhouse or non-greenhouse). The greenhouse edge and direction prediction tasks are not used during the inference phase.

[0014] Step 2: Large-scale land cover classification based on drone imagery:

[0015] Step 201, registering multispectral images of different channels to obtain a first multispectral image and a first visible light image of a specified size after registration, wherein the image spatial sizes of the first multispectral image and the first visible light image are consistent and match the spatial size of the input image of the multi-task greenhouse extraction network;

[0016] The greenhouse interpreter obtained in step 1 interprets the first multispectral image to obtain a prediction result of the first multispectral image (i.e., a small image prediction result);

[0017] Step 202: performing image stitching on the obtained first visible light images to obtain a second visible light image (large visible light image), and saving the camera pose information and point cloud information of each first visible light image obtained during the stitching process;

[0018] In step 203, the prediction result of the first multispectral image is used as the stitched input image for image stitching, and the camera pose information and point cloud information of the corresponding first visible light image are used as the camera pose and point cloud information between the stitched input image. The prediction result of the first multispectral image is subjected to image stitching processing to obtain the greenhouse interpretation result corresponding to the second visible light image.

[0019] Furthermore, in step 1, the greenhouse segmentation task prediction of the multi-task prediction network is used to perform binary classification of greenhouses and backgrounds at the pixel level; the greenhouse edge segmentation task prediction of the multi-task prediction network is used to perform binary classification of greenhouse edges and backgrounds at the pixel level; the greenhouse edge direction segmentation task prediction of the multi-task prediction network is used to perform multi-classification of backgrounds and multiple greenhouse edge directions at the pixel level; among them, multiple greenhouse edge directions refer to dividing the value range of the greenhouse edge direction into several direction categories at equal intervals.

[0020] Furthermore, in step 1, the prediction results of the greenhouse segmentation task prediction, greenhouse edge segmentation task prediction, and greenhouse edge direction segmentation task prediction of the multi-task prediction network are respectively subjected to the Softmax function to obtain the output probability of each category.

[0021] Furthermore, in step 1, the sample labels of the training data set are set as follows:

[0022] The sample labels predicted in the greenhouse segmentation task use manually labeled greenhouse masks;

[0023] The sample labels predicted for the greenhouse edge segmentation task are: the edges of the labeled greenhouse mask are calculated using the Canny operator;

[0024] The sample labels predicted by the greenhouse edge direction segmentation task are: the vertical and horizontal gradient values ​​at the greenhouse edge are calculated using the Sobel operator, and then the edge direction is obtained using the inverse cosine function.

[0025] Furthermore, in step 201 , the first multispectral image includes 5 channels, and the first visible light image includes 3 channels.

[0026] Furthermore, in step 202, the image stitching of the first visible light image is specifically performed as follows:

[0027] Calculate and save the camera posture information of each first visible light image, including position parameters and rotation angle parameters;

[0028] Calculate and save the point cloud information of each first visible light image;

[0029] An orthophoto of the second visible light map is generated based on the point cloud information.

[0030] The technical solution provided by the present invention brings at least the following beneficial effects:

[0031] Through a large number of tests, the interpretation accuracy of the greenhouse of the present invention is better than the existing methods, and it can achieve high-precision registration performance in a variety of scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 This is an overall framework diagram of the multi-task greenhouse extraction network in an embodiment of the present invention;

[0034] Figure 2 Schematic diagram of qualitative results of greenhouse segmentation based on multispectral imagery in an embodiment of the present invention on a small image;

[0035] Figure 3 This is a framework diagram for large-scale land cover classification based on drone images in an embodiment of the present invention;

[0036] Figure 4 Schematic diagram of the qualitative results of LISP in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0038] Most current research interprets small images captured by drones, interpreting only a small area within the sensor's field of view during a single drone shot. It's important to note that the field of view of a drone's sensor during a single shot is limited, and the coverage of a small drone image is very limited. Although increasing the drone's flight altitude can achieve a larger field of view, the image resolution decreases accordingly. Furthermore, greater sensor parallax results in greater image distortion at the edges. Given that land cover classification typically requires interpretation of a wide range of features, methods based on small image interpretation can only produce results for small areas, failing to meet practical requirements. Therefore, research on large-scale land cover classification methods is crucial.

[0039] Currently, the most commonly used method for interpreting large images is the segmentation and grouping method. The process is as follows: the large image is divided into multiple smaller images, and then predictions are made on each of these smaller images. For example, a 6000×2000 image can be divided into 6×2 1000×1000 smaller images. Each smaller image is interpreted separately using MTGEN. The interpretations of the smaller images are then combined based on their position within the larger image to form the interpretation of the larger image. However, the segmentation and grouping method produces noticeable stitching artifacts, resulting in poor image quality. This is because MTGEN does not have sufficient information at the edges of the smaller images, resulting in poor predictions at these edges. Consequently, when the predictions of the smaller images are combined into the larger image, the predictions at the boundaries between the smaller images are poor, resulting in noticeable stitching artifacts. Furthermore, the image quality of the resulting Pix4D stitched image is degraded compared to the original image, and this reduced image quality also leads to reduced MTGEN prediction performance. In summary, the segmentation and grouping method suffers from two issues: noticeable stitching artifacts and reduced image quality.

[0040] The process of the present invention is divided into two parts: a greenhouse interpretation step based on small drone imagery and a large-scale land cover classification step based on drone imagery.

[0041] Greenhouse interpretation steps based on drone thumbnails (Multi-Task Greenhouse Extraction Network): Considering the shape characteristics of greenhouses in remote sensing images, we can strengthen the supervision of greenhouse edges to enhance the network's learning of shapes. The overall framework of the Multi-Task Greenhouse Extraction Network (MTGEN) used in this invention is as follows: Figure 1As shown in Figure 2, MTGEN is a fully convolutional network. The entire process can be divided into three parts: image feature extraction module, multi-task prediction module, and loss calculation module.

[0042] First, the input image is passed through the image feature extraction module (preferably the backbone network HRNet) to extract image features, resulting in a feature map of a specified size (e.g., 180, 560, 560, where 180 represents the channel dimension of the feature map and 560×560 represents the input spatial size). The final feature map size can be set to 560×560, which is consistent with the input image spatial size. The input image size can be adjusted based on the computer's hardware resources.

[0043] Then, since a greenhouse is generally a regular long strip, that is, a greenhouse generally has four sides, and the direction of each side is basically the same, three tasks are designed in the embodiment of the present invention to better constrain the optimization direction of the network, including the greenhouse segmentation task (main task), the greenhouse edge segmentation task (auxiliary task 1) and the greenhouse edge direction segmentation task (auxiliary task 2).

[0044] Finally, each prediction task has its corresponding true label, and all three tasks use cross entropy loss to constrain the optimization network.

[0045] As a possible implementation method, in an embodiment of the present invention, the greenhouse interpretation step based on the drone thumbnail image is specifically implemented as follows:

[0046] (1) Image feature extraction

[0047] Considering the characteristics of greenhouses in drone images, this embodiment uses HRNet as the backbone network, mainly based on the following two reasons: (1) HRNet, as a backbone network, is widely used in many fields such as semantic segmentation and target detection, and has achieved good results; (2) HRNet can always maintain high-resolution feature maps, thereby obtaining high-resolution output results. Considering that many greenhouses are slender and only about 3-5 pixels wide, the high-resolution output of HRNet is very important for greenhouse segmentation. The core of HRNet's ability to maintain high-resolution output lies in the combination of feature maps of different resolutions. In order to obtain a new feature map, the low-resolution feature map is upsampled using bilinear interpolation, and then the number of channels and size are adjusted using 1x1 convolution; the high-resolution feature map is downsampled using convolution with a stride of 2, and then the number of channels and size are adjusted using 1x1 convolution; feature maps of the same resolution do not need to be upsampled or downsampled and can be directly connected to the convolution operation. The above three feature maps are fused by adding to obtain a new feature map. In addition, although the HRNet network is large and the inference time is long, the greenhouse segmentation task is generally performed on the server, and the inference time requirement is not high.

[0048] (2) Multi-task prediction

[0049] As shown in Table 1, the input features of multi-task prediction are the output features of HRNet, with a size of (180, 560, 560). Based on the input features, MTGEN uses three sets of convolutions to predict three different tasks in parallel: greenhouse prediction, greenhouse edge prediction, and greenhouse edge direction prediction.

[0050] Greenhouse prediction indicates the probability that a pixel belongs to a greenhouse, and its prediction results have two categories (greenhouse and background). When predicting greenhouses, MTGEN uses convolutional layers of (180, 3, 3, 45) and (45, 1, 1, 2) to convolve the input features, respectively. (180, 3, 3, 45) means using a 3×3 convolution to reduce the input feature dimension from 180 to 45, resulting in an output feature map size of (45, 560, 560). (45, 1, 1, 2) means using a 1×1 convolution to obtain the prediction result, resulting in a prediction size of (2, 560, 560).

[0051] Greenhouse edge prediction indicates whether a pixel belongs to a greenhouse edge, with two prediction categories (greenhouse edge and background). For greenhouse edge prediction, MTGEN convolves the input features with (180, 3, 3, 45) and (45, 1, 1, 2), respectively, resulting in a prediction size of (2, 560, 560). Greenhouse edge direction prediction indicates the edge direction of a pixel, with 37 prediction categories: 0 indicates that the pixel is background, and 1-36 indicate that the pixel's edge direction lies in the interval [0°, 10°), [10°, 20°), ... [350°, 360°), respectively.

[0052] When predicting the greenhouse edge direction, MTGEN uses (180, 3, 3, 45) and (45, 1, 1, 37) to convolve the input features, and the prediction result size is (37, 560, 560).

[0053] In the embodiment of the present invention, the true labels of the three prediction tasks are obtained as follows: greenhouse prediction requires manual labeling of the greenhouse mask; greenhouse edge prediction uses the Canny operator to calculate the edge based on the greenhouse mask; the greenhouse edge direction uses the Sobel operator to calculate the vertical and horizontal gradient values ​​at the greenhouse edge, and then uses the inverse cosine function to obtain its edge direction. The prediction results of the three tasks are respectively passed through the Softmax function to obtain the output probability of each category. Softmax can normalize the value to 0-1 and ensure that the sum of the probabilities of each category is 1. The larger the value of the prediction result, the greater the probability of belonging to that category. Softmax is shown in formula (1):

[0054]

[0055] Where M represents the number of categories of prediction results; z i represents the predicted value of the i-th category; p i represents the predicted probability of the i-th class.

[0056] Table 1 Structure of the result prediction part of the multi-task greenhouse extraction network

[0057]

[0058] (3) Loss calculation

[0059] Greenhouse prediction, greenhouse edge prediction and greenhouse edge direction prediction are all pixel-level segmentation tasks, and their loss functions are denoted as L mask , L edge and L ori These three losses are cross entropy losses, and their calculations are shown in formula (2):

[0060]

[0061] Where N represents the number of pixels in the entire image, i represents a specific pixel, M represents the number of categories, and L mask and L edge The number of categories is 2, L ori The number of categories is 37, y is an indicator of either 0 or 1, y ic Indicates whether the true label of pixel i is category c, p ic is the probability predicted by MTGEN that the pixel i belongs to category c, w c represents the loss weight of category c.

[0062] Taking into account the proportion of different categories, L mask The loss weights of the greenhouse and background are both 1, L edge The loss weights of the edge and background of the greenhouse are 1 and 10, respectively. ori The loss weight of the middle background is 1, and the weights of the 36 different edge directions are all 360. The total loss of the entire network is L = L mask +L edge +L ori , the loss weight of each task is 1. Figure 2 The qualitative results of greenhouse segmentation based on multispectral imagery are presented in multiple scenarios, including greenhouses of different shapes and sizes and different backgrounds. Figure 2 (a), (b), and (c) represent the prediction results for a slender greenhouse, (d) represents the prediction result for a greenhouse in a high-reflection area caused by water, and (e) and (f) represent the prediction results for a greenhouse with a larger area.

[0063] If a small drone image is fed into MTGEN as input, the probability of each pixel in the image being a greenhouse can be directly obtained. However, due to the very limited range of the small image, interpreting the small image can only provide interpretation results for a very small area. Therefore, directly interpreting the small image with MTGEN cannot meet the needs of users who need to interpret objects over a larger area. Users can use Pix4D to stitch together many small drone images covering a larger area into a large image, but the large image cannot be directly fed into MTGEN for interpretation. The stitched large image is large, and the computer's video memory capacity is limited. If the large image is fed directly into MTGEN, the network will fail to interpret the image due to exceeding the video memory capacity. In other words, when using MTGEN to perform inference on a large image, the video memory required far exceeds the computer's video memory capacity, resulting in prediction failure for the large image.

[0064] Currently, the most commonly used method for interpreting large images is the segmentation and grouping method. The process is as follows: the large image is divided into multiple smaller images, and then predictions are made on each of these smaller images. For example, a 6000×2000 image can be divided into 6×2 1000×1000 smaller images. Each smaller image is interpreted separately using MTGEN. The interpretations of the smaller images are then combined based on their position within the larger image to form the interpretation of the larger image. However, the segmentation and grouping method produces noticeable stitching artifacts, resulting in poor image quality. This is because MTGEN does not have sufficient information at the edges of the smaller images, resulting in poor predictions at these edges. Consequently, when the predictions of the smaller images are combined into the larger image, the predictions at the boundaries between the smaller images are poor, resulting in noticeable stitching artifacts. Furthermore, the image quality of the resulting Pix4D stitched image is degraded compared to the original image, and this reduced image quality also leads to reduced MTGEN prediction performance. In summary, the segmentation and grouping method suffers from two issues: noticeable stitching artifacts and reduced image quality.

[0065] In order to solve the two problems in the segmentation and grouping method, the embodiment of the present invention proposes a large image interpretation method based on Pix4D (remote sensing data processing software) to splice small images (a large image interpretation method based on small image interpretation and Pix4D, LISP). LISP has made targeted improvements to the problems of the segmentation and grouping method. Its overall framework is as follows Figure 3 As shown, it includes three main steps, that is, the large-scale land cover classification steps based on drone images provided by the embodiment of the present invention specifically include the following steps:

[0066] A: Thumbnail Prediction: First, we register multispectral images from different channels to obtain the registered multispectral thumbnail (5 channels) and visible light thumbnail (3 channels). Then, we use MTGEN to interpret the multispectral thumbnail to obtain the thumbnail prediction result.

[0067] B: Spatial information calculation of small images: Use Pix4D to stitch the small visible light images together to create a large visible light image. This process uses Pix4D to analyze the camera pose of each small image and the point cloud information between the small images.

[0068] Pix4D's process of stitching small images into a large image can be divided into three sub-steps. First, Pix4D calculates the camera pose for each small image. Each small image has its own corresponding camera pose. The camera pose, also known as the camera extrinsic parameters, consists of six parameters: three position parameters (altitude, longitude, and latitude) and three rotation angle parameters. By calculating the camera pose, we can determine the relative positional relationship between the images, which serves as the basis for subsequent point cloud solution. Next, Pix4D calculates the point cloud information between the small images. Generally speaking, a point cloud is a collection of points in a 3D coordinate system, each point containing a 3D coordinate and color information. During 3D reconstruction, point clouds are generated through matching synonyms and aerial triangulation. Since users only need to generate orthophotos and do not need height information or 3D point information, the point cloud solution here only requires a set of synonyms. In other words, the point cloud only stores the coordinate position of a 3D point in different images. Finally, Pix4D uses this point cloud information to generate the orthographic image of the large image. Pix4D uses the point cloud information between the small images as a basis, fusing the information from different images and generating an orthophoto through mapping and texturing. This step yields the camera pose information and point cloud information between the small images, as well as the resulting stitched large image. The camera pose information and point cloud information are used in the third step.

[0069] C: Thumbnail Prediction Results Stitching: When using Pix4D to stitch thumbnail prediction results, you need to provide the input image, the image's camera pose, and the point cloud information between the images. The input image is the greenhouse probability of the thumbnail predicted by MTGEN. That is, the value of each pixel in the input image is the probability of that pixel being a greenhouse. The camera pose information and point cloud information can be directly used from the second step, without repeated calculation. This is because the camera pose and point cloud information for the thumbnail probability are the same as the thumbnail information in the second step. It should be noted that compared to the visible light thumbnail, the thumbnail prediction result only contains greenhouse probability information and loses information such as the texture and shape of the ground object. Therefore, directly using the thumbnail prediction results to calculate the camera pose and point cloud information is not very accurate. Considering that the thumbnail prediction results correspond one-to-one with the visible light thumbnail used in the second step, this embodiment of the present invention pre-calculates the thumbnail's spatial information and point cloud information using the visible light thumbnail. This information is provided to Pix4D for stitching the thumbnail prediction results. After stitching the thumbnail greenhouse probabilities using Pix4D, an orthophoto of the greenhouse probability over a larger area can be obtained. In this embodiment of the present invention, a threshold of 0.5 is used to perform binary classification to obtain the greenhouse segmentation result for the large image. That is, when the greenhouse probability is greater than or equal to 0.5, the pixel is considered to be a greenhouse; otherwise, the pixel is considered to be background. Through the above three steps, the probability of splicing small images using Pix4D can be used to obtain the large image interpretation result. It is worth mentioning that LISP is a large image interpretation method that is applicable not only to the greenhouse interpretation task of the embodiment of the present invention, but also to other large image interpretation tasks.

[0070] Figure 4 The greenhouse interpretation effect of LISP in more scenarios is demonstrated, among which, Figure 4 Figures (4-a) and (4-b) show the greenhouse interpretation results in different scenarios. The first row of images in (4-a) and (4-b) shows a large visible light image of the greenhouse, obtained by stitching original small images collected by drones using Pix4D. The second row shows the greenhouse interpretation results using the LISP method. LISP has a good interpretation effect on both densely packed, elongated greenhouses and relatively wide greenhouses. In summary, after using multispectral imagery and the MTGEN network to interpret the small images, using LISP to obtain the segmentation results of the large image can effectively interpret the large image, thereby obtaining large-scale interpretation results.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

[0072] The above are only some embodiments of the present invention. For those skilled in the art, several modifications and improvements can be made without departing from the inventive concept of the present invention, which all fall within the scope of protection of the present invention.

Claims

1. A large-scale greenhouse interpretation method based on drone images, characterized by: The following steps are involved: Step 1: Build and train a multi-task greenhouse extraction network; A multi-task greenhouse extraction network is used to perform image feature extraction and multi-task prediction on UAV images with image size smaller than a specified value; The multi-task greenhouse extraction network includes an image feature extraction network and a multi-task prediction network, wherein the image feature extraction network is used to extract image features of the input image, and the multi-task prediction network includes three task prediction branches: a greenhouse segmentation task prediction branch, a greenhouse edge segmentation task prediction branch, and a greenhouse edge direction segmentation task prediction branch. The three task prediction branches respectively perform greenhouse segmentation task prediction, greenhouse edge segmentation task prediction, and greenhouse edge direction segmentation task prediction based on the extracted image features; Based on the set training data set, the multi-task greenhouse extraction network is trained and learned by using cross entropy loss on the network parameters of the multi-task greenhouse extraction network. When the preset training end condition is met, the training ends, and a greenhouse interpreter is obtained based on the trained image feature extraction network and the greenhouse segmentation task prediction branch in the multi-task prediction network; Step 2: Large-scale land cover classification based on drone imagery: Step 201, registering multispectral images of different channels to obtain a first multispectral image and a first visible light image of a specified size after registration, wherein the image spatial sizes of the first multispectral image and the first visible light image are consistent and match the spatial size of the input image of the multi-task greenhouse extraction network; The greenhouse interpreter obtained in step 1 interprets the first multispectral image to obtain a prediction result of the first multispectral image; Step 202: performing image stitching on the obtained first visible light images to obtain a second visible light image, and saving the camera pose information and point cloud information of each first visible light image obtained during the stitching process; In step 203, the prediction result of the first multispectral image is used as the stitched input image for image stitching, and the camera pose information and point cloud information of the corresponding first visible light image are used as the camera pose and point cloud information between the stitched input image. The prediction result of the first multispectral image is subjected to image stitching processing to obtain the greenhouse interpretation result corresponding to the second visible light image.

2. The method according to claim 1, wherein In step 1, the greenhouse segmentation task prediction of the multi-task prediction network is used to perform binary classification of greenhouses and backgrounds at the pixel level; the greenhouse edge segmentation task prediction of the multi-task prediction network is used to perform binary classification of greenhouse edges and backgrounds at the pixel level; the greenhouse edge direction segmentation task prediction of the multi-task prediction network is used to perform multi-classification of backgrounds and multiple greenhouse edge directions at the pixel level; among them, multiple greenhouse edge directions refer to dividing the value range of the greenhouse edge direction into several direction categories at equal intervals.

3. The method according to claim 1, wherein In step 1, the prediction results of the greenhouse segmentation task prediction, greenhouse edge segmentation task prediction, and greenhouse edge direction segmentation task prediction of the multi-task prediction network are respectively subjected to the Softmax function to obtain the output probability of each category.

4. The method according to claim 1, wherein In step 1, the sample labels of the training dataset are set as follows: The sample labels predicted in the greenhouse segmentation task use manually labeled greenhouse masks; The sample labels predicted for the greenhouse edge segmentation task are: the edges of the labeled greenhouse mask are calculated using the Canny operator; The sample labels predicted by the greenhouse edge direction segmentation task are: the vertical and horizontal gradient values ​​at the greenhouse edge are calculated using the Sobel operator, and then the edge direction is obtained using the inverse cosine function.

5. The method according to claim 1, wherein In step 201 , the first multispectral image includes 5 channels, and the first visible light image includes 3 channels.

6. The method according to claim 1, wherein In step 202, the image stitching of the first visible light image is specifically performed as follows: Calculate and save the camera posture information of each first visible light image, including position parameters and rotation angle parameters; Calculate and save the point cloud information of each first visible light image; An orthophoto of the second visible light map is generated based on the point cloud information.

Citation Information

Patent Citations

  • Height estimation and semantic segmentation multi-task prediction method for single-view remote sensing image

    CN115546649A

  • A method for aerial imagery acquisition and analysis

    WO2017077543A1