Crop identification system for aerial photography
Patent Information
- Application Number
- TW114100197
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2026-07-16
- Estimated Expiration
- 2045-01-02
AI Technical Summary
Traditional crop identification methods using the Normalized Difference Vegetation Index (NDVI) require near-infrared sensing, leading to high costs and computational inefficiencies in convolutional neural networks (CNNs), and gradient vanishing issues hinder optimal training results.
A drone-based crop identification system utilizing a receiving module, data processing module, and deep learning module with a Residual U-Net Architecture to process visible light images, incorporating adaptive thresholding and residual blocks to stabilize gradients and improve training efficiency.
The system achieves accurate and efficient crop type identification and growth status assessment in visible light images, overcoming gradient vanishing and computational inefficiencies, with high accuracy and reduced training time.
Abstract
Description
[Technical Field]
[0001] This invention provides a crop identification technology based on aerial photography, specifically a technology for identifying crops by analyzing aerial photography data using deep learning. [Previous Technology]
[0002] Traditional crop identification is a crucial technology for crop growth management in smart agriculture. However, the Normalized Difference Vegetation Index (NDVI), commonly used in traditional techniques, requires sensing values of near-infrared light (outside the visible light range), resulting in high application costs and hindering widespread adoption. Furthermore, in traditional convolutional neural networks, gradients may gradually decrease and disappear as the number of network layers increases, making it difficult for the convolutional neural network to achieve the desired training effect. Increased network layers may also lead to degradation. Thus, the computational performance of general convolutional neural networks not only fails to improve but may even decrease, wasting computation time. [Summary of the Invention]
[0003] Based on the need for crop identification technology using visible light, this invention provides a drone crop identification system, comprising: a receiving module for receiving a visible light image of a target area generated by a drone; a data processing module for generating vegetation index information corresponding to the visible light image; and a deep learning module for receiving the vegetation index information and generating a deep learning model based on the Residual U-Net Architecture to process the vegetation index information, identify crop types in the visible light image, and establish a label segmentation image for the crop types. The receiving module, the data processing module, and the deep learning module are connected by signal lines.
[0004] In one embodiment, the planting area of a crop type includes a field boundary in the target area, and the data processing module defines at least one field segmentation image of the crop type in the visible light image based on the field boundary. The vegetation index information includes field vegetation index information generated based on at least one field segmentation image.
[0005] In one embodiment, the field boundary can be defined based on the field information of this crop type in the target area from the geographic information system (GIS) data.
[0006] In one embodiment, the data processing module generates field vegetation index information for each field segment in a visible light image by means of an adaptive thresholding operation or a light intensity normalization operation.
[0007] In one embodiment, the deep learning module generates a labeled segmented image of crop type or crop growth status based on field vegetation index information.
[0008] In one embodiment, the deep learning model of the RES U-NET architecture includes: an encoder, a bridge layer, and a decoder, wherein the encoder generates multiple encoding computation layers with successively reduced spatial resolution based on field vegetation index information, and the encoder further includes a residual block for transmitting data across encoding computation layers.
[0009] In one embodiment, at least one of the bridging layer and the decoder further includes another residual block for transmitting data across the bridging layer or the decoding computation layer.
[0010] In one embodiment, the deep learning module further includes a model training unit. The model training unit inputs a field segmentation image of a known crop type in the target area and the field vegetation index information of the corresponding field segmentation image into the deep learning module to train the deep learning model to identify the crop type in the field segmentation image and establish a label segmentation image of the crop type, and generate the identification accuracy or loss function. The deep learning model adjusts its parameters to improve the accuracy or reduce the identification loss function.
[0011] In one embodiment, the vegetation index information includes: Green Red Vegetation Index (GRVI), Visible Atmospherically Resistant Index (VARI), Excess Green Index (ExG), Excess Red Index (ExR), Triangular Greenness Index (TGI), or Color Index of Vegetation Extractable (CIVE).
[0012] In one embodiment, the aerial photography vehicle includes: a drone or a satellite.
[0013] In one embodiment, the deep learning module establishes a species label for a crop type in a visible light image, and the label segmentation image of the species label corresponds to the distribution range of the crop type in the visible light image.
[0014] According to one perspective, the present invention provides a drone crop identification system, comprising: a receiving module for receiving a visible light image generated by a drone over a target area; a data processing module for generating a Green Red Vegetation Index (Green Red Vegetation Index) corresponding to the visible light image; and a deep learning module for receiving the Green Red Vegetation Index and generating a deep learning model based on the RESU U-NET architecture to process the Green Red Vegetation Index, identify crop types in the visible light image, and establish crop type label segmentation images. The receiving module, the data processing module, and the deep learning module are connected by signal lines.
[0015] According to one viewpoint, the present invention provides a method for identifying crops by aerial photography, comprising: providing an aerial photography vehicle to generate a visible light image for a target area, the target area including at least one field segmentation image of a known crop type for training; providing a data processing module to generate at least one field vegetation index information of the at least one field segmentation image based on the visible light image; and providing a deep learning module to receive the field vegetation index information and generate a deep learning model based on the RES U-NET architecture for processing the field vegetation index information, identifying crop types in the field segmentation image, and establishing a label segmentation image of the crop type, wherein the deep learning model includes: an encoder, a bridging layer, a decoder, and at least one residual block, the residual block being used to transmit data across multiple encoding computation layers of the encoder, the bridging layer, or multiple decoding computation layers of the decoder.
[0016] In one embodiment, the encoding computation layer, bridging layer and decoding computation layer include multiple batch normalization layers, multiple rectified linear unit layers and multiple convolutional layers. The aerial crop identification method also provides a gradient judgment module. When the gradient judgment module determines that gradient vanishing occurs in the batch normalization layer, rectified linear unit layer and convolutional layer, or gradient fluctuation occurs between adjacent batch normalization layers, the residual block transmits data across the batch normalization layer, rectified linear unit layer or convolutional layer.
[0017] The aforementioned data can be represented in various forms such as scalars, vectors, tensors, and matrices.
Implementation Method
[0018] The foregoing description and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the preferred embodiments with reference to the accompanying drawings.
[0019] Referring to Figure 1, regarding the aforementioned technical needs, the present invention provides a drone crop identification system 100, comprising: a receiving module 10, receiving a visible light image Imv generated by a drone Av for a target area Ta (Figure 2 illustrates the drone Av taking a picture of the target area Ta); a data processing module 20, generating vegetation index information Itv corresponding to the visible light image Imv; and a deep learning module 30, receiving the vegetation index information Itv, generating a deep learning model 30m according to the Residual U-Net Architecture, for processing the vegetation index information Itv, identifying crop species Ct in the visible light image Imv, and establishing a label segmentation image Imsl for the crop species Ct. The receiving module 10, the data processing module 20, and the deep learning module 30 are connected by signal lines. The deep learning model 30m can identify not only the crop species Ct in the visible light image Imv, but also the growth status of the crop. The labeled segmented image Imsl is a distribution area in the visible light image Imv containing labels established based on the crop species Ct, or labels representing the growth status of the crop species Ct. This area is distributed across the segmented image range in the visible light image Imv (refer to Figure 8, i.e., the labeled segmented image Imsl for crop species Ct, which will be explained in subsequent embodiments). The deep learning model 30m with the RES U-NET architecture has many advantages: it avoids gradient vanishing and accelerates the convergence of the deep learning model 30m using a smaller number of parameters, resulting in better deep learning performance, which will be described in detail in subsequent embodiments. The receiving module 10, data processing module 20, and deep learning module 30 of this invention are connected by signal lines; the signals are physical or chemical signals.
[0020] Referring to Figures 3 and 4, in one embodiment, the planting area of crop type Ct includes a field boundary in the target area Ta (in Figure 3, crop type Ct is taken as bananas in a banana field; in Figure 4, the banana field in Figure 3 is surrounded by a thick black border to illustrate the field boundary, and outside the field boundary is a palm tree field). The data processing module 20 defines at least one field segmentation image Imsf corresponding to bananas in the visible light image Imv based on the field boundary (refer to Figure 5, which can be divided into multiple small image blocks of 1024*1024 based on computational needs). (Figure 6, only the image of the banana field is retained, and the image of the palm tree field is not retained). The vegetation index information Itv includes the field vegetation index information Itvf generated based on the at least one field segmentation image Imsf. The field boundary is defined based on the relevant planting area of this banana in the target area Ta. Pre-defining field boundaries reduces interference from non-banana crop information during model training, enabling faster generation of crop labels within the planted fields. Simultaneously, this approach compares the intersection-over-union ratio between the planted fields within the field boundaries (banana fields in this example) and the generated label segmentation image Imsl, improving the accuracy of deep learning. Therefore, this approach can simultaneously determine the crop type Ct and the planted field of that crop type Ct. Furthermore, the distribution of crop type Ct2 is approximate, and the calculation of the intersection-over-union ratio is essentially an approximate value. In this embodiment, bananas are used as an example; users can change the crop type for analysis as needed.
[0021] In one embodiment, the field boundary can be defined based on the field information of crop type Ct in the target area Ta from the Geographic Information System (GIS) data. Specifically, this embodiment can use a combination of "field boundary + GRVI index" to create label data (crop type Ct, or the growth stage of crop type Ct) for each small map patch, dividing it into training and testing data for crop type Ct determination. In other embodiments, the field boundary can be determined using edge detection algorithms (gradient calculation or double thresholding), segmentation algorithms (thresholding, region growing, or K-Means clustering), or principal component analysis (converting multiple band data in the visible light image Imv into principal components to highlight the field boundary.
[0022] Referring to Figure 7, in one embodiment, the data processing module 20 generates field vegetation index information Itvf for each field segmentation image Imsf based on the visible light image Imv using an adaptive thresholding operation or a light intensity normalization operation. The vegetation index information Itv in this embodiment is generated based on the visible light image Imv. Various shooting conditions that generate the visible light image Imv, such as different times of day (noon or evening) or weather conditions (sunny or cloudy), will affect the stability of reflected light reception in the visible light image Imv. Alternatively, the visible light image Imv may be locally affected by lighting conditions, such as shadows, smoke, cloud interference, atmospheric aerosols, or water vapor. When the data processing module 20 generates field vegetation index information Itvf for each field segment image Imf in the visible light image Imv (e.g., the Green-Red Vegetation Index GRVI, whose value is based on the reflectance intensity of visible red and visible green light), it can correct the field vegetation index information Itvf of each field segment image Imf as needed through adaptive thresholding or light intensity normalization. The adaptive thresholding operation is specialized for processing images with uneven light distribution. In the adaptive thresholding operation, for each pixel of the visible light image Imv, a threshold is dynamically determined based on the pixel values in its neighborhood (local region), such as: local average, local weighted average, local median, etc. Then, each pixel in the visible light image Imv is compared with its calculated local threshold and binarized to determine the content of the field vegetation index information Itvf. Light intensity normalization is a technique used to reduce interference caused by changes in illumination (such as shadows, overexposure, and uneven brightness) in images, thereby making the crop reflectance in the images more prominent. The basic calculation principle of light intensity normalization is: I = Rf•Li; where I is the observed image intensity value, Rf is the crop reflectance, and Li is the ambient light intensity. In light intensity normalization, ambient light conditions (such as brightness, shadows, and ambient light) can be significantly reduced, allowing for more accurate calculation of crop reflectance. Furthermore, adaptive thresholding enhances the stability of vegetation index information (Itv, such as the green-red vegetation index) and reduces the impact of illumination changes on the monitoring results of crop reflectance. The green-red vegetation index is calculated as: GRVI = (G - R) / (G + R), where G is the reflectance value in the green band and R is the reflectance value in the red band.
[0023] In one embodiment, the deep learning module 30 generates a label segmentation image Imsl of crop type Ct or the growth status of crop type Ct based on the field vegetation index information Itvf (refer to Figure 8, where the white blank area shows the label distribution of the corresponding crop type Ct). Both images can be generated using the deep learning model 30m of the RESU-NET architecture. The crop's growth status includes physiological states (photosynthetic efficiency, respiration rate, etc.), growth stages (budding stage, seedling stage, rapid development stage of leaves and stems, reproductive growth stage, maturity stage, senescence stage, etc.), and health status (nutritional status, pests and diseases, leaf color, etc.).
[0024] Referring to Figure 9, in one embodiment, the deep learning model 30m according to the RES U-NET architecture includes: an encoder 31, a bridge layer 32, and a decoder 33. The encoder 31 generates multiple encoding computation layers 31L with successively reduced spatial resolution based on the field vegetation index information Itvf (received from the input of the deep learning model 30m). The encoder 31 may further include residual blocks for passing data across the encoding computation layers 31L. The shortcut connection of residual blocks across computation layers allows data to be more directly connected to previous layers in the backward propagation (i.e., in the opposite direction of the arrow), mitigating the gradient vanishing problem. Referring to Figure 10A, an embodiment of the residual block operation is illustrated, showing a portion of the original encoding computation layers 31L. Figure 10B shows the changes between the encoding computation layers 31L after the addition of residual blocks. Furthermore, the shortcut connections generated by the residual blocks provide a path that "skips" several layers, allowing the deep learning model 30m to flexibly utilize the parameters of each layer in the deep network and reducing degradation. Each encoding computation layer 31L contains feature data corresponding to each successively decreasing spatial resolution, where the spatial resolution (length per pixel) is higher the smaller the number. For example, the spatial resolution of the visible light image Imv produced by a high-pixel camera can be as low as 0.1 meters per pixel. The larger the value of meters per pixel, the lower the spatial resolution. The bridging layer 32 is connected to the encoding computation layer 31L with the lowest spatial resolution in the encoder 31 to extract global features of the field segmentation image Imsf. The decoder 33 receives the global features, successively upsamples them, and makes skip connections with the encoding computation layers 31L to generate multiple decoding computation layers 33L corresponding to the spatial resolution of each encoding computation layer 31L, thereby generating a label segmentation image Imsl with global features. The spatial resolution of the label segmentation image Imsl, representing global features, corresponds to the spatial resolution of the field segmentation image Imsf. The number of encoding computation layer 31L and decoding computation layer 33L can be set to a fixed value during the construction phase of the deep learning model 30m, or dynamically adjusted. This dynamic adjustment can be based on the feature complexity of the input data or the gradient distribution during training to address the vanishing gradient problem, improve training stability, and enhance feature fusion across computational layers.
[0025] Referring again to Figure 9, in one embodiment, the encoding computation layer 31L, the bridging layer 32, and the decoding computation layer 33L further include multiple batch normalization layers (BN), multiple rectified linear unit layers (ReLU), and multiple convolutional layers (Conv). The data from the batch normalization layers (BN) have a standard normal distribution with zero mean and unit variance, thereby improving the training stability and efficiency of the model. The rectified linear unit layers (ReLU) use an activation function to turn all data less than zero to zero, while keeping the positive values unchanged, in order to mitigate gradient vanishing. The convolutional layers (Conv) extract features from the data through convolution operations.
[0026] In one embodiment, at least one of the bridging layer 32 and the decoder 33 includes another residual block for transmitting data across the bridging layer 32 or the decoding computation layer 33L. The description of this other residual block is given in the aforementioned residual block section. The bridging layer 32 extracts global features from visible light image Imv or field segmentation image Imsf. The bridging layer 32 is located at the deepest part of the RES U-NET architecture, where gradients may decrease significantly. The other residual block of the bridging layer 32 can effectively transmit gradients and maintain the stability of feature learning. For example, skip connections in the residual block allow the RES U-NET architecture to perform different operations simultaneously (e.g., using different convolutional kernels to perform dimensionality increase and decrease operations respectively, generating features of different dimensions), thus efficiently compressing features and restoring decoding details. Regarding the other residual block of the decoder 33, this other residual block can fuse shallower features into the decoding process through skip connections, enhancing the restoration of decoding details. When the decoder 33 experiences gradient vanishing or gradient instability, resulting in ineffective feature updates, this additional residual block improves training performance by directly transmitting gradients. Therefore, the improvement effect of the additional residual block in the bridging layer 32 or the decoder 33 is significant. Importantly, in the foregoing embodiment, the number of residual blocks is illustrated using one as an example. In practice, the number of residual blocks can be adjusted as needed.
[0027] Referring to Figures 11A or 11B, in one embodiment, the deep learning module 30 further includes a model training unit 34. The model training unit 34 inputs a field segmentation image Imsf of a known crop type Ct for training within the target region Ta, and the corresponding field vegetation index information Itvf of the field segmentation image Imsf, into the deep learning module 30 to train the deep learning model 30m. This model identifies the crop type Ct in the field segmentation image Imsf and establishes a label segmentation image Imsl for this crop type Ct, generating the identification accuracy or loss function. The deep learning model 30m adjusts its parameters (e.g., the number of residual blocks and the number of computational layers, the upsampling and downsampling methods, the number of decoding and encoding layers, the weights during training, the calculation method of the loss function, etc.) to improve the accuracy or reduce the identification loss function. The accuracy can be determined based on the proportion of crop species Ct identified, while the loss function can be selected based on the task objective (e.g., image segmentation), data characteristics (e.g., classification imbalance, data noise), the weights of the field segmentation image Imsf, or the reduction in gradient values. After adjustment, the deep learning model 30m can be input with vegetation index information Itv (or field vegetation index information Itvf), field segmentation image Imsf, etc., to identify unknown crop species and create label segmentation images Imsl for other identified crop species Ct. In Figures 11A and 11B, the signal connection between the model training unit 34 and the deep learning model 30m is different. In Figure 11A, the model training unit 34 leads the training operation of the deep learning model 30m; in Figure 11B, the model training unit 34 assists the training operation of the deep learning model 30m.
[0028] Visible light is electromagnetic wave that can be seen by humans, and its wavelength range is generally in the range of 350–800 nm. Visible light vegetation indices are basically based on the reflectance characteristics of red, green, and blue light bands to estimate crop species (Ct), health status, growth conditions, or cover density. In one embodiment, the vegetation index information Itv includes: Green Red Vegetation Index (GRVI), Visible Atmospherically Resistant Index (VARI), Excess Green Index (ExG), Excess Red Index (ExR), Triangular Greenness Index (TGI), or Color Index of Vegetation Extractable (CIVE). If necessary, these indices can also be combined into a comprehensive vegetation index. For example, the Green Red Vegetation Index and the Color Index of Vegetation can be combined to identify the Ct of mixed-planted crops.
[0029] In one embodiment, the aerial vehicle Av includes: a drone (e.g., a high-altitude camera) or a satellite, primarily used to provide visible light images.
[0030] In one embodiment, the deep learning module 30 establishes a species label for crop species Ct in the visible light image Imv, and the label segmentation image Imsl of this species label corresponds to the distribution range of crop species Ct in the visible light image Imv.
[0031] According to another perspective, the present invention provides a drone crop identification system 100, comprising: a receiving module 10, receiving a visible light image Imv generated by a drone Av for a target area Ta; a data processing module 20, generating a Green Red Vegetation Index corresponding to the visible light image Imv; and a deep learning module 30, receiving the Green Red Vegetation Index information, generating a deep learning model 30m according to the RES U-NET architecture, for processing the Green Red Vegetation Index information, identifying crop species Ct in the visible light image Imv, and establishing a label segmentation image Imsl for crop species Ct. The receiving module 10, the data processing module 20, and the deep learning module 30 are connected by signal lines. For a description of each component, please refer to the relevant descriptions of the aforementioned modules; further details are omitted here.
[0032] Referring to FIG12, according to one viewpoint, the present invention provides a method for identifying crops by aerial photography, comprising: providing an aerial photography vehicle Av to generate a visible light image Imv for a target area Ta, the target area Ta including at least one field segmentation image Imsf for training a known crop type (S1); providing a data processing module 20 to generate at least one field vegetation index information Itvf of at least one field segmentation image Imsf based on the visible light image Imv (S2); and providing a deep learning module 30 to receive the field vegetation index information Itvf, generate a deep learning model 30m based on the RES U-NET architecture, for processing the field vegetation index information Itvf, identifying the crop type Ct in the field segmentation image Imsf, and establishing a label segmentation image Imsl for the crop type Ct (S3). The deep learning model 30m includes an encoder 31, a bridging layer 32, a decoder 33, and at least one residual block 40. The residual block is used to transmit data between multiple encoding computation layers 31L of the encoder 31, the bridging layer 32, or multiple decoding computation layers 33L of the decoder 33.
[0033] In one embodiment, the encoding computation layer 31L, bridging layer 32, and decoding computation layer 33L contain multiple batch normalization layers (BN), multiple rectified linear unit layers (ReLU), and multiple convolutional layers (Conv). The aerial crop identification method also provides a gradient judgment module. When the gradient judgment module determines that gradient vanishing occurs in the batch normalization layers (BN), rectified linear unit layers (ReLU), or convolutional layers (Conv), or gradient fluctuations occur between adjacent batch normalization layers (BN), residual blocks are used to transfer data across the encoding computation layer 31L, bridging layer 32, or the decoding computation layer 33L of the decoder 33. For example, the batch normalization layer (BN) has a standard normal distribution with zero mean and unit variance. When the input data distribution of the batch normalization layer (BN) varies too much or the gradient is close to zero, residual blocks are needed for connection. For example, when the gradient in the negative region of the ReLU layer in a linear unit layer is corrected to zero, some neurons may output zero for an extended period during training, resulting in a zero gradient. In this case, the activation function can be adjusted by inputting data from other layers through residual blocks to activate the gradient. Similarly, when the gradient of the Conv layer in a convolutional layer is zero, feature information may be lost. In this situation, inputting data across computational layers through residual blocks can help maintain gradient stability.
[0034] Referring to Figure 13, in the aforementioned embodiment, the process of step S3 in Figure 12 can be broken down as follows: providing a deep learning module to receive field vegetation index information, generating a deep learning model based on the RES U-NET architecture, and processing the field vegetation index information (step S31) and identifying crop types in the field segmentation image by the deep learning module, and establishing label segmentation images of crop types (step S5). In this embodiment, step S4 is added between steps S31 and S5: providing a gradient judgment module, when gradient vanishing occurs in the batch normalization layer, the corrected linear unit layer, and the convolutional layer, or when gradient fluctuations occur between adjacent batch normalization layers, residual blocks are used to transmit data between the batch normalization layer, the corrected linear unit layer, or the convolutional layer. During deep learning, residual block operations can be performed to accelerate the identification of crop types in the field segmentation image and the establishment of label segmentation images of crop types.
[0035] Referring to Figures 14A and 14B, which show that according to the technique of the present invention, after approximately 20 complete training iterations (epochs, one learning iteration of the entire training dataset), the accuracy can reach 95% and the loss function can be as low as 0.35. Subsequent training can continue to improve the accuracy, but the improvement slows down. In short, the deep learning technique of the present invention has the effect of quickly achieving stable learning.
[0036] In one embodiment, the aforementioned receiving module 10, data processing module 20, and deep learning module 30 may be disposed within at least one controller of at least one device. The controller may include: a central processing unit (CPU), a neural network processing unit (NPU), a tensor processing unit (TPU), a microcontroller (MCU), a programmable logic controller (PLC), an instruction set architecture processor (ISA), a microprocessor, an application specific integrated circuit (ASIC), a digital signal processor (DSP), an arithmetic logic unit (ALU), a graphics processing unit (GPU), an image signal processor (ISP), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), or other similar elements and combinations thereof. In one embodiment, the controller also includes a means of executing stored code or a circuit, such as a Network Interface Controller (NIC).
[0037] The technology in this case, through the aforementioned deep learning training and prediction, can be applied to the accurate identification and management of crops, the estimation of crop planting area, and the management of crop production stages.
[0038] The present invention has been described above with reference to preferred embodiments. However, the above description is only for the purpose of enabling those skilled in the art to easily understand the content of the present invention, and is not intended to limit the scope of the present invention or the disclosed technology. Any person skilled in the art can make equivalent embodiments by combining, modifying or altering the above-disclosed technical content without departing from the scope of the technical solution of this application. [Simplified Explanation of the Diagram]
[0039] Figure 1 illustrates a schematic diagram of an aerial crop identification system according to an embodiment of the present invention.
[0040] Figure 2 illustrates a schematic diagram of a drone taking pictures according to an embodiment of the present invention.
[0041] Figures 3 and 4 illustrate schematic diagrams of planting areas and field boundaries for crop types according to an embodiment of the present invention.
[0042] Figure 5 illustrates a schematic diagram of the conversion of visible light images into vegetation index information and field area outlines therein, according to an embodiment of the present invention.
[0043] Figure 6 illustrates a schematic diagram of vegetation index information and field segmentation image corresponding to a crop type according to an embodiment of the present invention.
[0044] Figure 7 illustrates a schematic diagram of a data processing module according to an embodiment of the present invention.
[0045] Figure 8 illustrates a schematic diagram of a label segmentation image corresponding to a crop type according to an embodiment of the present invention.
[0046] Figure 9 illustrates a schematic diagram of a deep learning model based on the RES U-NET architecture according to an embodiment of the present invention.
[0047] Figures 10A and 10B illustrate a flowchart of a crop identification method using aerial photography according to an embodiment of the present invention.
[0048] Figures 11A and 11B illustrate schematic diagrams of deep learning modules according to two embodiments of the present invention.
[0049] Figure 12 illustrates a flowchart of the aerial crop identification method according to an embodiment of the present invention.
[0050] Figure 13 illustrates a flowchart of the aerial crop identification method according to another embodiment of the present invention.
[0051] Figures 14A and 14B illustrate the relationship between the number of complete training cycles (Epoch) and accuracy and loss function, respectively, in one embodiment of the deep learning technology according to the present invention.
Claims
1. A drone crop identification system, comprising: a receiving module for receiving a visible light image of a target area generated by a drone; a data processing module for generating vegetation index information corresponding to the visible light image; and a deep learning module for receiving the vegetation index information, generating a deep learning model based on the RES U-NET architecture, for processing the vegetation index information, identifying crop types in the visible light image, and establishing a labeled segmentation image of the crop type; wherein... The planting area of this crop type includes a field boundary in the target area. The field boundary is defined based on the field information of this known crop type in the target area from Geographic Information System (GIS) data. The data processing module defines a field segmentation image of this crop type in the visible light image based on the field boundary. The vegetation index information includes field vegetation index information generated based on the field segmentation image. There is a signal connection between the receiving module, the data processing module, and the deep learning module.
2. The aerial crop identification system as described in claim 1, wherein the data processing module generates field vegetation index information of the field segmentation images based on the visible light images by means of an adaptive thresholding method or light intensity normalization.
3. The aerial crop identification system as described in claim 1, wherein the deep learning module generates the labeled segmented image of the crop species or the growth status of the crop species based on the field vegetation index information.
4. The aerial crop identification system as described in claim 1, wherein the deep learning model of the RES U-NET architecture includes: an encoder, a bridge layer, and a decoder, wherein the encoder generates multiple encoding computation layers with progressively decreasing spatial resolution based on the field vegetation index information, wherein the encoder further includes a residual block for transmitting data across the encoding computation layers; the bridge layer is connected to the lowest spatial resolution encoding computation layer in the encoder to extract global features of the field segmentation image; and the decoder receives the global features, progressively upsamples them, and makes skip connections with the encoding computation layers to generate multiple decoding computation layers corresponding to the spatial resolutions of the encoding computation layers, thereby generating the label segmentation image of the global features.
5. The aerial crop identification system as claimed in claim 4, wherein at least one of the bridging layer and the decoder further includes another residual block for transmitting data across the bridging layer or the decoding computation layers.
6. The aerial crop identification system as described in claim 1, wherein the deep learning module further includes a model training unit, which inputs the field segmentation image of the known crop type in the target area for training, and the field vegetation index information corresponding to the field segmentation image into the deep learning module to train the deep learning model to identify the crop type in the field segmentation image and establish the label segmentation image of the known crop type, and generate the corresponding accuracy or loss function. The deep learning module adjusts the parameters in the deep learning model to improve the accuracy or reduce the loss function.
7. The aerial crop identification system as described in claim 1, wherein the vegetation index information includes: Green Red Vegetation Index (GRVI), Visible Atmospherically Resistant Index (VARI), Excess Green Index (ExG), Excess Red Index (ExR), Triangular Greenness Index (TGI), or Color Index of Vegetation Extractable (CIVE).
8. The aerial crop identification system as claimed in claim 1, wherein the aerial vehicle comprises: a drone or a satellite.
9. The aerial crop identification system as claimed in claim 1, wherein the deep learning module establishes a species tag for the crop type in the visible light image, and the tag segmentation image of the species tag corresponds to the distribution range of the crop type in the visible light image.
10. A drone crop identification system, comprising: a receiving module for receiving a visible light image of a target area generated by a drone; a data processing module for generating a Green-Red Vegetation Index (Green-Red Vegetation Index) corresponding to the visible light image; and a deep learning module for receiving the Green-Red Vegetation Index and generating a deep learning model based on the RES U-NET architecture to process the Green-Red Vegetation Index, identify crop species in the visible light image, and establish a labeled segmentation image of the crop species; wherein, The planting area of this crop type includes a field boundary in the target area, which is defined based on the field information of this crop type in the target area from the Geographic Information System (GIS) data; wherein, the receiving module, the data processing module, and the deep learning module are connected by a signal line.
11. A method for aerial crop identification, comprising: providing an aerial photography vehicle to generate a visible light image of a target area, the target area including at least one field segmentation image of a known crop type for training; providing a data processing module to generate at least one field vegetation index information of the at least one field segmentation image based on the visible light image and a field boundary in the target area, wherein the field boundary is defined based on field information of the known crop type in the target area from Geographic Information System (GIS) data; and providing a deep learning module to receive the field vegetation index information, generate a deep learning model based on the RES U-NET architecture, for processing the field vegetation index information, identifying the crop type in the field segmentation image, and establishing a labeled segmentation image of the crop type, wherein... The deep learning model includes an encoder, a bridging layer, a decoder, and at least one residual block, which is used to pass data between multiple encoding computation layers of the encoder, the bridging layer, or the multiple decoding computation layers of the decoder.
12. The aerial crop identification method as described in claim 11, wherein the encoding computation layer, the bridging layer, and the decoding computation layer include multiple batch normalization layers, multiple rectified linear unit layers, and multiple convolutional layers, wherein the aerial crop identification method further provides a gradient judgment module, wherein when the gradient judgment module determines that gradient vanishing occurs in the batch normalization layers, the rectified linear unit layers, or the convolutional layers, or gradient fluctuations occur between adjacent batch normalization layers, the residual block transmits data across the batch normalization layers, the rectified linear unit layers, or the convolutional layers.