Crop organ phenotype analysis method for unmanned aerial vehicle RGB image super-resolution reconstruction based on semantic perception
Through the super-resolution reconstruction method based on semantic perception, the problem of difficulty in taking into account high precision and high efficiency when a drone collects crop organ phenotype images at different flight altitudes is solved, and a more efficient and high-quality crop organ phenotype analysis is achieved.
Patent Information
- Application Number
- CN202510300910.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
The phenotypic images of crop organs collected by drones at different flight altitudes are difficult to take into account both high precision and high efficiency. Traditional super-resolution reconstruction algorithms ignore semantic guidance, resulting in reduced reconstruction efficiency and waste of computing resources.
The super-resolution reconstruction method based on semantic perception is adopted to construct semantic boot data sets and degradation data sets, train HRNet segmentation networks and super-segment reconstruction networks, use degradation models to simulate the degradation process of drone images, and improve the reconstruction quality through coarse-refined network architecture and semantic information guidance.
It significantly improves the spatial resolution of drone images, improves the accuracy and efficiency of phenotype extraction of crop organs, reduces computing resource consumption, and achieves higher reconstruction authenticity and accuracy.
Smart Images

Figure CN120147133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for phenotypic analysis of crop organs, and particularly to a method for phenotypic analysis of crop organs based on semantic perception and super-resolution reconstruction of UAV RGB images. Background Art
[0002] Crop organ phenotype refers to the morphological characteristics of crop organs during growth, such as the size and shape of leaves, the number and size of flowers, etc. These phenotypic information can reflect the genetic characteristics and growth status of crops, and are of great significance for breeding and agricultural production. The application of unmanned aerial vehicles (UAVs) in crop organ phenotypic analysis is becoming increasingly widespread. The RGB cameras carried by them can quickly and efficiently collect crop organ phenotypic data in large-scale areas, significantly improving the efficiency. The images taken by UAVs at low flight altitudes have higher spatial resolution, but lower temporal resolution and stronger wind disturbance to the plant canopy. The images taken at high flight altitudes have higher temporal resolution, but lower spatial resolution, reducing the accuracy of organ-level phenotype extraction. Therefore, it is difficult to simultaneously achieve high precision and high efficiency when using UAVs equipped with RGB cameras to extract organ-level phenotypic traits.
[0003] Image super-resolution reconstruction refers to the process of restoring a high-resolution image from a low-resolution image, aiming to improve the perceptual quality of the image and enhance downstream computer vision tasks (such as classification, segmentation, and detection). This technology can improve the accuracy of phenotypic analysis by enhancing the spatial resolution of UAV RGB images. There are mainly three traditional super-resolution reconstruction algorithms. One is the interpolation method that uses local pixel points around the pixel point to be reconstructed to predict the pixel value, such as bicubic interpolation and bilinear interpolation; the second is the reconstruction method that uses the prior knowledge inside the image, such as iterative back-projection method and projection onto convex sets method; the third is to model the mapping relationship between low-resolution and high-resolution paired images based on methods such as sparse coding to achieve reconstruction.
[0004] The image super-resolution reconstruction technology based on deep learning has attracted much attention in crop phenotyping analysis. This technology requires the construction of low-resolution to high-resolution paired images. There are two traditional construction strategies. One is to adjust the camera focal length to photograph the same object or use dual cameras with different resolutions to photograph simultaneously. However, these strategies require complex post-processing of image registration. Moreover, it is very difficult to construct a dataset of paired low-resolution and high-resolution images in phenotyping research because factors such as wind, light, and time in the field environment may cause significant differences in images taken at different time points. There are scale changes in the drone images of the same ground object at different flight altitudes, which will lead to less than ideal super-resolution reconstruction results. The other construction strategy is to simulate degradation, such as using algorithms like bicubic interpolation to generate simulated low-resolution images from high-resolution images. However, there is a domain gap between the simulated low-resolution images and the low-resolution images collected by drones. The drone images undergo a series of degradation operations (such as compression) by the image signal processor, resulting in poor performance of existing deep learning super-resolution methods based on simulated low-resolution images.
[0005] Current research usually ignores the role of semantic guidance in improving the reconstruction speed and quality. The importance of different regions in drone images for crop organ phenotyping analysis varies. For example, the region belonging to the wheat ear is important for crop organ phenotyping analysis, while the soil region is the opposite. Most super-resolution reconstruction methods use reconstruction models with the same parameter scale for all regions, resulting in a decrease in reconstruction efficiency and waste of computing resources. Therefore, methods such as ClassSR and ARM classify image regions into "difficult", "easy", or other categories according to texture complexity and use models with different parameter scales for reconstruction to improve speed. However, texture complexity cannot fully reflect the significance of different regions for crop organ phenotyping research. For example, the texture complexity of the weed region is similar to that of wheat leaves. Identifying the importance of different regions based on semantic categories for phenotyping research provides a potential solution. In addition to using semantic guidance to improve the SR speed, it is also important to incorporate semantic information to improve the reconstruction quality. Semantic information, as prior knowledge, has been proven to be able to effectively restore the true texture consistent with the high-resolution image. Existing research mainly focuses on the convolutional neural network architecture or uses semantic loss as an additional objective function to incorporate semantic priors. There are few methods to incorporate semantic information into the Transformer-based architecture to improve the reconstruction quality. Summary of the Invention
[0006] Object of the Invention: The object of the present invention is to provide a method for crop organ phenotyping analysis of super-resolution reconstruction of drone RGB images based on semantic perception to improve the spatial resolution of images collected from high flight altitudes, thereby improving the accuracy of crop organ phenotyping extraction.
[0007] Technical solution: A method for crop organ phenotype analysis based on semantic-aware super-resolution reconstruction of UAV RGB images according to the present invention includes the following steps:
[0008] (1) Use a UAV equipped with an RGB camera to collect images at different flight altitudes, and construct a semantic-guided dataset and a degradation dataset;
[0009] (2) Construct a degradation model, which consists of a blur kernel, noise, and interpolation operations;
[0010] (3) Train an HRNet segmentation network based on the semantic-guided dataset, calculate the S e score and extract semantic features, and train a super-resolution reconstruction network based on the degradation dataset;
[0011] (4) Reconstruct the low-resolution RGB image collected by the UAV based on a semantic-aware super-resolution algorithm to generate a super-resolution image;
[0012] (5) Use a deep learning algorithm to extract organ phenotypes in the super-resolution reconstructed image.
[0013] Preferably, the specific steps of constructing the degradation model in step 2 include: explicitly learning the blur kernel from the low-resolution UAV image using KernelGAN, and expanding the degradation space using a typical blur kernel; adding Gaussian noise, Poisson noise, and JPEG noise to simulate the degradation of the image during transmission or storage; using bicubic interpolation and bilinear interpolation to reduce the size of the high-resolution image and lower the image resolution.
[0014] Preferably, the formula of the degradation model in step 2 is as follows:
[0015]
[0016] k = F G (I LR );
[0017] where I scaled represents the scaled high-resolution image, I LR represents the generated low-resolution image, ↓ s represents the downsampling operation with a magnification of s, k and n respectively represent the blur kernel and additive noise, x represents the image patch extracted from the LR image I LR , G and D respectively represent the generator and discriminator, R represents the regularization term, and F G represents the convolution operation.
[0018] Preferably, the specific steps of training the HRNet segmentation network based on the semantic guidance dataset in step 3 include: the semantic segmentation annotation tool marks the data, increases the scale of the dataset by rotation, flipping, and scaling, and trains the HRNet segmentation network based on the dataset.
[0019] Preferably, the training of the super-resolution reconstruction network based on the degradation dataset in step 3 specifically includes a PSNR-guided stage and a perception-guided stage. The PSNR-guided stage uses an L1 loss function; the perception-guided stage initializes the generator with the weights of the PSNR-guided model and simultaneously adopts L1 loss, perceptual loss, and GAN loss functions, with the weights set to 1, 0.1, and 0.1 respectively, and uses the trained HRNet segmentation network to extract semantic features and calculate the S e score.
[0020] Preferably, the super-resolution reconstruction of the RGB image collected by the drone in step 4 includes a generator and a discriminator. The generator adopts a coarse-refined architecture and consists of a coarseNet for coarse reconstruction, an SE gating module for evaluating whether the coarsely reconstructed image needs to be optimized, and a refinedNet for optimizing the coarsely reconstructed image. The discriminator consists of an encoder and a decoder.
[0021] Preferably, the encoder maps the super-resolution image to a high-dimensional feature space through convolutional operations and the LeakyReLU activation function, and performs three downsampling operations. The downsampling operations include a convolutional layer, spectral regularization, and the LeakyReLU activation function; the decoder performs three upsampling operations. The upsampling operations include bilinear interpolation, a convolutional layer, the LeakyReLU activation function, and spectral regularization. The size of the sampled feature map is used to splice the feature map in the downsampling stage of the encoder with the corresponding feature map in the upsampling stage through skip connections, and the authenticity value of each pixel in the super-resolution image is calculated through a convolutional layer.
[0022] Preferably, the specific formula for the coarse reconstruction is as follows:
[0023] I coarse = f up (Conv(F shallow + F c_deep ));
[0024] where f up represents the Pixel Shuffle reconstruction module, Conv represents the convolutional operation, F shallow represents the shallow low-frequency feature, F c_deep represents the deep high-frequency feature, and I coarse represents the coarsely super-resolved image.
[0025] Preferably, the specific formula for determining whether the coarsely reconstructed image needs to be optimized is as follows:
[0026]
[0027] Among them, p represents the number of pixels with a gray value greater than zero in the semantic segmentation mask image, w and h respectively represent the width and height of the mask image, S e ∈[0,1], and t represents the threshold of the proportion of important regions in the UAV image.
[0028] Preferably, the specific formula for optimizing the coarsely reconstructed image is as follows:
[0029] F concat = Concat(LReLU(LN(Conv 1 (F seg )), LReLU(LN(F refined )));
[0030] F pre_f = LN(Conv 2 (F concat ));
[0031] Q c = Reshape(DConv(Conv(LReLU(F pre_f ))));
[0032]
[0033] F wrap = Conv(LN(F refined + F attention ));
[0034] F r_deep = Conv(Gelu(Conv(LN(F seg ))) × GeLu(Conv(LN(F wrap )))) + F wrap ;
[0035] I refined = f up (Conv(F shallow + F r_deep ));
[0036] Among them, F seg represents the semantic features extracted from semantic segmentation, F refined represents the super-resolution features, LReLU represents the LeakyReLU activation function, Concat, Conv 1 and Conv 2represent concatenation, convolution, and convolution operations respectively, F concat represents the preprocessed features, the Reshape operation is used to change the shape of the tensor, DConv represents the two-dimensional depthwise separable convolution, F attention represents generating the semantic-aware attention map, SoftMax represents the Softmax activation function, Q c 、K c and V c represent the query, key, and value in the cross-attention respectively, d represents the dimension of the query / key, + represents element-wise addition, F r_deep represents generating the feature map as the input for the next HTB-SA module, GeLu represents the Gaussian Error Linear Units activation function, × represents element-wise multiplication, f_up represents the PixelShuffle reconstruction module, I refined represents refining and reconstructing the SR image patch.
[0037] Advantages: Compared with the prior art, the present invention has the following remarkable advantages: 1. Using a degradation model to degrade a high-resolution image to generate a paired low-resolution image, by simulating the degradation operations such as compression and noise experienced by real UAV images during the acquisition process, so as to generate simulation data closer to the low-resolution images acquired by UAVs; 2. Through a multi-scale scaling data augmentation strategy, scaling the high-resolution image to generate an image with approximately invariant spatial resolution, which can simulate the scale change of the same ground object in images at different heights; 3. Implementing the reconstruction of different regions of different importance using models with different parameter scales through a coarse-refined network architecture, and introducing semantic information as prior knowledge to improve the reconstruction fidelity, significantly improving the reconstruction efficiency and authenticity, reducing the consumption of computing resources, and improving the accuracy of crop organ phenotype extraction. Brief Description of the Drawings
[0038] Figure 1 is the schematic flow chart of the present invention;
[0039] Figure 2 is the schematic diagram of the degradation model of the present invention;
[0040] Figure 3 is the schematic diagram of the semantic-aware super-resolution reconstruction algorithm of the present invention;
[0041] Figure 4 is the schematic diagram of the effective improvement of the image spatial resolution by the super-resolution reconstruction algorithm of the present invention, where (a) represents the NIQE, FID, and HyperIQA index diagrams of the 10 - 40m super-resolution image; (b) represents that the algorithm proposed by the present invention can effectively restore the wheat organ details in the 10 - 40m image;
[0042] Figure 5Schematic diagram for effectively improving the accuracy of crop organ phenotype analysis in the present invention. Detailed implementation manners
[0043] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0044] The research area is located at the Baima Experimental Station of Nanjing Agricultural University, China (119°18′71″ E, 31°62′00″ N). During 2022 to 2023, a total of 1116 plots were planted, including 558 wheat varieties, with two replicates for each variety, and the nitrogen fertilizer application rate was 240 kg / ha. The area of each plot is 1.5 m × 1.5 m, with six rows of wheat planted, the row spacing is 0.25 m, and the distance between plots is 0.5 m. A DJI Matric 300 RTK drone equipped with a Zenmuse P1 RGB camera (resolution 8192×5436 pixels) was used, in waypoint flight mode, the camera angle was 90°, and the acquisition height range was from 3 m to 40 m.
[0045] (1) From April 19 to May 29, 2023, a drone equipped with an RGB camera was used to collect RGB images once a week in waypoint flight mode. A semantic-guided dataset and a degradation dataset were constructed based on the drone images, and the two datasets were used for the development and verification of super-resolution reconstruction algorithms.
[0046] The semantic-guided dataset is used to extract semantic features and screen regions important for organ phenotype research in the images to calculate the S e score and extract semantic prior knowledge, including the following steps:
[0047] 1. Crop the high-resolution drone images collected at a flight height of 3 m into image patches of size 256×256×3.
[0048] 2. Through an annotation tool based on the Segment Anything large model, pixels are labeled into three categories: wheat spikes, canopy regions, and background (such as soil and weeds).
[0049] 3. Finally, the dataset scale is increased through common data augmentation strategies, such as rotation, flipping, and scaling.
[0050] The degradation dataset is used to train the model to enhance the spatial resolution of drone images, including the following steps:
[0051] 1. Crop the high-resolution drone images collected at a flight height of 3 m into image patches of size 1024×1024×3.
[0052] 2. The proposed multi-scale scaling strategy is used to scale the high-resolution image patches to solve the problem that the same ground object in drone images at different flight heights has scale changes, resulting in poor reconstruction effects.
[0053] 3. Each image patch is scaled to r ∈ {0.75, 0.5, 0.39, 0.33} times the original size, and the output sizes are 768×768×3, 512×512×3, 400×400×3, and 341×341×3;
[0054] 4. During the training process, it is dynamically cropped from the HR image patch to a size of 256×256×3, and then the proposed degradation model is used to generate paired LR samples with a size of 64×64×3 and a downsampling ratio of 4.
[0055] The multi-scale scaling strategy enables the super-resolution network to learn multi-scale features of the same ground object, thus achieving higher reconstruction quality. Specifically, the high-resolution image is scaled to a smaller size using predefined parameters, and the formula is as follows:
[0056] I scaled = I HR × r, 0 < r < 1; (1)
[0057] where I scaled represents the scaled high-resolution image patch, I HR represents the high-resolution image patch after cropping, and r represents the scaling parameter, whose value is restricted between 0 and 1.
[0058] (2) Construct a degradation model, which consists of a blur kernel, noise, and interpolation operations. This model assumes that the low-resolution image is obtained through the following process:
[0059]
[0060] where I LR represents the generated low-resolution image, ↓ s represents the downsampling operation with a magnification factor of s, k represents the blur kernel, and n represents the additive noise;
[0061] This blur kernel includes a learned kernel and a typical kernel. The learned kernel aims to explicitly simulate the degradation process unique to UAV images. The blur kernel is explicitly learned from low-resolution UAV images captured at a high flight altitude through KernelGAN in an unsupervised manner from low-resolution UAV images at a flight altitude of 15m. The formula is as follows:
[0062]
[0063] where x represents the LR image I LRThe image patches extracted, where G and D represent the generator and discriminator respectively, and R is the regularization term. Since the generator of KernelGAN is a linear network without non-linear activation, the blur kernel k required for degradation can be explicitly extracted through a convolution with a stride of 1, and the formula is as follows:
[0064] k = F G (I LR ); (4)
[0065] where F G represents the convolution operation, and typical kernels include isotropic and anisotropic Gaussian kernels and 2D sinc filters.
[0066] At the same time, typical blur kernels are used to expand the degradation space. Typical blur kernels include isotropic and anisotropic Gaussian kernels and 2D sinc filters. Typical noises are added to simulate the degradation of images during transmission or storage. Typical noises include Gaussian noise, Poisson noise, and JPEG noise. During the degradation process, bicubic interpolation and bilinear interpolation sampling methods are used to reduce the size of high-resolution images.
[0067] (3) Train the HRNet segmentation network based on the semantic-guided dataset to calculate the S e score and extract semantic features. HRNet is implemented based on Python 3.9 and Pytorch 1.10.2. The training process lasts for 200 Epochs, the initial learning rate is set to 0.01, and the batch size is 8. The SGD algorithm is used as the optimizer, with a momentum of 0.937 and a weight decay of 0.0005. All training and testing are completed on a Windows server equipped with an NVIDIA GeForce RTX 3090 GPU (24GB video memory) and an Intel(R) Xeon(R) W-2245 CPU (128GB memory);
[0068] Based on the degradation dataset, a two-stage strategy is used to train other parts of the super-resolution reconstruction network, including the PSNR-guided stage and the perception-guided stage. During this period, the S e score is set to 1 to ensure that coarseNet and refineNet are trained using all samples.
[0069] In both stages, the size of the training image patches is 64×64×3, the reconstruction magnification is 4, and an image of 256×256×3 is generated. During the training process, the constructed degradation model is used to generate low-resolution-high-resolution paired images. In the PSNR-guided stage, the generator is trained for 100,000 iterations with a learning rate of 2×10 -4 and a Batch size of 24. In the perception-guided stage, the generator network parameters are initialized with the weights of the PSNR-guided stage, and the learning rate is 1×10-4 With a batch size of 8, train for 100,000 iterations.
[0070] Meanwhile, introduce a discriminator and train it using the same strategy as the generator. The parameters of the generator and the discriminator are updated alternately. The optimizer is Adam, with β 1 = 0.9 and β 2 = 0.99.
[0071] In the perception-guided stage, coarseNet and refineNet are optimized by one optimizer and one discriminator respectively, and the parameters are not shared. Introduce EMA (exponential moving average) to achieve more stable training and better performance. The super-resolution network is implemented based on Python 3.9 and Pytorch 1.10.2. The network is trained and evaluated on an Ubuntu server equipped with an NVIDIA GeForce A100 GPU (80GB video memory) and an Intel(R) Xeon(R) Gold 6348 CPU (200GB memory).
[0072] (4) Crop the low-resolution UAV RGB image into tiles of size 256×256×3, record the coordinates of each tile in the original image, enhance the resolution through the trained super-resolution reconstruction model, and splice the reconstructed images according to the recorded coordinates;
[0073] Super-resolution reconstruction mainly consists of two parts: a generator and a discriminator:
[0074] (a) Generator. The generator adopts a coarse-refined architecture, aiming to accelerate reconstruction and improve the reconstruction quality through semantic prior guidance. It includes a coarse sub-network (coarseNet), an SE gating module, and a refined sub-network (refinedNet).
[0075] 1. The Coarse sub-network is used for coarse reconstruction
[0076] First, extract shallow low-frequency features F LR ∈R H×W×C from the low-resolution image patch I shallow through a convolutional layer, and map it to a high-dimensional feature space;
[0077] Second, extract deep high-frequency features F shallow from F c_deep, The HTB module is the basic building block of the method proposed in this invention. It includes k hybrid attention blocks (HAB) and a skip connection. The HAB aims to activate more pixels at the local scale and channel dimension through the attention mechanism, and is composed of a layer normalization operation (Layer Normalization, LN), a CAB module, a multi-head self-attention module (standard and shifted window multihead self-attention module, (S)W-MSA), a skip connection, and another LN layer;
[0078] Then, reusable super-resolution features F are extracted from the last HTB module of coarseNet reusable , for subsequent feature fusion in refinedNet. The low-frequency information in the shallow features is transmitted to the reconstruction module through the skip connection, which helps the HTB module focus on high-frequency information and stabilize training. A convolutional layer is introduced before upsampling to integrate the inductive bias of the convolutional operation into the Transformer-based network;
[0079] Finally, through upsampling module reconstruction, a coarse super-resolution image I is obtained coarse , as follows:
[0080] I coarse = f up (Conv(F shallow + F c_deep )); (5)
[0081] where f up represents the Pixel Shuffle reconstruction module, and Conv represents the convolutional operation.
[0082] 2. SE Gating Module
[0083] The SE gating module evaluates whether the coarse super-resolution image I needs further optimization based on the S e score. The S coarse score is calculated based on the segmentation mask generated by the pre-trained HRNet, as follows: e
[0084]
[0085] where p represents the number of pixels with a gray value greater than zero in the semantic segmentation mask image, w and h represent the width and height of the mask image respectively. As the flight altitude increases, the spatial resolution of the low-score image decreases, affecting the segmentation accuracy and resulting in inaccurate S e score values. Therefore, the segmentation operation is performed on the coarse super-resolution image I coarse performed based on S e score (S e ∈ [0, 1]) to screen the coarse super-resolution image patches that need to be refined, as follows:
[0086]
[0087] where t is the threshold of the proportion of important regions in the UAV image, and each S e coarse super-resolution image patch I with a score exceeding the threshold t coarse will be further optimized by refinedNet.
[0088] 3. refinedNet is used to optimize the coarse super-resolution image patch I coarse to further improve the reconstruction quality.
[0089] First, extract high-frequency features F reusable from F through n hybrid Transformer blocks with semantic-aware fusion modules (HTB-SA). r_deep Compared with HTB, HTB-SA further introduces a semantic-aware fusion module (SAFM) to integrate semantic information.
[0090] The SAFM module fuses the super-resolution features with the class-related prior knowledge extracted from the semantic segmentation network through the cross-attention mechanism, which specifically includes the following steps:
[0091] The first step is to pre-fuse the semantic features F seg extracted from semantic segmentation with the super-resolution features F refined , as follows:
[0092] F concat = Concat(LReLU(LN(Conv 1 (F seg ))), LReLU(LN(F refined ))); (8)
[0093] F pre_f = LN(Conv 2 (F concat )); (9)
[0094] where LReLU represents the LeakyReLU activation function, and Concat, Conv 1 and Conv 2 represent concatenation, convolution, and convolution operations respectively.
[0095] The second step is to process the pre-processed features F concat, as the query (Q c ) in the cross - attention, is as follows:
[0096] Q c = Reshape(DConv(Conv(LReLU(F pre_f )))); (S4)
[0097] Among them, the Reshape operation is used to change the shape of the tensor, and DConv represents two - dimensional depth - separable convolution.
[0098] The third step is to generate the key K refined and the value V c based on the super - resolution feature F c .
[0099] The fourth step is to fuse the semantic feature and the super - resolution feature based on the cross - attention mechanism to generate the semantic - aware attention map F attention , and the calculation is as follows:
[0100]
[0101] Among them, SoftMax represents the Softmax activation function, Q c , K c and V c represent the query, key, and value in the cross - attention respectively, and d represents the dimension of the query / key.
[0102] The fifth step is to fuse the super - resolution feature and the semantic - aware attention map, as follows:
[0103] F wrap = Conv(LN(F refined + F attention )); (11)
[0104] Among them, + represents element - wise addition.
[0105] The sixth step is to generate the feature map F r_deep as the input of the next HTB - SA module:
[0106] F r_deep = Conv(Gelu(Conv(LN(F seg )))×GeLu(Conv(LN(F wrap ))))+ F wrap (12)
[0107] Among them, GeLu represents the Gaussian Error Linear Units activation function, × represents element - wise multiplication, and the fine - reconstruction SR image patch I is generated through skip connections, convolutional layers, and up - sampling modules.refined , as follows:
[0108] I refined = f up (Conv(F shallow + F r_deep )); (13)
[0109] where f up represents the Pixel Shuffle reconstruction module.
[0110] (b) Discriminator
[0111] The discriminator is used to constrain the generator during training to reconstruct images with better quality. The discriminator consists of an encoder and a decoder, takes the super-resolution image as input, and outputs the authenticity value of each pixel.
[0112] In the encoder, first, the super-resolution image is mapped to a high-dimensional feature space through convolutional operations and the Leaky ReLU activation function, and three downsampling operations are performed, including convolutional layers, spectral normalization, and the Leaky ReLU activation function.
[0113] In the decoder, first, three upsampling operations are performed, including bilinear interpolation, convolutional layers, the Leaky ReLU activation function, and spectral normalization, to upsample the feature map size while reducing the number of channels. The feature maps in the downsampling stage of the encoder are concatenated with the corresponding feature maps in the upsampling stage through skip connections. Finally, the authenticity value of each pixel in the super-resolution image is calculated through a convolutional layer.
[0114] (5) Crop the RGB images collected by waypoint flight or the stitched drone RGB route flight images into overlapping tiles, record the coordinates of the tiles in the original images, and use the proposed super-resolution reconstruction algorithm to enhance the spatial resolution of the low-resolution tiles;
[0115] During this period, dynamically determine whether the refinedNet is needed to further optimize the reconstruction result according to the S e score of the tiles. After all the tiles are reconstructed, stitch the super-resolution orthoimage according to the corresponding coordinates of each super-resolution tile in the low-resolution image.
[0116] Then use various deep learning algorithms to extract organ-level phenotypes from the super-resolution orthoimage, such as using object detection algorithms to detect the number of wheat ears in the wheat breeding plot.
[0117] (6) Use deep learning algorithms to extract organ phenotypes from the super-resolution orthoimage and perform analysis;
[0118] Taking the calculation of the number of wheat ears in a single breeding plot as an example, the process of organ phenotype extraction based on super-resolution reconstructed images is demonstrated, including the following steps: 1. Extract a single breeding plot from the super-resolution image using the MaskRCNN instance segmentation algorithm; 2. Detect the wheat ears in a single breeding plot using the YOLOV8 object detection algorithm to obtain the number of wheat ears.
[0119] Figure 4 It can be seen that the super-resolution reconstruction algorithm designed in the present invention effectively improves the spatial resolution of the image. (a) shows the NIQE, FID, and HyperIQA metrics of the 10-40m super-resolution image, which are 5.547, 193.627, and 0.461 respectively, while the corresponding values of the low-resolution image are 19.628, 245.526, and 0.331. Compared with the low-resolution image, the NIQE and FID metrics of the super-resolution image decreased by 71.37% and 21.53% respectively, and the HyperIQA metric increased by 39.36%. (b) shows that the algorithm proposed in the present invention can effectively restore the details of wheat organs in the 10-40m image.
[0120] Figure 5 It shows that the present invention can effectively improve the accuracy of crop organ phenotype analysis, taking the detection and calculation of wheat ears as an example. At the same flight altitude, the detection accuracy of the super-resolution image is better than that of the low-resolution image. The detection accuracy of the super-resolution image gradually decreases with the increase of the flight altitude, but is always higher than that of the low-resolution image.
Claims
1. A crop organ phenotyping method based on semantic-aware UAV RGB image super-resolution reconstruction, characterized in that: The following steps are involved: (1) Use a drone equipped with an RGB camera to collect images at different flight altitudes and construct a semantic guidance dataset and a degradation dataset; (2) constructing a degradation model, wherein the degradation model is composed of a blur kernel, noise, and interpolation operations; (3) Train the HRNet segmentation network based on the semantic guidance dataset and calculate S e Score and extract semantic features, and train a super-resolution reconstruction network based on the degraded data set; (4) Reconstruct the low-resolution RGB images collected by the drone based on the semantic perception super-resolution algorithm to generate super-resolution images; (5) Use deep learning algorithms to extract organ phenotypes in super-resolution reconstructed images.
2. The crop organ phenotyping method according to claim 1, characterized in that: The specific steps of constructing the degradation model described in step 2 include: using KernelGAN to explicitly learn the blur kernel from the low-resolution UAV image, and using the typical blur kernel to expand the degradation space; adding Gaussian noise, Poisson noise and JPEG noise to simulate the degradation of the image during transmission or storage; using bicubic interpolation and bilinear interpolation to reduce the size of the high-resolution image and reduce the image resolution.
3. The crop organ phenotyping method according to claim 1, characterized in that: The degradation model formula described in step 2 is as follows: k=F G (I LR ); Among them, I scaled represents the scaled high-resolution image, I LR Represents the generated low-resolution image,↓ s represents a downsampling operation with a magnification of s, k and n represent blur kernel and additive noise respectively, and x represents the LR image I LR The image blocks extracted from , G and D represent the generator and discriminator respectively, R represents the regularization term, F G Represents a convolution operation.
4. The crop organ phenotyping method according to claim 1, characterized in that: Step 3 trains the HRNet segmentation network based on the semantically guided dataset, and the specific steps include: marking the data with a semantic segmentation annotation tool, increasing the size of the dataset by rotating, flipping and scaling; and training the HRNet segmentation network based on the dataset.
5. The crop organ phenotyping method according to claim 1, characterized in that: Step 3 trains the super-resolution reconstruction network based on the degraded data set, specifically including a PSNR-guided stage and a perception-guided stage. The PSNR-guided stage uses the L1 loss function; the perception-guided stage uses the weight initialization generator of the PSNR-guided model, and uses the L1 loss, perceptual loss and GAN loss functions at the same time, with the weights set to 1, 0.1 and 0.1 respectively, and uses the trained HRNet segmentation network to extract semantic features and calculate S e Fraction.
6. The crop organ phenotyping method according to claim 1, characterized in that: In step 4, super-resolution reconstruction of the RGB image collected by the drone is performed, including a generator and a discriminator. The generator adopts a coarse-refined architecture, which is composed of a coarseNet for coarse reconstruction, an SE gating module for evaluating whether the coarse reconstructed image needs to be optimized, and a refinedNet for optimizing the coarse reconstructed image. The discriminator is composed of an encoder and a decoder.
7. The crop organ phenotyping method according to claim 6, characterized in that: The encoder maps the super-resolution image to a high-dimensional feature space through convolution operations and Leaky ReLU activation functions, and performs three downsampling operations, wherein the downsampling operations include convolution layers, spectral regularization, and Leaky ReLU activation functions; the decoder performs three upsampling operations, wherein the upsampling operations include bilinear interpolation, convolution layers, Leaky ReLU activation functions, and spectral regularization, samples the feature map size, splices the feature map of the encoder downsampling stage with the corresponding feature map of the upsampling stage through jump connections, and calculates the authenticity value of each pixel in the super-resolution image through convolution layers.
8. The crop organ phenotyping method according to claim 6, characterized in that: The specific formula for the coarse reconstruction is as follows: I coarse =f up (Conv(F shallow +F c_deep )); Among them, f up represents the Pixel Shuffle reconstruction module, Conv represents the convolution operation, and F shallow represents shallow low-frequency features, F c_deep represents deep high-frequency features, I coarse Represents a coarse super-resolved image.
9. The crop organ phenotyping method according to claim 6, characterized in that: The specific formula for judging whether the coarse reconstructed image needs to be optimized is as follows: Where p represents the number of pixels with grayscale values greater than zero in the semantic segmentation mask image, w and h represent the width and height of the mask image, respectively. e ∈[0,1], t represents the threshold of the proportion of important areas in the UAV image.
10. The crop organ phenotyping method according to claim 6, characterized in that: The specific formula for optimizing the coarse reconstruction image is as follows: F concat =Concat(LReLU(LN(Conv1(F seg ))),LReLU(LN(F refined ))); F pre_f =LN(Conv2(F concat )); Q c =Reshape(DConv(Conv(LReLU(F pre_f ))))); F wrap =Conv(LN(F refined +F attention )); F r_deep =Conv(GeLu(Conv(LN(F seg )))×GeLu(Conv(LN(F wrap ))))+F wrap ; I refined =f up (Conv(F shallow +F r_deep )); Among them, F seg represents the semantic features extracted from semantic segmentation, F refined represents super-resolution features, LReLU represents LeakyReLU activation function, Concat, Conv1 and Conv2 represent concatenation, convolution and convolution operations respectively, concat Represents preprocessing features, Reshape operation is used to change the shape of the tensor, DConv represents two-dimensional depth separable convolution, F attention represents the generation of semantically aware attention maps, SoftMax represents the Softmax activation function, Q c , K c and V c denotes the query, key, and value in the cross attention, d denotes the dimension of the query / key, + denotes element-wise addition, and F r_deep Indicates the generation of feature maps as the input of the next HTB-SA module, GeLu represents the Gaussian Error Linear Units activation function, × represents the element-by-element product, f_up represents the Pixel Shuffle reconstruction module, I refined Represents a finely reconstructed SR image patch.