Unsupervised field adaptive unmanned aerial vehicle wheat disease accurate detection method
By constructing source domain and target domain data sets, using unsupervised adaptive teacher-student networks and adaptive semantic segmentation networks to generate pseudo-labels, the problems of large data demand and high marking in the existing technology are solved, and efficient and accurate detection of wheat disease detection is achieved.
Patent Information
- Application Number
- CN202510172481.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing wheat disease detection methods require a large number of image data sets for different wheat diseases to train models, resulting in large data demand, high marking volume, and high training and detection cost, making it difficult to expand and apply to other wheat diseases.
The unsupervised field adaptation method is adopted to construct the source domain and target domain data sets of wheat diseases. Through the unsupervised adaptive teacher-student network and adaptive semantic segmentation network, pseudo-labels of the target domain are generated, and the global and local information is captured in combination with the dual-branch structure to conduct wheat disease lesion area detection.
It reduces the pressure of data labeling, improves the accuracy and efficiency of wheat disease detection, and realizes accurate detection with low labeling dependence.
Smart Images

Figure CN120259912A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wheat disease detection, and relates to, but is not limited to, an unsupervised domain adaptation method for accurate detection of wheat diseases by drones. Background Art
[0002] Wheat stripe rust, wheat yellow dwarf disease, and wheat scab are the main diseases that seriously harm wheat and have a wide impact on its production. Traditional manual visual detection methods for wheat diseases have many deficiencies. Currently, detection methods for visually identifying wheat diseases have been reported. For example, a multi-scale detection method for wheat yellow dwarf disease based on drone multispectral images with the patent application number 202410649459.8. This method includes constructing a dataset to be detected composed of complete wheat yellow dwarf images, and using a trained dual-branch multi-scale model constructed by a dual-branch scale encoder and a lightweight decoder to detect and identify the wheat yellow dwarf disease lesion area in the wheat yellow dwarf images. Although this method can accurately segment wheat yellow dwarf disease, the dataset used in this method must be constructed using wheat yellow dwarf images and the model is trained with wheat yellow dwarf image data. Then, when extended to other wheat diseases, it is necessary to construct training sets with images of each wheat disease to train the model. Obviously, obtaining images of various wheat diseases is a huge task. Therefore, there is a need to develop a precise wheat disease detection method with a small data requirement, a small labeling amount, and a low training and detection cost. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides an unsupervised domain adaptation method for accurate detection of wheat diseases by drones. The method of the present invention can reduce the data annotation pressure and achieve accurate detection of wheat diseases. By designing a field-scale segmentation model with low annotation dependence through technologies such as domain adaptation, knowledge distillation, and dual-branch, the model can generate pseudo-labels of target domain data by learning the feature characteristics of the source domain data with labels, and then capture the global and local information of the target domain images through a dual-branch structure. At the same time, feature fusion and multi-scale fusion are considered, which helps to improve the accuracy and efficiency of the model in the detection task of wheat disease lesion areas.
[0004] The technical solution of the embodiment of the present invention is implemented as follows: An unsupervised domain adaptation method for accurate detection of wheat diseases by drones, and the method is as follows: Construct a source domain dataset and a target domain dataset for wheat diseases. The source domain dataset and the target domain dataset are constructed by drones collecting multiple images during the complete infection cycle of wheat diseases; the source domain dataset of wheat diseases includes a dataset with wheat disease category labels, and the target domain dataset of wheat diseases includes a dataset without wheat disease category labels; the wheat diseases in the source domain dataset and the target domain dataset are of different types of diseases; Label the background area, healthy wheat area, and wheat disease lesion area of the images in the source domain dataset to obtain the source domain label image dataset; Construct the dataset to be detected from the source domain label image dataset and the target domain dataset; Construct a low-annotation-dependent wheat disease field-scale segmentation model based on an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network; wherein, the unsupervised adaptive teacher-student network includes a cross-domain mixing module and a teacher-student network for generating pseudo-labels of the target domain; the adaptive semantic segmentation network includes a decoder with a Segformer-ResNet concurrent encoder and an optimized feature fusion module. The Segformer-ResNet concurrent encoder reconciles the dimensional differences between the feature maps of ResNet and the feature embeddings of Segformer during the encoding stage through a feature alignment module, thereby achieving effective fusion of multi-dimensional features; the decoder of the optimized feature fusion module introduces depthwise separable convolutions to enhance the expression of context information and fuse multi-level features to improve the segmentation performance; Input the dataset to be detected into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each image of the target domain dataset.
[0005] Preferably, the source domain dataset and the target domain dataset are constructed from multiple images within the complete infection cycle of wheat diseases, specifically including: the source domain dataset and the target domain dataset are constructed from multiple groups of spectral images within the complete infection cycle of wheat diseases collected by drones.
[0006] Preferably, the source domain dataset and the target domain dataset are constructed from multiple groups of spectral images within the complete infection cycle of wheat diseases collected by drones, including: using a spectral drone to image the wheat pest and disease images in a preset experimental field to obtain multiple groups of multispectral images; the preset experimental field contains multiple different wheat varieties and covers the complete disease-susceptible cycle of wheat diseases; each group of the multispectral images includes an RGB image and multispectral band images of five bands.
[0007] Preferably, the cross-domain mixing module includes a sample mixing module and a data augmentation module. The sample mixing module mixes the images and corresponding labels in the source domain dataset with the spectra in the target domain dataset to generate mixed samples and pseudo-labels. The mixed samples, pseudo-labels, and the data in the source domain dataset are simultaneously data-augmented by the data augmentation module and input into the next network; The teacher-student network includes a teacher network and a student network. The student network receives the mixed sample data and the data in the source domain dataset, outputs the exponential moving average parameters to the teacher network, and outputs features to the subsequent network. The teacher network receives the mixed sample data from the cross-domain mixing module and the exponential moving average parameters of the student network, and outputs features to the subsequent network.
[0008] Preferably, the SegFormer branch includes four SegFormer blocks connected in sequence. The output of the previous SegFormer block serves as the input of the subsequent SegFormer block at the same time, and is used to generate feature maps of different scales. The ResNet branch part includes four identical convolutional blocks connected in sequence, and batch normalization and ReLU activation functions are added between adjacent convolutional blocks. The feature alignment module is used to connect the feature maps generated by the third and fourth blocks of the two branches along the channel dimension, and perform feature fusion on the channel features and the feature maps to obtain a fused feature map. The optimized feature fusion module is obtained by replacing the original multi-layer perceptron in the adaptive semantic segmentation network with a 1×1 convolution and bilinear upsampling, and then introducing a dilated convolution group.
[0009] Preferably, inputting the dataset to be detected into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each image in the target domain dataset includes: Input each image in the dataset to be detected into the cross-domain mixing module and the data augmentation module to generate mixed samples and pseudo-labels. Process the data in the source domain dataset and the mixed sample data through the teacher-student network, and then input them into the SegFormer branch and the ResNet branch to extract wheat disease features. Fuse the third and fourth feature maps output by each channel of the SegFormer branch and the ResNet branch through the feature alignment module to obtain a fused feature map. First, align the size and number of channels of the output of the encoder through the optimized feature fusion module, and then feed the features into the dilated convolution group for context-aware feature fusion. Connect and input the output of the optimized feature fusion module into a 1×1 convolution to obtain the final semantic segmentation output, and finally segment the target domain wheat disease lesion areas in the target domain dataset.
[0010] Preferably, mark the background areas, healthy wheat areas, and wheat disease lesion areas in the images of the source domain dataset. After marking, use the random enhancement method to augment the marked source domain dataset to obtain a source domain label image dataset with a preset number.
[0011] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include: The present invention adopts unsupervised domain adaptation technology. First, a source domain dataset and a target domain dataset of wheat diseases are constructed. Then, only the background area, healthy wheat area, and wheat disease lesion area in the images of the source domain dataset need to be labeled, which can reduce the data annotation pressure. Then, a dataset to be detected is constructed from the source domain label image dataset and the target domain dataset. Finally, the dataset to be detected is input into the wheat disease field-scale segmentation model to detect the wheat disease lesion area in each spectral image of the target domain dataset. The low-annotation-dependent wheat disease field-scale segmentation model used is constructed based on an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network through technologies such as domain adaptation, knowledge distillation, and dual branches, enabling the model to generate pseudo-labels of the target domain data by learning the characteristics of the labeled source domain data, and then capturing the global and local information of the target domain images through the dual-branch structure, while considering feature fusion and multi-scale fusion, which helps to improve the accuracy and efficiency of the model in the task of detecting wheat disease lesion areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where: Figure 1 It is a schematic flowchart of an unsupervised domain adaptation-based precise detection method for wheat diseases by an unmanned aerial vehicle provided in an embodiment of the present invention; Figure 2 It is an example diagram of a multi-spectral image collected by a spectral unmanned aerial vehicle provided in an embodiment of the present invention; Figure 3 It is a schematic diagram of the overall architecture of a low-annotation-dependent field-scale segmentation model provided in an embodiment of the present invention; Figure 4 It is a schematic diagram of an unsupervised adaptive teacher-student network structure provided in an embodiment of the present invention; Figure 5 It is a schematic diagram of an adaptive semantic segmentation network structure provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but not to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0015] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.
[0016] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0017] The embodiments of the present invention provide an unsupervised domain adaptation method for precise detection of wheat diseases by drones, which is applied to an electronic device. The electronic device includes, but is not limited to, a mobile phone, a laptop computer, a tablet computer, a handheld Internet device, a multimedia device, a streaming media device, a mobile Internet device, a wearable device, or other types of electronic devices. The functions implemented by this method can be realized by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium. The processor can be used to process the precise detection process of wheat diseases based on drone images and domain adaptation technology, and the memory can be used to store the data required and generated during the precise detection process of wheat diseases based on drone images and domain adaptation technology.
[0018] Figure 1The following is a schematic flowchart of an unsupervised domain adaptation method for precise detection of wheat diseases provided by an embodiment of the present invention. As Figure 1 shown, the method at least includes the following steps: Step S1, constructing a source domain dataset and a target domain dataset for wheat diseases. The source domain dataset and the target domain dataset are constructed by multiple images within the complete infection cycle of wheat diseases collected by an unmanned aerial vehicle (UAV). The source domain dataset of wheat diseases includes a dataset with wheat disease category labels, and the target domain dataset of wheat diseases includes a dataset without wheat disease category labels; Here, the multiple images within the complete infection cycle of wheat diseases can be multispectral images, specifically multispectral images obtained by using a spectral UAV to image different wheat varieties under complex backgrounds. A multispectral image refers to a multispectral band image including multiple bands.
[0019] Specifically, use a spectral UAV to image wheat pest and disease images in a preset experimental field to obtain multiple sets of spectral images. The preset experimental field contains multiple different wheat varieties and covers the complete disease-susceptible cycle of wheat diseases. Each set of the multispectral images includes 1 RGB image and multispectral band images of 5 bands, namely Blue, Green, Red, RE, and NIR. As Figure 2 shown, the multispectral images of wheat rust, wheat scab, and wheat yellow dwarf disease are listed in the figure. The source domain dataset of wheat diseases can be constructed by selecting the multispectral images of wheat scab, wheat yellow dwarf disease, or wheat rust. The types of wheat diseases in the target domain dataset may be unknown and do not necessarily match the types of wheat diseases in the source domain dataset.
[0020] Step S2, marking the background area, healthy wheat area, and wheat disease lesion area in the images of the source domain dataset to obtain a source domain label image dataset; Step S3, constructing a dataset to be detected from the source domain label image dataset and the target domain dataset; Step S4, constructing a low-labeling-dependent wheat disease field-scale segmentation model based on an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network; Among them, the unsupervised adaptive teacher-student network (DATS) includes a cross-domain mixing module and a teacher-student network for generating pseudo-labels of the target domain; the adaptive semantic segmentation network (ASSFormer) includes a decoder with a Segformer-ResNet concurrent encoder and an optimized feature fusion module. The Segformer-ResNet concurrent encoder reconciles the dimensional differences between the feature maps of ResNet and the feature embeddings of Segformer during the encoding stage through a feature alignment module, thereby achieving effective fusion of multi-dimensional features; the decoder of the optimized feature fusion module introduces depthwise separable convolutions to enhance the expression of context information and fuse multi-level features to improve the segmentation performance. The cross-domain mixing module uses the ClassMix method to mix images from the source domain and the target domain to obtain mixed images and generate corresponding pseudo-labels. During the training process of the teacher-student network, no gradients are propagated back to the teacher network, and the weights of the teacher network are updated using the exponential moving average (EMA) parameters from the student network. Through the unsupervised adaptive teacher-student network, the model can generate pseudo-labels of the target domain data by learning the features of the labeled source domain data, thus solving the problem of data annotation.
[0021] The SegFormer branch uses the encoder of SegFormer, integrating multiple SegFormer blocks to produce feature maps of different scales.
[0022] The ResNet branch uses the ResNet101 structure, including multiple identical Conv blocks, with batch normalization and ReLU activation functions added in the middle, as well as skip connections to enhance local feature extraction.
[0023] The Feature Align Model (FAM) effectively aligns the channel dimensions and spatial dimensions of feature maps from different sources, enhancing the model's ability to identify lesion areas.
[0024] Meanwhile, the optimized feature fusion module can adjust the size and number of channels of the encoder output feature maps, while fusing context-aware features through dilated convolution groups and obtaining the final semantic segmentation output.
[0025] Step S5: Input the dataset to be detected into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each spectral image of the target domain dataset.
[0026] Here, after inputting the source domain labeled image dataset and the target domain dataset into the low-labeling-dependent field-scale segmentation model for wheat diseases, an unsupervised adaptive teacher-student network is used to extract pseudo-labels of the target domain data, and an adaptive semantic segmentation network is used for wheat disease detection in the target domain.
[0027] In an embodiment of the present invention, first, a multi-spectral unmanned aerial vehicle is used to collect multiple groups of spectral images during the complete infection cycle of different types of wheat diseases, and a source domain dataset and a target domain dataset for wheat diseases are constructed; then, the background area, healthy wheat area, and wheat disease lesion area in the spectral images of the source domain dataset are marked to obtain a source domain labeled image dataset, and a dataset to be detected is constructed from the source domain labeled image dataset and the target domain dataset; finally, the dataset to be detected is input into a wheat disease field-scale segmentation model constructed based on an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network to detect the wheat disease lesion area in each spectral image of the target domain dataset. The present invention designs a low-labeling-dependent field-scale segmentation model through techniques such as domain adaptation, knowledge distillation, and dual-branch, enabling the model to generate pseudo-labels of the target domain data by learning the characteristics of the labeled source domain data, and then capturing the global and local information of the target domain images through a dual-branch structure, while considering feature fusion and multi-scale fusion, which helps to improve the accuracy and efficiency of the model in the task of detecting wheat disease lesion areas.
[0028] In some possible embodiments, a DJI Phantom 4 multi-spectral unmanned aerial vehicle is used to collect wheat pest and disease images in complex backgrounds on-site, and then the collected multi-spectral images are calibrated. The calibration process includes phase calibration, dark current correction, exposure time and sensor gain correction, phase correction, and lens distortion correction to obtain a source domain dataset and a target domain dataset. Then, the images in the source domain dataset are marked, and after marking, a random augmentation method is used to augment the marked source domain dataset to obtain a preset number of source domain labeled image datasets. Four methods, namely random rotation, mirror transformation, random cropping, and Cutout, are used to augment the collected dataset, and finally 7,190 groups of multi-spectral images are obtained.
[0029] Exemplarily, the five sample collection sites are Fengming Wheat Field in Qishan County, Baoji City, Lijiazhuang Wheat Field in Meixian County, Baoji City, Caoxinzhuang Wheat Field in Yangling District, Xianyang City, Dongcheng Wheat Field in Xingping District, Xianyang City, and Wuhe Wheat Field in Huayin District, Weinan City. The five experimental fields contain more than 50 different wheat varieties, covering the complete infection cycles of wheat yellow dwarf disease, scab, and rust. Subsequently, a DJI Phantom 4 multi-spectral unmanned aerial vehicle is used for imaging, and multiple groups of spectral images of the complete infection cycles of wheat yellow dwarf disease, scab, and rust can be used as the source domain dataset.
[0030] In the embodiments of the present invention, the background region, healthy wheat region, and wheat disease lesion region of the images in the source domain dataset are marked to obtain the source domain label image dataset; specifically, it includes: The data in the source domain dataset is used with the normalized difference vegetation index DNVI to segment the wheat region in the image from the background region to obtain the NDVI label image; part of the data in the source domain dataset is manually labeled to segment the healthy wheat region and the wheat disease lesion region to obtain the manually labeled label image; the remaining data in the source domain dataset except for the manual labeling is semi-automatically labeled using the adaptive semantic segmentation network model to segment the wheat disease lesion region in the spectral image to obtain the wheat disease lesion region label image; the NDVI label image, the manually labeled label image, and the wheat disease lesion region label image are synthesized to obtain the source domain label image.
[0031] Further, part of the data in the source domain dataset is manually labeled to segment the healthy wheat region and the wheat disease lesion region to obtain the manually labeled label image; the remaining data in the source domain dataset except for the manual labeling is semi-automatically labeled using the adaptive semantic segmentation network model to segment the wheat disease lesion region in the spectral image to obtain the wheat disease lesion region label image; it includes: S201. Manually label 30% of the data in the source domain dataset, and segment the healthy wheat region and the wheat disease lesion region in the spectral image to obtain the initial dataset; S202. Use the initial dataset to input into the adaptive semantic segmentation network model for initial training, and add the initial dataset to the labeled dataset; S203. Input the remaining 20% of the data in the source domain dataset into the adaptively semantically segmented network model after initial training to generate a temporary dataset; S204. Manually adjust the temporary dataset, and add the adjusted temporary dataset to the labeled dataset; S205. Use the labeled dataset to perform secondary training on the adaptively semantically segmented network model after initial training to generate a secondary temporary dataset; S2026. Manually adjust the secondary temporary dataset, and add the adjusted secondary temporary dataset to the labeled dataset; S207. Repeat the steps of S203 - S206 to perform multiple trainings on the adaptive semantic segmentation network model until all the data in the source domain dataset is labeled for the wheat disease lesion region to obtain the wheat disease lesion region label image.
[0032] In the above semi-automatic annotation process, the adaptive semantic segmentation network model was trained multiple times. Subsequently, the dataset to be detected is input into the wheat disease field-scale segmentation model constructed by the unsupervised adaptive teacher-student network and the adaptive semantic segmentation network after multiple trainings, and the results of detecting the wheat disease lesion areas in each spectral image in the target domain dataset are more accurate.
[0033] Furthermore, synthesizing the NDVI label image, the manually annotated label image, and the wheat disease lesion area label image to obtain the source domain label image includes: Taking the background area in the NDVI label image as the background area in the obtained source domain label image; Intersecting the wheat area in the NDVI label image with the healthy wheat area in the manually annotated label image to obtain the healthy wheat area in the source domain label image; Merging the wheat disease incidence areas in the wheat disease lesion area label image and the wheat disease incidence areas in the manually annotated label image as the wheat disease incidence areas in the source domain label image.
[0034] Here, the wheat disease field-scale segmentation model with low annotation dependence provided by the present invention integrates an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network. The unsupervised adaptive teacher-student network consists of a cross-domain mixing module and a teacher-student network, and is used to generate target domain pseudo-labels. At the same time, the adaptive semantic segmentation network adopts a SegFormer-ResNet concurrent encoder, and reconciles the dimensional differences between the feature maps of ResNet and the feature embeddings of SegFormer in the encoding stage through a feature alignment module, so as to achieve effective fusion of multi-dimensional features. The decoder introduces depthwise separable convolutions to further enhance the expression of context information and fuse multi-level features to improve the segmentation performance. The overall architecture of the model is as Figure 3 shown, the architecture of the unsupervised adaptive teacher-student network is as Figure 4 shown, and the structure of the adaptive semantic segmentation network is as Figure 5 shown.
[0035] The present invention uses the ClassMix method to construct a cross-domain mixing module. This can mix the images and corresponding labels in the source domain dataset and the images in the target domain dataset, and generate mixed samples and pseudo-labels. At the same time, color jitter and Gaussian blur are used to perform data augmentation on the mixed samples and pseudo-labels. The cross-domain mixing module includes a sample mixing module and a data augmentation module. The sample mixing module mixes the spectral images and corresponding labels in the source domain dataset with the spectra in the target domain dataset to generate mixed samples and pseudo-labels. The mixed samples, pseudo-labels, and the data in the source domain dataset are simultaneously subjected to data augmentation and input into the next network; The present invention uses a teacher-student network for pseudo-label optimization to evaluate and solve the problems that using only pseudo-labels for segmentation network training may cause the network to implicitly learn to distinguish the target domain and the source domain instead of learning feature transfer between domains, and that when the network infers in the target domain, there are significant biases for pixels of smaller-scale classes. The teacher-student network includes a teacher network and a student network. The student network receives data from the mixed sample data and the source domain dataset, outputs exponential moving average parameters to the teacher network, and outputs features to the subsequent network; the teacher network receives the mixed sample data from the cross-domain mixing module and the exponential moving average parameters of the student network, and outputs features to the subsequent network.
[0036] The present invention uses the encoder of SegFormer as the SegFormer branch. This can generate high-resolution rough features and low-resolution fine features. Four SegFormers are integrated to extract the necessary features. The output of each SegFormerBlock serves as the input to the subsequent Block at the same time and helps generate feature maps of different scales.
[0037] The ResNet branch aims to correct the problem of insufficient local feature extraction ability of the SegFormer branch. The ResNet branch consists of four identical Conv blocks. To alleviate the problem of vanishing gradients and thus enhance the stability and convergence speed of network training, batch normalization and ReLU activation functions are added in the middle of the convolutional layer. By using the SegFormer and ResNet branches, feature maps of the same size can be obtained from the multi-spectral image. The SegFormer branch gives priority to global information extraction, while the ResNet branch is more suitable for local information extraction. The effective fusion of these different feature maps is crucial for obtaining significant semantic segmentation results. Therefore, a Feature Alignment Model (FAM) is developed. This process concatenates the third and fourth feature maps generated by the two branches along the channel dimension. First, the input image is processed in parallel by the Segformer branch and the ResNet branch; secondly, the patch embeddings passing through the third Conv Block are aligned in terms of the number of channels using 1×1 convolution, then downsampled, and then added to the patch embeddings of SegFormer; then, the patch embeddings are fed back to the fourth Conv Block of the ResNet branch after passing through the fourth SegFormerBlock; finally, the final output feature map of the fourth ResNetBlock is resized using the same strategy and then added to the patch embeddings of the fourth SegFormerBlock to obtain the fused feature map.
[0038] Considering the differences in the sizes and numbers of channels of different feature maps, before feature fusion, simpler 1×1 convolutions and bilinear upsampling are used to replace the original MLP in SegFormer to adjust the size and number of channels of the encoder output feature map to match the first feature map, and then they are concatenated. Next, the feature fusion module feeds the concatenated features into a group of dilated convolutions for context-aware feature fusion, and finally obtains the final semantic segmentation output.
[0039] In some possible embodiments, the optimized feature fusion module is obtained by replacing the original multi-layer perceptron in the semantic segmentation network with 1×1 convolutions and bilinear upsampling and then introducing a group of dilated convolutions.
[0040] Here, although the encoder generates four feature maps with different scales for semantic segmentation, the MLP decoder of SegFormer usually only analyzes local information. However, incorporating context information into the decoder can enhance the robustness of semantic segmentation, which is an ideal property required for unsupervised domain adaptation. Therefore, in the design of the decoder, the present invention introduces a group of dilated convolution modules for feature fusion. The present invention proposes a simple and efficient optimized feature fusion module. Compared with the original MLP, the new feature extraction module uses 1×1 convolutions and bilinear upsampling to replace the original MLP in SegFormer to adjust the size and number of channels of the encoder output feature map to match the first feature map, and then they are concatenated. Next, the concatenated features are fed into a group of dilated convolutions for context-aware feature fusion. In the group of dilated convolutions, three parallel convolution operations are first performed: 1×1 convolution and two 3×3 dilated convolutions with different dilation rates. The outputs of these operations are concatenated and input into a subsequent 1×1 convolution to obtain the final semantic segmentation output. The MLP designed by the present invention has both low computational complexity and high performance.
[0041] In some possible embodiments, inputting the dataset to be detected into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each image in the target domain dataset includes: inputting each multi-spectral image in the dataset to be detected into the cross-domain mixing module and the data augmentation module to generate mixed samples and pseudo-labels; processing the source domain data and the mixed sample data through the teacher-student network and then inputting them into the SegFormer branch and the ResNet branch to extract wheat disease features; fusing the third and fourth feature maps output by each channel of the SegFormer branch and the ResNet branch through the feature alignment module to obtain a fused feature map; first aligning the sizes and numbers of channels of the encoder outputs through the optimized feature fusion module, and then feeding the features into the dilated convolution group for context-aware feature fusion; connecting the output of the optimized feature fusion module and inputting it into a 1×1 convolution to obtain the final semantic segmentation output, and finally segmenting the wheat disease lesion areas in the target domain in the dataset to be detected.
[0042] Here, the source domain label image dataset and the target domain dataset are constructed into a dataset to be detected and input into the low-labeling-dependent field-scale segmentation model. First, an unsupervised adaptive teacher-student network is used to generate pseudo-labels for the target domain, and an adaptive semantic segmentation network is used to segment the wheat disease areas. Finally, the wheat disease in the target domain is segmented through the segmentation mask, so as to accurately monitor the wheat disease lesion areas in the target domain.
[0043] The following uses a specific experimental example to illustrate the above-mentioned unsupervised domain adaptation method for accurate detection of drone wheat diseases. However, it should be noted that this specific experimental example is only for better explaining the present invention and does not constitute an improper limitation to the present invention.
[0044] The present invention has developed a low-labeling-dependent field-scale segmentation model (DATS-ASSFormer), and used multi-spectral images and RGB images as the inputs of the model respectively, and adopted the same training method for comparative testing. The experimental results show that the prediction results of the model for multi-spectral images are better than those of RGB. When using multi-spectral images, the performance of the DATS-ASSFormer model is better than that when using RGB images in terms of both supervised (Oracle) and unsupervised domain adaptation (UDA) performance. For example, when the source domain data is an image of Fusarium head blight of wheat and the target domain data is an image of wheat yellow dwarf disease, the mIoU of RGB images in Oracle and UDA are 70.48% and 48.88% respectively, and the mIoU of multi-spectral images in Oracle and UDA are 79.25% and 62.97% respectively. That is to say, using multi-spectral images as input makes the model have higher performance than using RGB images.
[0045] To verify the effectiveness of the unsupervised adaptive teacher-student network provided by the present invention on the experimental results, comparative experiments on the teacher-student domain adaptation performance of different backbone networks were carried out using other well-known semantic segmentation models. The experimental results show that in the domain adaptation experiment from wheat yellow dwarf disease to wheat scab, the mIoU of SegFormer in Oracle and UDA are 78.23% and 60.14% respectively, and the relative domain adaptation performance (Rel.) is 76.88%, which is 2.1%, 1.9% and 6.32% higher than that of DANet, Deeplabv2 and Deeplabv3 on Oracle respectively; on UDA, they are 7.03%, 8.82% and 11.92% higher respectively; on Rel., they are 7.12%, 9.65% and 9.82% higher respectively. It can be seen that the unsupervised adaptive teacher-student network provided by the present invention has stronger segmentation performance for the lesion areas of wheat diseases.
[0046] To verify the effectiveness of the adaptive semantic segmentation network provided by the present invention on the experimental results, comparative experiments on the influence of different encoder-decoder combinations on the domain adaptation performance were carried out using the SegFormer and Deeplabv3 semantic segmentation models. The experimental results show that in the domain adaptation experiment from wheat yellow dwarf disease to wheat scab, the combination of the encoder and decoder of the ASSFormer provided by the present invention performs optimally, and the mIoU in Oracle and UDA are 86.70% and 70.16% respectively, and Rel. is 80.92, which is 0.43%, 1.71% and 1.58% higher than that of the sub-optimal combination of ASSFormer encoder plus SegFormer decoder on Oracle, UDA and Rel. respectively. It can be seen that the adaptive semantic segmentation network provided by the present invention has stronger segmentation performance for the lesion areas of wheat diseases.
[0047] Through experiments, the DATS-ASSFormer model designed by the present invention achieved accurate segmentation of wheat diseases. In the domain adaptation experiments from wheat rust to wheat yellow dwarf, the mIoU on Oracle and UDA were 86.72% and 69.1% respectively; in the domain adaptation experiments from wheat scab to wheat yellow dwarf, the mIoU on Oracle and UDA were 79.25% and 62.97% respectively; in the domain adaptation experiments from wheat yellow dwarf to wheat rust, the mIoU on Oracle and UDA were 88.67% and 68.01% respectively; in the domain adaptation experiments from wheat scab to wheat rust, the mIoU on Oracle and UDA were 78.46% and 61.45% respectively; in the domain adaptation experiments from wheat yellow dwarf to wheat scab, the mIoU on Oracle and UDA were 84.87% and 66.13% respectively; in the domain adaptation experiments from wheat rust to wheat scab, the mIoU on Oracle and UDA were 84.18% and 65.18% respectively. In summary, the DATS-ASSFormer model performed excellently in the semantic segmentation of multi-spectral images and in the segmentation of wheat diseases, demonstrating the adaptability and robustness of the model in different wheat disease scenarios. The successful development of the DATS-ASSFormer model provides substantial technical support for future precise agricultural crop disease monitoring and management, showing the broad prospects of the application of deep learning technology in the agricultural field.
[0048] It should be noted that in the embodiments of the present invention, if the above-mentioned unsupervised domain adaptation method for precise detection of wheat diseases by drones is implemented in the form of software functional modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable an electronic device to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0049] Correspondingly, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in any one of the above-described unsupervised domain adaptation-based precise detection methods for wheat diseases by drones are implemented. Correspondingly, in an embodiment of the present invention, a computer program product is further provided. When the computer program product is executed by a processor of an electronic device, it is used to implement the steps in any one of the above-described unsupervised domain adaptation-based precise detection methods for wheat diseases by drones.
[0050] Based on the same inventive concept, an embodiment of the present invention provides an electronic device for implementing the unsupervised domain adaptation-based precise detection method for wheat diseases by drones described in the above method embodiment. The electronic device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, the steps in any one of the unsupervised domain adaptation-based precise detection methods for wheat diseases by drones in the embodiments of the present invention are implemented.
[0051] The memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed by the processor and each module in the electronic device. It can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). When the processor executes the program, the steps of the unsupervised domain adaptation-based precise detection method for wheat diseases described in any of the above are implemented. The processor generally controls the overall operation of the electronic device.
[0052] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the descriptions of the method embodiments of the present invention for understanding.
[0053] It should be understood that the term "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present invention. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the order numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0054] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.
[0055] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the couplings, direct couplings or communication connections between the various components shown or discussed can be through some interfaces, and the indirect couplings or communication connections of devices or units can be electrical, mechanical or other forms.
[0056] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0057] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0058] The methods disclosed in several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0059] The above is only the implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An unsupervised domain adaptation method for precise detection of wheat diseases by unmanned aerial vehicles, characterized in that a source domain dataset and a target domain dataset of wheat diseases are constructed. The source domain dataset and the target domain dataset are constructed by multiple images within the complete infection cycle of wheat diseases collected by unmanned aerial vehicles. The source domain dataset of wheat diseases includes a dataset with wheat disease category labels, and the target domain dataset of wheat diseases includes a dataset without wheat disease category labels. The wheat diseases in the source domain dataset and the target domain dataset are of different types. The background area, healthy wheat area, and wheat disease lesion area in the images of the source domain dataset are marked with labels to obtain a source domain label image dataset. A dataset to be detected is constructed from the source domain label image dataset and the target domain dataset. A wheat disease field-scale segmentation model with low annotation dependence is constructed based on an unsupervised adaptive teacher-student network and an adaptive semantic segmentation network. Among them, the unsupervised adaptive teacher-student network includes a cross-domain mixing module and a teacher-student network, which are used to generate pseudo-labels for the target domain. The adaptive semantic segmentation network includes a decoder with a Segformer-ResNet concurrent encoder and an optimized feature fusion module. The Segformer-ResNet concurrent encoder reconciles the dimensional differences between the feature maps of ResNet and the feature embeddings of Segformer in the encoding stage through a feature alignment module, so as to achieve effective fusion of multi-dimensional features. The decoder of the optimized feature fusion module introduces depthwise separable convolutions to enhance the expression of context information and fuse multi-level features to improve the segmentation performance. The dataset to be detected is input into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each image of the target domain dataset.
2. The method according to claim 1, wherein The source domain dataset and the target domain dataset are constructed by multiple images within the complete infection cycle of wheat diseases, specifically including: the source domain dataset and the target domain dataset are constructed by multiple groups of spectral images within the complete infection cycle of wheat diseases collected by unmanned aerial vehicles.
3. The method according to claim 2, characterized in that The source domain dataset and the target domain dataset are constructed by multiple groups of spectral images within the complete infection cycle of wheat diseases collected by unmanned aerial vehicles, including: using a spectral unmanned aerial vehicle to image the wheat pest and disease images in a preset experimental field to obtain multiple groups of multispectral images. The preset experimental field contains multiple different wheat varieties and covers the complete disease-susceptible cycle of wheat diseases. Each group of the multispectral images includes an RGB image and multispectral band images of five bands.
4. The method according to claim 2, wherein The cross-domain mixing module includes a sample mixing module and a data augmentation module. The sample mixing module mixes the images and corresponding labels in the source domain dataset with the spectra in the target domain dataset to generate mixed samples and pseudo-labels. The mixed samples, pseudo-labels, and the data in the source domain dataset are simultaneously augmented by the data augmentation module and input into the next network. The teacher-student network includes a teacher network and a student network. The student network receives the mixed sample data and the data in the source domain dataset, outputs the exponential moving average parameters to the teacher network, and outputs features to the subsequent network. The teacher network receives the mixed sample data from the cross-domain mixing module and the exponential moving average parameters of the student network, and outputs features to the subsequent network.
5. The method according to any one of claims 1 to 4, characterized in that The SegFormer branch includes four SegFormer blocks connected in sequence. The output of the previous SegFormer block serves as the input to the subsequent SegFormer block simultaneously, and is used to generate feature maps of different scales. The ResNet branch part includes four identical convolutional blocks connected in sequence. Batch normalization and ReLU activation functions are added between adjacent convolutional blocks. The feature alignment module is used to connect the feature maps generated by the third and fourth blocks of the two branches along the channel dimension, and perform feature fusion between the channel features and the feature maps to obtain a fused feature map. The optimized feature fusion module is obtained by replacing the original multi-layer perceptron in the adaptive semantic segmentation network with a 1×1 convolution and bilinear upsampling, and then introducing a dilated convolution group.
6. The method according to any one of claims 1 to 4, characterized in that, Inputting the dataset to be detected into the wheat disease field-scale segmentation model to detect the wheat disease lesion areas in each image in the target domain dataset, including: Inputting each image in the dataset to be detected into the cross-domain mixing module and the data augmentation module to generate mixed samples and pseudo-labels. Processing the data in the source domain dataset and the mixed sample data through the teacher-student network, and then inputting them into the SegFormer branch and the ResNet branch to extract wheat disease features. Fusing the third and fourth feature maps output by each channel of the SegFormer branch and the ResNet branch through the feature alignment module to obtain a fused feature map. First aligning the size and number of channels of the encoder output through the optimized feature fusion module, and then feeding the features into the dilated convolution group for context-aware feature fusion. Connecting and inputting the output of the optimized feature fusion module into a 1×1 convolution to obtain the final semantic segmentation output, and finally segmenting the wheat disease lesion areas in the target domain in the target domain dataset.
7. The method according to any one of claims 1 to 4, characterized in that Marking the background areas, healthy wheat areas, and wheat disease lesion areas in the images in the source domain dataset. After marking, a random augmentation method is used to augment the marked source domain dataset to obtain a source domain labeled image dataset with a preset number.
Citation Information
Patent Citations
Wheat yellow dwarf disease multi-scale detection method based on unmanned aerial vehicle multi-spectral image
CN118537731A