Bursaphelenchus xylophilus tree detection method based on unmanned aerial vehicle vision
By constructing a target detection model based on the Color-ViT-YOLOv8 algorithm, the problem of missed detection and missed detection in the detection of pine nematode disease tree is solved, and efficient and accurate disease tree detection is achieved to adapt to complex environments and different disease tree stages.
Patent Information
- Application Number
- CN202510450434.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing drone remote sensing technology relies on human visual judgment in the detection of pine nematode disease trees. It has high labor intensity, strong subjectivity, and limited accuracy. It is prone to false detection and missed detection in complex scenarios.
The object detection model based on the Color-ViT-YOLOv8 algorithm is adopted. By replacing the backbone network of YOLOv8 as a ViT module, and adding the Color module before the ViT module, the channel selection and fusion technology of multi-dimensional color space is used to build a target detection model, extract color features at different development stages and environments, reduce background interference, and improve detection accuracy.
It effectively solves the problems of false detection and missed detection in complex scenarios, improves the accuracy and stability of pine nematode disease tree detection, adapts to the development stages of different diseased trees and ambient lighting changes, and enhances the robustness of the model.
Smart Images

Figure CN120375233A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and particularly relates to a method for detecting pine wilt disease trees based on UAV vision. Background Technique
[0002] The UAV remote sensing technology has developed rapidly and plays a crucial role in fields such as geological exploration, emergency rescue, and agricultural management. In forest protection work, UAVs can provide a brand-new macroscopic perspective and an efficient data acquisition means: on the one hand, compared with traditional manual patrols, UAVs have a broad aerial perspective, can cover large areas of pine forest regions, and obtain multi-dimensional geographical information; on the other hand, UAVs are highly flexible and can adapt to diverse shooting requirements and complex and changeable environmental conditions. Therefore, the development of disease tree detection technology based on UAV remote sensing information is crucial for future efficient and economical forest protection work.
[0003] Currently, some studies have used UAV remote sensing images to identify pine wilt disease trees. This method can collect pine forest images over large areas in a relatively short time, greatly improving the monitoring efficiency. However, in actual operation, image analysis still highly relies on human visual judgment, with high labor intensity, strong subjectivity, and limited accuracy. In view of the above deficiencies, there is an urgent need to develop a reliable, efficient, and intelligent method for detecting pine wilt disease trees. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a method for detecting pine wilt disease trees based on UAV vision. The technical problems to be solved by the present invention are realized through the following technical solutions: A method for detecting pine wilt disease trees based on UAV vision, comprising: Obtaining pine forest images and geographical information collected by UAV aerial photography; Fusing all pine forest images using geographical information to obtain a fused pine forest image; Dividing the fused pine forest image into multiple image blocks of the same size; performing data cleaning on the multiple image blocks of the same size to obtain an image block set; and performing label annotation of pine wilt disease trees on each image block in the image block set, and obtaining a training set according to the labeled image block set; Constructing an object detection model based on the Color-ViT-YOLOv8 algorithm; wherein the object detection model replaces the backbone network in YOLOv8 with a ViT module, and adds a Color module before the ViT module; the Color module processes the input image using channel selection and fusion techniques in a multi-dimensional color space; The target detection model is trained using the training set, and a trained target detection model is obtained after a training stop condition is reached, so as to be used for detecting pine wood nematode diseased trees in a pine forest image to be tested.
[0005] Beneficial effects of the present invention: The embodiment of the present invention proposes a method for detecting pine wilt diseased trees based on drone vision, which is implemented based on the target detection model of the Color-ViT-YOLOv8 algorithm. The color model is combined with the ViT detection model and the YOLOv8 detection model to extract different color features for different development stages of diseased trees, effectively dealing with the effects of different development stages of diseased trees, changes in ambient light, and different growth cycles of trees. At the same time, the model uses complex multidimensional color space fusion technology to extract more diverse color information, avoiding the limitations of a single color space, and has significant advantages in dealing with complex environments and different development stages of diseased trees. It can solve the current problems of false detection and missed detection in complex scenes in the detection of pine wilt diseased trees. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 A schematic diagram of a flow chart of a method for detecting pine wood nematode diseased trees based on drone vision provided by an embodiment of the present invention; Figure 2 A schematic diagram of a process of a method for detecting pine wood nematode diseased trees based on drone vision provided by an embodiment of the present invention; Figure 3 A schematic diagram of the specific structure of the target detection model provided by an embodiment of the present invention; Figure 4 Output results of pine wood nematode diseased trees in nine channels; Figure 5 This is a comparison of the original image of the pine wood nematode diseased tree and the fusion result of the three channels of R, H, and A; Figure 6 This is the flow chart of diseased tree detection based on Color-ViT-YOLOv8. DETAILED DESCRIPTION
[0007] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0008] The present invention provides a method for detecting pine wood nematode diseased trees based on drone vision. Figure 1 and Figure 2 As shown, the method may include the following steps: S1, obtain the pine forest images and geographic information collected by drone aerial photography; Specifically, S1 includes: Obtain pine forest images and geographical information collected by aerial photography at different positions in the target pine forest area when a drone of the same model flies within a preset height range from the ground within a preset time period.
[0009] The target pine forest area is a large pine forest area. A drone of the same model continuously flies over the target pine forest area and conducts aerial photography for different positions. For each moment of collection, pine forest images and geographical information of the corresponding area can be collected. During the aerial photography process, the drone can fly along a certain trajectory to complete traversing all positions in the target pine forest area.
[0010] The preset time period is the working time period for aerial photography. To ensure sufficient light and exclude the influence of irrelevant factors, the preset time period can be selected as a part of the daytime. For example, in an optional implementation manner, the preset time period is from 10:00 to 17:00; that is, the collection time . Of course, it is also reasonable to select a time period between 10:00 and 17:00, such as from 10:00 to 12:00.
[0011] To ensure the consistency of parameters such as the resolution and size of the collected images, the height from the ground when collecting pine forest images should be kept within the preset height range The value can be set according to needs. For example, in an optional implementation manner, the preset height range is from 60 to 100 meters. That is . Of course, a fixed value within the preset height range can also be selected, such as 80 meters, so as to keep the drone flying at the same height.
[0012] Store the collected pine forest images in TIF format, store the geographical information in TFW format, and store the image coordinate system in WCG84.
[0013] S2. Use the geographical information to fuse all the pine forest images to obtain a fused pine forest image; The purpose of S2 is to obtain a fused pine forest image by fusing all the collected pine forest images at different positions, which represents the complete image of the target pine forest area. Specifically, S2 may include: S21. Use the corresponding geographical information to obtain the image boundary coordinates of each pine forest image; Specifically, calculate the image boundary coordinates by parsing the coordinate information of the TFW file.
[0014] S22. Continuously perform image fusion processing on the pine forest images with adjacent image boundary coordinates until a fused pine forest image is obtained.
[0015] Among them, performing image fusion processing on two adjacent pine forest images with image boundary coordinates includes: 1) Use the image boundary coordinates to determine whether there is a discontinuity in the edge regions of two adjacent pine forest images with respect to the image boundary coordinates; If the values of the image boundary coordinates of the two pine forest images are continuous and there are no gaps, then there is no discontinuity in the edge region.
[0016] 2) If there is no discontinuity, splice and fuse the two adjacent pine forest images with respect to the corresponding image boundary coordinates; 3) If there is a discontinuity, after splicing and fusing the two adjacent pine forest images with respect to the corresponding image boundary coordinates, convert the pixel points in the discontinuous part from the three-channel RGB to the four-channel RGBA, and define the transparency of the A channel as 0.
[0017] It can be understood that two adjacent pine forest images with respect to the image boundary coordinates can be subjected to image fusion processing, and then another adjacent pine forest image with respect to the image boundary coordinates can be fused. Repeat this process until a fused pine forest image is obtained.
[0018] S3. Divide the fused pine forest image into multiple image blocks of the same size; perform data cleaning on the multiple image blocks of the same size to obtain an image block set; and perform label annotation of pine wilt disease trees on each image block in the image block set, and obtain a training set based on the labeled image block set; Dividing the fused pine forest image into multiple image blocks of the same size is to perform image subframing on the fused pine forest image. The specific process includes: According to a preset step size and a preset image block size, perform overlapping cutting on the fused pine forest image to obtain multiple divided image blocks.
[0019] The preset image block size can be expressed as , and the preset step size can be expressed as ; ; Among them, is a proportionality coefficient related to the size of the image block; is the length of the image block.
[0020] The preset image block size can be set as needed. For example, it can be ; The value of can be 0.5, then the preset step size
[0021] In S3, preprocessing the multiple image blocks of the same size to obtain an image block set includes: 1) Delete the image blocks with noise in the multiple image blocks of the same size; 2) Convert the retained image patches from four-channel RGBA to three-channel RGB, and form a set of image patches consisting of all the converted image patches.
[0022] Since there will be two types of image patches after the aerial drone images are mosaicked: noise-free image patches and noisy image patches. It is necessary to perform the above data cleaning on the mosaicked images to obtain a set of image patches.
[0023] Perform label annotation of pine wilt disease trees on each image patch in the set of image patches. It can be completed with the help of an annotation tool and expert experience. The set of annotated image patches can be divided into a training set and a test set according to a ratio. The division ratio can be selected according to needs, for example, it can be 8:2.
[0024] S4. Construct an object detection model based on the Color-ViT-YOLOv8 algorithm; wherein, in the object detection model, the ViT module is used to replace the backbone network in YOLOv8, and a Color module is added before the ViT module; the Color module processes the input image by using the channel selection and fusion technology in the multi-dimensional color space. The present invention is an improvement based on the YOLOv8 as the basic framework. The backbone network in YOLOv8 is replaced with a ViT module (fully called Vision Transformer module) to improve the global feature extraction ability, and a Color module is added before the ViT module.
[0025] In the YOLOv8 model, the traditional CNN backbone network is mainly used to extract local features. Although it is efficient, it has problems such as insufficient global information and weak long-distance dependence. The present invention replaces the backbone network with a ViT module, which can significantly improve the global feature extraction ability. The ViT captures long-distance dependencies through the self-attention mechanism, solves the local feature extraction limitation of the traditional CNN (Convolutional Neural Network) in complex backgrounds, and enhances the adaptability to different development stages of diseased trees. In addition, the ViT can also effectively reduce background interference, improve the robustness of the model in complex environments, and make the detection of diseased trees more accurate and stable. For the specific structure of the object detection model of the present invention, please refer to Figure 3 as shown. Figure 3 For the abbreviations of each module in
[0026] At present, certain progress has been made in the research on detection methods for trees infected with Bursaphelenchus xylophilus. However, the existing color models mainly rely on static feature extraction and still have certain limitations, urgently requiring a technological breakthrough. The Color module in the embodiments of the present invention can extract different color features for different development stages of diseased trees, effectively addressing the impacts of different development stages of diseased trees, environmental light changes, and different tree growth cycles, and avoiding the limitations of traditional static color models. At the same time, the Color module adopts channel selection and fusion techniques in a multi-dimensional color space, which can extract more diverse color information, avoid the limitations of a single color space, and has significant advantages in dealing with complex environments.
[0027] Specifically, the Color module is used for: 1) For the input image, select the output of the R channel in the RGB color space, select the output of the H channel in the HSV color space, and select the output of the A channel in the LAB color space; To highlight the features of diseased trees and reduce background interference, the Color module selects specific channels with a large difference in the foreground diseased trees and the background in different color spaces to adapt to the diseased tree detection task. In the RGB color space, the R channel can capture the red information of the lesion, which has a significant effect in the detection of Bursaphelenchus xylophilus-infected trees. In different development stages of diseased trees, there will be changes in reddish-brown or withered yellow, and the R channel can effectively capture this color feature; in the HSV color space, the H channel can capture stable hue information. Under different lighting conditions, the color of diseased trees may have light and dark changes, and the H channel is not easily affected by brightness and can adapt to different lighting conditions; in the LAB color space, the A channel can capture the red-green contrast characteristics, which can effectively distinguish between non-diseased and diseased tree areas, reduce background interference, and improve the model's attention to the diseased tree area.
[0028] Please refer to Figure 4 for understanding, Figure 4 which is the output result of the diseased tree in nine channels.
[0029] 2) Normalize the outputs of the R channel, the H channel, and the A channel to the same channel value range and then fuse them to obtain a new fused three-channel image; Fuse the selected R, H, and A channels for output. The channel value ranges in different color spaces are different. Normalize them to the same channel value range [0, 255] and fuse them into a new three-channel image. Compared with a single color space, the fused image can enhance the features of diseased trees, highlight the diseased areas, and reduce environmental interference.
[0030] Please refer to Figure 5 for understanding, Figure 5 which is the original image of the Bursaphelenchus xylophilus-infected tree (please refer to Figure 5 Figure (a) ofFigure 5 (b) Comparison chart of .
[0031] 3) converting the fused new three-channel image into a grayscale image, and using the OTSU (Nobuyuki Otsu method) method to adaptively calculate the optimal threshold to obtain a binary image; This step can further enhance the significance of the diseased tree area, reduce background interference, improve the distinction of the diseased area, and make the diseased tree area clearer in the binarization result. The OTSU method can adaptively select the optimal threshold and can stably and clearly segment the diseased tree area under different conditions.
[0032] For the specific process, please refer to OTSU related technology, which will not be described in detail here.
[0033] 3) performing morphological operations on the binary image to correct the connectivity of the diseased area and remove noise; Morphological processing can be used to correct the connectivity of the diseased area, remove noise and improve segmentation accuracy.
[0034] 4) The image after morphological operation is combined with the area screening strategy to remove invalid targets and finally retain the qualified detection area.
[0035] The area screening strategy can be to ignore areas that are too small. This can be achieved by using a set area threshold. Invalid targets can be removed through screening, and the remaining targets are used as qualified detection areas.
[0036] The above process can be understood by referring to 6. Figure 6 This is the flow chart of diseased tree detection based on Color-ViT-YOLOv8. Figure 6 The image after the result marking is the image after the qualified detection area is finally retained, which will be used as the input of the subsequent model structure.
[0037] S5, using the training set to train the target detection model, and after reaching a training stop condition, obtaining a trained target detection model for use in detecting pine wood nematode diseased trees in the pine forest image to be tested.
[0038] Wherein, the training stop condition is: In continuous Within training rounds, the detection accuracy change of the target detection model is always less than or equal to the threshold .
[0039] It is expressed as: ; in, Indicates that the model Accuracy during round training; is a preset threshold; if within consecutive training rounds, the performance change of the model is always less than or equal to the threshold , it indicates that the accuracy of the model on the training set tends to be stable, the training stops, and the current model is saved as the final training result. Among them, and can be set as needed. For example, is 0.1%, is 50.
[0040] For the specific training process, it can be understood by referring to the training process of a general model and will not be elaborated here.
[0041] After obtaining the trained object detection model, the pine wilt disease trees can be detected in the to-be-tested pine forest image. The to-be-tested pine forest image can be from the test set or other pine forest images that need to be detected, so as to obtain the detection result. It is expressed as: ; Among them, represents the to-be-tested pine forest image; represents model inference; represents the final detection result of pine wilt disease trees; is the probability that the preselected box is a diseased tree; is the set diseased tree probability threshold; when the predicted probability of the preselected box exceeds the threshold , it will be marked as a diseased tree. The flow chart of diseased tree detection based on Color-ViT-YOLOv8 is as shown in Figure 6 .
[0042] In an optional implementation, after detecting pine wilt disease trees in the to-be-tested pine forest image, the method further includes: If there are differences in the external environmental conditions between the to-be-tested pine forest image and the pine forest images in the training set, after processing the detection result corresponding to the to-be-tested pine forest image, it is used as new training data for iterative training of the object detection model, so as to update the model parameters.
[0043] The present invention can perform inference on the to-be-tested pine forest image using the trained object detection model to obtain the final detection result; then judge the to-be-tested pine forest image that has been inferred. If it is consistent with the images in the original training set in terms of external environmental conditions such as time, light, and weather, it will not be made into a training set; if there are differences, the detection result of this to-be-tested pine forest image will be processed to make a new training set, and then these new training sets will be mixed with the original training set to repeatedly iterate and train the model to update the model parameters.
[0044] An embodiment of the present invention proposes a method for detecting pine wilt disease trees based on UAV vision, which is implemented based on an object detection model of the Color-ViT-YOLOv8 algorithm. The Color module and the ViT module are introduced into the YOLOv8 detection model to extract different color features for different development stages of diseased trees, effectively coping with the impacts of different development stages of diseased trees, environmental light changes, and different tree growth cycles. At the same time, the model adopts a complex multi-dimensional color space fusion technology, which can extract more diverse color information, avoiding the limitations of a single color space, and having significant advantages in dealing with complex environments and different development stages of diseased trees, and being able to solve the problems of false detection and missed detection in the detection of pine wilt disease trees in complex scenarios.
[0045] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0046] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A detection method for pine wilt disease trees based on UAV vision, characterized in that, Including: Obtain the pine forest images and geographical information collected by drone aerial photography; Fuse all the pine forest images using the geographical information to obtain a fused pine forest image; Divide the fused pine forest image into multiple image blocks of the same size; perform data cleaning on the multiple image blocks of the same size to obtain an image block set; and label each image block in the image block set with the label of the pine wilt disease tree, and obtain a training set according to the labeled image block set; Construct an object detection model based on the Color-ViT-YOLOv8 algorithm; wherein the object detection model replaces the backbone network in YOLOv8 with a ViT module, and adds a Color module before the ViT module; the Color module processes the input image using the channel selection and fusion technology of the multi-dimensional color space; Use the training set to train the object detection model, and obtain a trained object detection model after reaching the training stop condition, so as to detect the pine wilt disease tree in the pine forest image to be measured.
2. The method according to claim 1, wherein The obtaining of the pine forest images and geographical information collected by drone aerial photography includes: Obtain the pine forest images and geographical information collected by aerial photography at different positions of the target pine forest area when a drone of the same model flies within a preset height range from the ground within a preset time period.
3. The method according to claim 2, wherein The preset time period is from 10 o'clock to 17 o'clock; the preset height range is from 60 to 100 meters.
4. The method according to claim 1, characterized in that The fusing of all the pine forest images using the geographical information to obtain a fused pine forest image includes: Use the corresponding geographical information to obtain the image boundary coordinates of each pine forest image; Continuously perform image fusion processing on the pine forest images with adjacent image boundary coordinates until a fused pine forest image is obtained.
5. The method according to claim 4, characterized in that Performing image fusion processing on two adjacent pine forest images with image boundary coordinates includes: Use the image boundary coordinates to judge whether there is a discontinuity in the edge area of two adjacent pine forest images with image boundary coordinates; If there is no discontinuity, splice and fuse the two adjacent pine forest images with image boundary coordinates according to the corresponding image boundary coordinates; If there is a discontinuity, after splicing and fusing the two adjacent pine forest images with image boundary coordinates according to the corresponding image boundary coordinates, convert the pixel points of the discontinuous part from three-channel RGB to four-channel RGBA, and define the transparency of the A channel as 0.
6. The method according to claim 1, wherein Dividing the fused pine forest image into multiple image blocks of the same size includes: According to the preset step size and preset image block size, perform overlapping cutting on the fused pine forest image to obtain multiple divided image blocks.
7. The method according to claim 1, characterized in that, Performing data cleaning on the multiple image blocks of the same size to obtain an image block set includes: Delete the image blocks with noise in the multiple image blocks of the same size; Convert the remaining image blocks from four-channel RGBA to three-channel RGB, and form an image block set from all the converted image blocks.
8. The method according to claim 1, wherein The Color module is used for: For the input image, select the output of the R channel in the RGB color space, select the output of the H channel in the HSV color space, and select the output of the A channel in the LAB color space; Fuse the outputs of the R channel, the H channel, and the A channel after normalizing them to the same channel value range to obtain a new fused three-channel image; Convert the fused new three-channel image into a grayscale image and adaptively calculate the optimal threshold using the OTSU method to obtain a binary image; Perform morphological operations on the binary image to correct the connectivity of the disease areas and remove noise; Combine the area screening strategy with the image after morphological operations to remove invalid targets and finally retain the eligible detection areas.
9. The method according to claim 1, wherein The training stop condition is: Within consecutive training rounds, the change in the detection accuracy of the target detection model is always less than or equal to a threshold .
10. The method according to claim 1, characterized in that After detecting pine wilt disease trees in the pine forest image to be tested, the method further includes: If there are differences in the external environmental conditions between the pine forest image to be tested and the pine forest images in the training set, process the detection results corresponding to the pine forest image to be tested and use them as new training data for iterative training of the object detection model, thereby realizing model parameter update.
Citation Information
Patent Citations
Image enhancement system and method for fault detection
CN112184588A
Method and system for constructing pine wood nematode disease recognition model in unmanned aerial vehicle remote sensing image for complex interference environment
CN118521887A
Scene-agnostic regression on pixel-level annotations
EP4471663A1