Crop disease, pest and weed identification and pesticide application system and method based on multi-modal fusion
The multimodal fusion crop disease, pest and weed identification system, which combines hyperspectral imaging and RGB imaging technology, enables accurate identification and targeted application of pesticides for crop diseases, pests and weeds. This solves the problem of difficulty in accurately distinguishing biological species in existing technologies and improves the accuracy of pesticide application and operational efficiency.
Patent Information
- Application Number
- CN202511966596.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-17
AI Technical Summary
Existing crop disease, pest and weed identification and application systems are unable to accurately distinguish the species of organisms causing stress, making it impossible to achieve precise application of pesticides.
The crop disease, pest and weed identification and application system adopts multimodal fusion, including a multimodal data acquisition unit, an edge computing and intelligent identification unit, and a precise decision-making and targeted application unit. It acquires multimodal data of crops through a hyperspectral imaging module and a global shutter RGB industrial camera, obtains geographical location information by combining a GNSS/IMU module, identifies diseases, pests and weeds using a lightweight multimodal fusion identification model, and achieves precise spraying through the targeted application unit.
It has achieved a shift from "single identification" to "three-in-one identification of diseases, pests, and weeds," which has improved the accuracy of pesticide application, reduced the amount of pesticides used, lowered production costs, and increased operational efficiency.
Smart Images

Figure CN121686520A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart agriculture technology and relates to a system and method for identifying and applying pesticides to crop diseases, pests and weeds based on multimodal fusion. Background Technology
[0002] Currently, the main methods for controlling crop diseases, pests, and weeds (hereinafter referred to as "disease-pest-weed") include manual inspection and pesticide application, visible light-based automatic identification, hyperspectral remote sensing monitoring, single-variable pesticide application, and cloud-based processing models. Manual inspection and pesticide application relies heavily on experience and visual judgment, involving manual operation of sprayers. This method is extremely inefficient, highly subjective, unable to detect diseases early, and leads to serious pesticide overuse. Visible light-based automatic identification mainly uses RGB cameras and image recognition algorithms to identify crop diseases, pests, and weeds, but this method can only identify visible lesions, large pests, and weeds, and cannot perform early disease diagnosis. Hyperspectral remote sensing monitoring technology mainly uses hyperspectral cameras for remote sensing analysis to identify crop diseases, pests, and weeds. While this method excels at physiological inversion, it lacks accuracy in identifying pest and weed species, and the equipment is expensive and the analysis is complex. Single-variable pesticide application is based on prescription maps or a single sensor (such as NDVI) for variable spraying, but this method has a slow response, cannot distinguish between diseases, pests, and weeds, and can only adjust the dosage, not the type of pesticide. The cloud-based processing model transmits image data back to the cloud server for processing, which is highly dependent on the network and has high latency, making it unable to meet the real-time control requirements of mobile platforms.
[0003] In summary, existing technical solutions mostly address only surface symptoms. For example, visible light-based systems struggle to detect internal physiological changes in crops, while hyperspectral technology, although capable of sensing physiological stress, struggles to accurately identify the organism causing the stress (e.g., whether it's a disease or a pest). Furthermore, the disconnect between identification and execution means that even with accurate identification, truly targeted treatment is impossible. Summary of the Invention
[0004] One objective of this invention is to provide a crop pest, disease and weed identification and application system based on multimodal fusion, which solves the problem that existing crop pest, disease and weed identification and application systems are unable to accurately distinguish the species of organisms causing stress, thus making it difficult to apply pesticides accurately.
[0005] Another objective of this invention is to provide a method for identifying and applying pesticides to crop diseases, pests, and weeds based on multimodal fusion.
[0006] The first technical solution adopted in this invention is a crop disease, pest, and weed identification and application system based on multimodal fusion, including a multimodal data acquisition unit, an edge computing and intelligent recognition unit, and a precise decision-making and targeted application unit. The multimodal data acquisition unit is used to collect visible light information, fine spectral information of the canopy, and geographical location information of crops. The edge computing and intelligent recognition unit is used to preprocess the collected multimodal data, then perform cross-modal attention fusion, compare the information features with the stored disease, pest, and weed database, and determine the disease, pest, and weed information of crops in the image. The precise decision-making and targeted application unit is used to complete precise pesticide delivery based on the disease, pest, and weed information of crops.
[0007] The multimodal data acquisition unit includes a hyperspectral imaging module, a high-resolution visible light imaging module, and a synchronization and positioning module. The hyperspectral imaging module uses a pushbroom hyperspectral camera with a spectral range of 400-1000nm and a spectral resolution of ≤5nm to capture fine spectral information of the crop canopy. The high-resolution visible light imaging module uses a global shutter RGB industrial camera to acquire visible light information of the crops, including color, texture, shape, and morphological structure information. The synchronization and positioning module is a GNSS / IMU module, including a GNSS / RTK receiver and an IMU inertial measurement unit, to acquire the positions of the pushbroom hyperspectral camera and the global shutter RGB industrial camera, and obtain the geographical location information of the crops through time alignment.
[0008] The edge computing and intelligent recognition unit is installed in the airborne Jetson device. The edge computing and intelligent recognition unit includes a data receiving unit, a data preprocessing unit, and a lightweight multimodal fusion recognition model. The data receiving unit is used to receive visible light information of crops, fine spectral information of the canopy, and geographical location information of crops through a wireless network. The data preprocessing unit is used to preprocess the received information data and input the preprocessed data into the lightweight multimodal fusion recognition model. The lightweight multimodal fusion recognition model compares the information features with the stored pest and weed database, identifies and outputs whether the crops in the image coordinate system are affected by pests and weeds. If there are no pests or weeds, it outputs healthy crops; if there are pests or weeds, it outputs information about the pests and weeds, including the affected area, individual pests, and weed communities.
[0009] The lightweight multimodal fusion recognition model includes a dual-branch feature extraction network, a cross-modal attention fusion module, and a task head and output module. The dual-branch feature extraction network extracts features from hyperspectral and visible light images respectively. The cross-modal attention fusion module is connected to the dual-branch feature extraction network and adaptively weights and fuses the features from the two branches using learnable dynamic weights. The task head and output module is connected to the cross-modal attention fusion module and performs semantic segmentation in parallel based on the fused features to output masks for diseased and weed areas, and performs target detection to output bounding boxes of individual pests. The lightweight multimodal fusion recognition model achieves real-time inference on edge computing devices by employing one or more lightweight techniques, such as depthwise separable convolution, channel pruning, and model quantization.
[0010] The precision decision-making and targeted application unit includes an intelligent decision-making system, a control system, and a targeted application execution system. The intelligent decision-making system is installed on an offline computer and is used to generate application prescriptions based on information about pests, diseases, and weeds affecting crops. The control system is used to control the targeted application execution system to perform precise spraying according to the application prescriptions. The targeted application execution system includes a mixing tank, a delivery pipeline, and a nozzle array connected in sequence. The nozzle array is installed around the crop to be monitored. The mixing tank is a three-compartment tank used to mix fungicides, insecticides, and herbicides separately. The nozzle array consists of multiple independently controlled electromagnetic nozzles.
[0011] The second technical solution adopted in this invention is a method for identifying and applying pesticides to crop diseases, pests, and weeds based on multimodal fusion, comprising the following steps: Step 1: The drone flies autonomously over the crops along a preset route. During the flight, the multimodal data acquisition unit on the drone simultaneously collects visible light information of the crops, fine spectral information of the canopy, and geographical location information. Step 2: The collected data is transmitted to the airborne Jetson device in real time. The airborne Jetson device preprocesses and performs spatiotemporal registration on the collected multimodal data. Then, the spatiotemporally registered image is input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The output is information on crop diseases, pests and weeds and corresponding geographical location information. Step 3: Generate a pesticide prescription based on the information of pests and weeds affecting the crops, then prepare the pesticide, and finally spray the corresponding crop area precisely based on the geographical location information of the affected crops.
[0012] In step 1, the multimodal data acquisition unit includes a pushbroom hyperspectral camera, a global shutter RGB industrial camera, and a synchronization and positioning module. The pushbroom hyperspectral camera and the global shutter RGB industrial camera are securely mounted under the drone fuselage via a shock-absorbing gimbal. The synchronization and positioning module is the drone's built-in GNSS / IMU module.
[0013] A lightweight multimodal fusion recognition model was trained using a crop disease, pest, and weed database. This database contains information on various crop types and common diseases, pests, and weeds affecting those crops. The information includes affected areas, individual pests, and weed communities. Affected areas include disease type and severity; individual pests include pest type and quantity; and weed communities include weed species and coverage. The database is a multimodal image database, including hyperspectral and visible light images. Crop disease areas were identified on the hyperspectral images. Pixel-level segmentation and annotation allows for the inference of disease type and severity. Boundary boxes are used to annotate pests on crops in the visible light image, and instance segmentation and annotation are used to annotate weeds around the crops. The pest type and number can be inferred from the pest bounding box annotation, and the weed species and weed coverage can be inferred from the weed annotation. The input of the lightweight multimodal fusion recognition model is a hyperspectral image and a visible light image containing crops, and the output is the information on diseases, pests and weeds affecting the crops, including disease areas, individual pests and weed communities.
[0014] In step 2, the spatiotemporally registered hyperspectral image and visible light image are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The specific process is as follows: Step 2.1, Feature Extraction: The hyperspectral image is input into the hyperspectral branch feature extraction network, which adopts a lightweight network based on 3D convolution to extract the spectral spatial joint feature F_hs containing physiological and biochemical information. The visible light image is input into the visible light branch feature extraction network, which adopts a lightweight 2DCNN backbone network to extract the color, texture, shape and morphological structure spatial features F_rgb. Step 2.2, Adaptive fusion: F_hs and F_rgb are input into the cross-modal attention fusion module, the dynamic weight α related to the current image content is calculated, and weighted fusion is performed according to the formula "F_fused=α·F_hs+(1-α)·F_rgb" to obtain the comprehensive feature F_fused; Step 2.3, Target Recognition and Output: F_fused is simultaneously input to the segmentation head and the detection head. The segmentation head uses a semantic segmentation network based on the UNet architecture to output pixel-level segmentation masks for diseased and weeded areas. The detection head uses a lightweight variant of the single-stage target detection algorithm based on the YOLO structure to output the bounding boxes of individual pests. Finally, the model outputs structured recognition results, including the type and region of the disease, the type and location of the pest, the species and range of the weeds, and their corresponding image coordinates.
[0015] In step 3, the application prescription includes the type and dosage of the pesticide. According to the application prescription, the control system uses a PID control algorithm to precisely drive the metering pump and the mixing solenoid valve group in the mixing tank to achieve precise mixing of the pesticide. Based on the geographical location information of the crops affected by pests, diseases and weeds, the control system activates the corresponding electromagnetic nozzles to achieve precise spraying in specific areas.
[0016] The beneficial effects of this invention are as follows: (1) Qualitative change in identification ability: It has achieved a leap from "single identification" to "three-in-one identification of disease-insect-weed", and from "late diagnosis" to "early warning". The introduction of hyperspectral imaging makes it possible to detect diseases 7-14 days before the crops show visible symptoms, which wins valuable time for prevention and control and reduces yield loss from the source. (2) Improved the precision of pesticide application: It has realized the transformation from "flood irrigation" to "precision drip irrigation". Through targeted application and on-demand pesticide application, it is expected to reduce the amount of pesticides used by more than 50%, and even more than 70% in some scenarios (such as local pests). This directly reduces production costs and reduces agricultural non-point source pollution and pesticide residue risks in agricultural products from the root. (3) Improved system efficiency: It realizes "one machine for multiple uses and one inspection for multiple checks". The diagnosis and treatment of three core plant protection tasks can be completed in one field inspection, which greatly improves the efficiency of operation and saves manpower and time costs; (4) The technology is highly practical and has broad prospects for implementation: The lightweight model designed for edge computing solves the problem of "difficulty in implementing complex AI models in agricultural fields". The whole system relies on a mature agricultural machinery / drone platform, so that high technology is no longer limited to the laboratory and has the practical feasibility of large-scale industrialization and promotion. Attached Figure Description
[0017] Figure 1 This is a framework diagram of the crop disease, pest and weed identification and application system based on multimodal fusion of the present invention; Figure 2 This is a flowchart illustrating the method for identifying and applying pesticides to crops based on multimodal fusion, as per the present invention. Detailed Implementation
[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0019] Example 1 See Figure 1A multimodal fusion-based system for identifying and applying pesticides to crops, including a multimodal data acquisition unit, an edge computing and intelligent recognition unit, and a precise decision-making and targeted pesticide application unit. The multimodal data acquisition unit is used to collect visible light information, fine spectral information of the canopy, and geographical location information of crops. The edge computing and intelligent recognition unit is used to preprocess the collected multimodal data, then perform cross-modal attention fusion, and compare the information features with a stored database of crops, pests, and weeds to determine the pests, diseases, and weeds in the image. The precise decision-making and targeted pesticide application unit is used to complete precise pesticide delivery based on the pest, disease, and weed information of crops.
[0020] The multimodal data acquisition unit includes a hyperspectral imaging module, a high-resolution visible light imaging module, and a synchronization and positioning module. The hyperspectral imaging module uses a pushbroom hyperspectral camera with a spectral range of 400-1000nm and a spectral resolution of ≤5nm to capture fine spectral information of the crop canopy. The high-resolution visible light imaging module uses a global shutter RGB industrial camera to acquire visible light information of the crops, including color, texture, shape, and morphological structure information. The synchronization and positioning module is a GNSS / IMU module, including a GNSS / RTK receiver and an IMU inertial measurement unit, to acquire the positions of the pushbroom hyperspectral camera and the global shutter RGB industrial camera, and obtain the geographical location information of the crops through time alignment.
[0021] Example 2 A crop disease, pest, and weed identification and application system based on multimodal fusion includes a multimodal data acquisition unit, an edge computing and intelligent recognition unit, and a precise decision-making and targeted application unit. The multimodal data acquisition unit is used to collect visible light information, fine spectral information of the canopy, and geographical location information of crops. The edge computing and intelligent recognition unit is used to preprocess the collected multimodal data, then perform cross-modal attention fusion, compare the information features with a stored disease, pest, and weed database, and determine the disease, pest, and weed information of crops in the image. The precise decision-making and targeted application unit is used to complete precise pesticide delivery based on the disease, pest, and weed information of crops.
[0022] The multimodal data acquisition unit includes a hyperspectral imaging module, a high-resolution visible light imaging module, and a synchronization and positioning module. The hyperspectral imaging module uses a pushbroom hyperspectral camera with a spectral range of 400-1000nm and a spectral resolution of ≤5nm. It is used to capture fine spectral information of the crop canopy, forming a three-dimensional image cube containing two spatial dimensions and one spectral dimension. This provides a physical basis for diagnosing early crop diseases and nutrient stress. The high-resolution visible light imaging module uses a global shutter RGB industrial camera with a resolution of no less than 12 megapixels. The system acquires visible light information about crops, including color, texture, shape, and morphological structure information, providing a foundation for accurately identifying individual crop pests, typical lesions, and weed communities. The synchronization and positioning module is a GNSS / IMU module, including a GNSS / RTK receiver and an IMU inertial measurement unit. The IMU measures the linear motion acceleration and angular velocity of the pushbroom hyperspectral camera and the global shutter RGB industrial camera, thereby calculating the platform's real-time attitude angles (pitch, roll, and yaw). Its main functions are as follows: On the one hand, there is image geometric correction. The attitude angle is used to correct image distortion caused by changes in the flight attitude of the UAV (such as shaking or tilting), ensuring that the hyperspectral image and the visible light image are aligned in spatial geometry, thus providing a basis for pixel-level fusion.
[0023] On the other hand, there is precise geographic mapping. By combining GNSS location information with camera installation parameters, the attitude angle can accurately convert the pixel coordinates of the targets (disease areas, pests, weeds) identified in the image into absolute geographical locations in the geodetic coordinate system. This is a prerequisite for driving nozzles at specific locations to perform targeted pesticide application.
[0024] Hardware synchronization triggers ensure that the two cameras are exposed at the same microsecond level. The GNSS / IMU receiver is used to receive and process satellite signals, thereby determining the position of the receiver and the pushbroom hyperspectral camera and global shutter RGB industrial camera that carry the receiver. The geographical location information of the crops is obtained through time alignment.
[0025] The edge computing and intelligent recognition unit is installed in the airborne Jetson device. The edge computing and intelligent recognition unit includes a data receiving unit, a data preprocessing unit, and a lightweight multimodal fusion recognition model. The data receiving unit is used to receive visible light information of crops, fine spectral information of the canopy, and geographical location information of crops through a wireless network. The data preprocessing unit is used to preprocess the received information data and input the preprocessed data into the lightweight multimodal fusion recognition model. The lightweight multimodal fusion recognition model compares the information features with the stored pest and weed database, identifies and outputs whether the crops in the image coordinate system are affected by pests and weeds. If there are no pests or weeds, it outputs healthy crops; if there are pests or weeds, it outputs information about the pests and weeds, including the affected area, individual pests, and weed communities.
[0026] The lightweight multimodal fusion recognition model includes a two-branch feature extraction network, a cross-modal attention fusion module, and a task head and output module. The two-branch feature extraction network extracts features from hyperspectral and visible light images respectively. The cross-modal attention fusion module is connected to the two-branch feature extraction network and is used to adaptively weight and fuse the features from the two branches using learnable dynamic weights. The task head and output module is connected to the cross-modal attention fusion module and is used to perform semantic segmentation in parallel based on the fused features to output masks for diseased and weed areas, and to perform target detection to output bounding boxes of individual pests. The lightweight multimodal fusion recognition model achieves real-time inference on edge computing devices by employing one or more lightweight techniques, such as depthwise separable convolution, channel pruning, and model quantization.
[0027] Example 3 See Figure 2 A multimodal fusion-based system for identifying and applying pesticides to crops, including a multimodal data acquisition unit, an edge computing and intelligent recognition unit, and a precise decision-making and targeted pesticide application unit. The multimodal data acquisition unit is used to collect visible light information, fine spectral information of the canopy, and geographical location information of crops. The edge computing and intelligent recognition unit is used to preprocess the collected multimodal data, then perform cross-modal attention fusion, and compare the information features with a stored database of crops, pests, and weeds to determine the pests, diseases, and weeds in the image. The precise decision-making and targeted pesticide application unit is used to complete precise pesticide delivery based on the pest, disease, and weed information of crops.
[0028] The multimodal data acquisition unit includes a hyperspectral imaging module, a high-resolution visible light imaging module, and a synchronization and positioning module. The hyperspectral imaging module uses a pushbroom hyperspectral camera with a spectral range of 400-1000nm and a spectral resolution of ≤5nm to capture fine spectral information of the crop canopy. The high-resolution visible light imaging module uses a global shutter RGB industrial camera to acquire visible light information of the crops, including color, texture, shape, and morphological structure information. The synchronization and positioning module is a GNSS / IMU module, including a GNSS / RTK receiver and an IMU inertial measurement unit. The IMU inertial measurement unit is used to measure the linear motion acceleration and angular velocity of the pushbroom hyperspectral camera and the global shutter RGB industrial camera, thereby calculating the platform's real-time attitude angles (pitch angle, roll angle, yaw angle) for image geometric correction and accurate geographic mapping. Image geometric correction ensures that the hyperspectral image and the visible light image are spatially geometrically aligned, providing a basis for pixel-level fusion. Accurate geographic mapping is a prerequisite for driving nozzles at specific locations to perform targeted pesticide application. GNSS / IMU receivers are used to receive and process satellite signals, thereby determining the positions of the receiver and the pushbroom hyperspectral camera and global shutter RGB industrial camera that carry the receiver, and obtaining the geographical location information of the crops captured by the images through time alignment.
[0029] The edge computing and intelligent recognition unit is installed in the airborne Jetson device. The edge computing and intelligent recognition unit includes a data receiving unit, a data preprocessing unit, and a lightweight multimodal fusion recognition model. The data receiving unit is used to receive visible light information of crops, fine spectral information of the canopy, and geographical location information of crops through a wireless network. The data preprocessing unit is used to preprocess the received information data and input the preprocessed data into the lightweight multimodal fusion recognition model. The lightweight multimodal fusion recognition model compares the information features with the stored pest and weed database, identifies and outputs whether the crops in the image coordinate system are affected by pests and weeds. If there are no pests or weeds, it outputs healthy crops; if there are pests or weeds, it outputs information about the pests and weeds, including the affected area, individual pests, and weed communities.
[0030] The lightweight multimodal fusion recognition model comprises a sequentially connected dual-branch feature extraction network, a cross-modal attention fusion module, and a task head and output module. The dual-branch feature extraction network includes a hyperspectral branch feature extraction network and a visible light branch feature extraction network. The hyperspectral branch feature extraction network employs a lightweight 3D convolution-based network to simultaneously extract joint features from the spatial neighborhood and spectral dimensions. The visible light branch feature extraction network uses a lightweight 2D... The CNNbackbone network's cross-modal attention fusion module receives feature maps F_hs from the hyperspectral branch feature extraction network and F_rgb from the visible light branch feature extraction network. It performs global average pooling on F_hs and F_rgb respectively, resulting in two global feature vectors V_hs and V_rgb. These V_hs and V_rgb are concatenated and input into a small neural network consisting of fully connected layers and a sigmoid activation function. This network outputs a dynamic fusion weight coefficient α, where 0 < α < 1. The weighted fusion is performed according to the formula F_fused = α·F_hs + (1-α)·F_rgb, yielding the fused feature map F_fused. The innovation of the cross-modal attention fusion module lies in the fact that the weight α is not a fixed value but is automatically learned and generated by the network based on the content of the current input image. For example, for regions with subtle spectral changes caused by early-stage diseases, the network learns a larger α, relying more on hyperspectral features; for morphologically specific pests, it learns a smaller α, relying more on visible light features. This enables optimal feature fusion that adapts to content.
[0031] The task head consists of a segmentation head and a detection head. The segmentation head employs a semantic segmentation network based on the UNet architecture. This head takes the fused feature map F_fused as input, performs upsampling and skip connections, and outputs a pixel-level semantic segmentation mask to accurately identify the contours of disease-infected areas and weed-covered areas. The detection head uses lightweight variants of single-stage object detection algorithms such as YOLOv5 or YOLOv8. This head takes F_fused as input and outputs bounding boxes containing class confidence scores and location coordinates to locate individual pests. The output module integrates the outputs of the segmentation and detection heads to form a structured recognition result.
[0032] In the dual-branch feature extraction network, the hyperspectral branch employs a lightweight network constructed using depthwise separable 3D convolutions, such as a simplified network based on HybridSN. Unlike 2D convolutions, which only extract spatial features, 3D convolutions can simultaneously extract joint features from both the spatial neighborhood and the spectral dimension, fully leveraging the inherent structural information of the "image-spectrum integration" of hyperspectral data and being more sensitive to subtle spectral differences. For deployment on mobile platforms, this invention significantly reduces the number of parameters and computational cost by using depthwise separable 3D convolutions. The visible light branch uses a lightweight 2D CNN backbone network, either MobileNetV3 or ShuffleNetV2, focusing on efficiently extracting texture, edge, shape, and contextual features of targets from RGB images. These features are crucial for distinguishing targets that may be spectrally similar but have different morphologies (such as weeds and crops). During the model training and deployment phase, channel pruning and model quantization techniques are comprehensively applied to further compress the model size and improve inference speed.
[0033] To achieve real-time inference on edge devices, this invention integrates several lightweight techniques throughout the network design and post-training processing: depthwise separable convolution (significantly reducing computational cost), channel pruning (removing redundant neural network channels), and model quantization (converting FP32 precision models to INT8 precision). These techniques ensure that the final deployed model size is kept below 10MB, and the inference time for a single frame image on the Jetson AGX Orin platform is less than 100ms, fully meeting the requirements for real-time operation on mobile field platforms.
[0034] The precision decision-making and targeted application unit includes an intelligent decision-making system, a control system, and a targeted application execution system. The intelligent decision-making system is installed on an offline computer and is used to generate application prescriptions based on information about pests, diseases, and weeds affecting crops. The control system is used to control the targeted application execution system to perform precise spraying based on the application prescriptions.
[0035] The targeted pesticide application system includes a mixing tank, a delivery pipeline, and a nozzle array connected in sequence. The nozzle array is installed around the crop to be monitored. The mixing tank is a three-compartment tank used to mix fungicides, insecticides, and herbicides separately. The nozzle array consists of multiple independently controlled electromagnetic nozzles.
[0036] Example 4 A method for identifying and applying pesticides to crops based on multimodal fusion includes the following steps: Step 1: The drone flies autonomously above the crops along a preset route. During the flight, the multimodal data acquisition unit on the drone simultaneously collects visible light information of the crops below, fine spectral information of the canopy, and geographical location information. The multimodal data acquisition unit includes a pushbroom hyperspectral camera, a global shutter RGB industrial camera, and a synchronization and positioning module. The pushbroom hyperspectral camera and the global shutter RGB industrial camera are securely mounted under the drone fuselage via a shock-absorbing gimbal, and the synchronization and positioning module is the drone's built-in GNSS / IMU module.
[0037] Step 2: The collected data is transmitted to the airborne Jetson device in real time. The airborne Jetson device preprocesses and performs spatiotemporal registration on the collected multimodal data. Then, the registered image pairs are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The output is information on crop diseases, pests and weeds and corresponding geographical location information. A lightweight multimodal fusion recognition model was trained using a crop disease, pest, and weed database. This database contains information on various crop types and common diseases, pests, and weeds affecting those crops. The information includes diseased areas, individual pests, and weed communities. Diseased areas include disease type and severity; individual pests include pest type and quantity; and weed communities include weed species and coverage. The database is a multimodal image database, including hyperspectral and visible light images. Hyperspectral images show pixel-level segmentation and annotation of crop diseased areas, allowing for the inference of disease type and severity. Visible light images show bounding box annotations of pests on crops and instance segmentation and annotation of weeds surrounding crops. Pest bounding box annotations allow for the inference of pest type and quantity, while weed annotations allow for the inference of weed species and coverage.
[0038] The lightweight multimodal fusion recognition model takes hyperspectral and visible light images of crops as input and outputs information on diseases, pests, and weeds affecting the crops, including diseased areas, individual pests, and weed communities.
[0039] Step 3: Generate a pesticide prescription based on the information of pests and weeds affecting the crops, then prepare the pesticide, and finally spray the corresponding crop area precisely based on the geographical location information of the affected crops.
[0040] Example 5 A method for identifying and applying pesticides to crops based on multimodal fusion includes the following steps: Step 1: The drone flies autonomously over the crops along a preset route. During the flight, the multimodal data acquisition unit on the drone simultaneously collects visible light information of the crops, fine spectral information of the canopy, and geographical location information. The multimodal data acquisition unit includes a pushbroom hyperspectral camera, a global shutter RGB industrial camera, and a synchronization and positioning module. The pushbroom hyperspectral camera and the global shutter RGB industrial camera are securely mounted under the drone fuselage via a shock-absorbing gimbal, and the synchronization and positioning module is the drone's built-in GNSS / IMU module.
[0041] Step 2: The collected data is transmitted to the airborne Jetson device in real time. The airborne Jetson device preprocesses and performs spatiotemporal registration on the collected multimodal data. Then, the registered image pairs are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The output is information on crop diseases, pests and weeds and corresponding geographical location information. A lightweight multimodal fusion recognition model was trained using a crop disease, pest, and weed database. This database contains information on various crop types and common diseases, pests, and weeds affecting those crops. The information includes diseased areas, individual pests, and weed communities. Diseased areas include disease type and severity; individual pests include pest type and quantity; and weed communities include weed species and coverage. The database is a multimodal image database, including hyperspectral and visible light images. Hyperspectral images show pixel-level segmentation and annotation of crop diseased areas, allowing for the inference of disease type and severity. Visible light images show bounding box annotations of pests on crops and instance segmentation and annotation of weeds surrounding crops. Pest bounding box annotations allow for the inference of pest type and quantity, while weed annotations allow for the inference of weed species and coverage.
[0042] In step 2, the registered image pairs are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The specific process is as follows: Step 2.1, Feature Extraction: The hyperspectral image is input into the hyperspectral branch feature extraction network, which adopts a lightweight network based on 3D convolution to extract the spectral spatial joint feature F_hs containing physiological and biochemical information. The visible light image is input into the visible light branch feature extraction network, which adopts a lightweight 2DCNN backbone network to extract the color, texture, shape and morphological structure spatial features F_rgb. Step 2.2, Adaptive fusion: F_hs and F_rgb are input into the cross-modal attention fusion module, the dynamic weight α related to the current image content is calculated, and weighted fusion is performed according to the formula "F_fused=α·F_hs+(1-α)·F_rgb" to obtain the comprehensive feature F_fused; Step 2.3, Target Recognition and Output: F_fused is simultaneously input to the segmentation head and the detection head. The segmentation head uses a semantic segmentation network based on the UNet architecture to output pixel-level segmentation masks for diseased and weeded areas. The detection head uses a lightweight variant of the single-stage target detection algorithm based on the YOLO structure to output the bounding boxes of individual pests. Finally, the model outputs structured recognition results, including the type and region of the disease, the type and location of the pest, the species and range of the weeds, and their corresponding image coordinates.
[0043] Step 3: Generate a pesticide prescription based on the information of pests, diseases and weeds on the crops, including the type and dosage of pesticides. The control system uses a PID control algorithm to precisely drive the metering pump and mixing solenoid valve group in the mixing tank according to the pesticide prescription, so as to achieve precise mixing of pesticides and complete the pesticide preparation. Then, according to the geographical location information of the crops with pests, diseases and weeds, start the electromagnetic nozzles at the corresponding locations to achieve precise spraying in specific areas. The decision logic table is shown in Table 1.
[0044] Table 1 Decision Logic Table
[0045] Example 6 A method for identifying and applying pesticides to rice diseases, pests, and weeds based on multimodal fusion includes the following steps: Step 1: Select the DJI T50 agricultural drone as the mobile carrier platform. The drone flies autonomously above the crops according to the preset route. During the flight, the multimodal data acquisition unit on the drone synchronously collects the visible light information of the crops, the fine spectral information of the canopy and the geographical location information at a frequency of 2Hz. The multimodal data acquisition unit includes a pushbroom hyperspectral camera (such as the Curt S185), a global shutter RGB industrial camera (such as the Sony IMX477), and a synchronization and positioning module. The pushbroom hyperspectral camera and the global shutter RGB industrial camera are securely mounted under the drone fuselage using a shock-absorbing gimbal. The synchronization and positioning module is the drone's built-in GNSS / IMU module. The edge computing device (NVIDIA Jetson AGX Orin) is placed inside the drone's cabin, with proper shockproof and heat dissipation measures.
[0046] Step 2: The collected data is transmitted to the airborne Jetson device in real time. The airborne Jetson device preprocesses and performs spatiotemporal registration on the collected multimodal data. Then, the registered image pairs are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The output is information on crop diseases, pests and weeds and corresponding geographical location information. A lightweight multimodal fusion recognition model was trained using a crop disease, pest, and weed database. This database contains information on various crop types and common diseases, pests, and weeds affecting those crops. The information includes diseased areas, individual pests, and weed communities. Diseased areas include disease type and severity; individual pests include pest type and quantity; and weed communities include weed species and coverage. The database is a multimodal image database, including hyperspectral and visible light images. Hyperspectral images show pixel-level segmentation and annotation of crop diseased areas, allowing for the inference of disease type and severity. Visible light images show bounding box annotations of pests on crops and instance segmentation and annotation of weeds surrounding crops. Pest bounding box annotations allow for the inference of pest type and quantity, while weed annotations allow for the inference of weed species and coverage.
[0047] In step 2, the registered image pairs are input into the pre-trained lightweight multimodal fusion recognition model for real-time inference. The specific process is as follows: Step 2.1, Feature Extraction: The hyperspectral image is input into the hyperspectral branch feature extraction network, which adopts a lightweight network based on 3D convolution to extract the spectral spatial joint feature F_hs containing physiological and biochemical information. The visible light image is input into the visible light branch feature extraction network, which adopts a lightweight 2DCNN backbone network to extract the color, texture, shape and morphological structure spatial features F_rgb. Step 2.2, Adaptive fusion: F_hs and F_rgb are input into the cross-modal attention fusion module, the dynamic weight α related to the current image content is calculated, and weighted fusion is performed according to the formula "F_fused=α·F_hs+(1-α)·F_rgb" to obtain the comprehensive feature F_fused; Step 2.3, Target Recognition and Output: F_fused is simultaneously input to the segmentation head and the detection head. The segmentation head uses a semantic segmentation network based on the UNet architecture to output pixel-level segmentation masks for diseased and weeded areas. The detection head uses a lightweight variant of the single-stage target detection algorithm based on the YOLO structure to output the bounding boxes of individual pests. Finally, the model outputs structured recognition results, including the type and region of the disease, the type and location of the pest, the species and range of the weeds, and their corresponding image coordinates.
[0048] In this embodiment, the areas of early-stage rice blast lesions, the locations of individual rice planthoppers, and the distribution areas of barnyard grass were identified.
[0049] Step 3: The decision-making system generates a pesticide prescription based on the information of pests and diseases affecting the crops, including the type and dosage of pesticides. The control system, based on the pesticide prescription, uses a PID control algorithm to precisely drive the metering pump and mixing solenoid valve group in the mixing tank to achieve precise mixing of the pesticides and complete the pesticide preparation. It also activates the corresponding electromagnetic nozzles based on the geographical location information of the crops affected by pests and diseases to achieve precise spraying in specific areas.
[0050] The mixing tank is a custom-designed three-compartment tank, containing fungicide (azoxystrobin), insecticide (imidacloprid), and herbicide (pentoxifyllium). The mixing tank is connected to a delivery pipeline and a nozzle array. The nozzle array consists of 16 independently controlled electromagnetic nozzles, arranged around the crop to be monitored. The control system directs the electromagnetic nozzles to spray "fungicide + low water volume" on diseased areas, "insecticide + very low water volume" on planthoppers, and "herbicide + high water volume" on barnyard grass. The nozzles are completely shut off on healthy areas. The control system responds within milliseconds, driving the corresponding solenoid valves and nozzles to perform precise actions.
Claims
1. A crop disease, pest and weed identification and pesticide application system based on multi-modal fusion, characterized in that, The application relates to a precision spraying system based on multi-modal data acquisition, edge computing and intelligent identification, and precise decision and target spraying.
2. The crop disease, pest and weed identification and pesticide application system based on multi-modal fusion of claim 1, characterized in that, The multi-modal data acquisition unit comprises a hyperspectral imaging module, a high-resolution visible light imaging module and a synchronization and positioning module, the hyperspectral imaging module adopts a push-broom hyperspectral camera, the spectral range is 400-1000 nm, the spectral resolution is less than or equal to 5 nm, and the hyperspectral imaging module is used for capturing fine spectral information of a crop canopy, the high-resolution visible light imaging module adopts a global shutter RGB industrial camera, and the high-resolution visible light imaging module is used for acquiring visible light information of the crops, including color, texture, shape and morphological structure information, and the synchronization and positioning module is a GNSS / IMU module, including a GNSS / RTK receiver and an IMU inertial measurement unit, and the synchronization and positioning module is used for acquiring the positions of the push-broom hyperspectral camera and the global shutter RGB industrial camera, and acquiring the geographical position information of the photographed crops through time alignment.
3. The crop disease, pest and weed identification and pesticide application system based on multi-modal fusion of claim 2, characterized in that, The edge computing and intelligent identification unit is installed in an airborne Jetson device, the edge computing and intelligent identification unit comprises a data receiving unit, a data preprocessing unit and a lightweight multi-modal fusion identification model, the data receiving unit is used for receiving the visible light information of the crops, the fine spectral information of the crop canopy and the geographical position information through a wireless network, the data preprocessing unit is used for preprocessing the received information data, and the preprocessed data is input into the lightweight multi-modal fusion identification model, the lightweight multi-modal fusion identification model compares the data with a stored pest and disease library, identifies and outputs whether the crops are diseased or not in an image coordinate system, if the crops are not diseased, healthy crops are output, and if the crops are diseased, pest and disease information is output, including a disease area, pest individuals and weed colonies.
4. The crop disease, pest and weed identification and pesticide application system based on multi-modal fusion of claim 3, characterized in that, The lightweight multi-modal fusion identification model comprises a double-branch feature extraction network, a cross-modal attention fusion module and a task head and output module, the double-branch feature extraction network is used for extracting features from hyperspectral images and visible light images respectively, the cross-modal attention fusion module is connected with the double-branch feature extraction network and is used for adaptively weighting and fusing the features from the double branches through learnable dynamic weights, the task head and output module is connected with the cross-modal attention fusion module and is used for performing semantic segmentation based on the fused features to output a mask of a disease area and a weed area and performing target detection to output a bounding box of pest individuals, and the lightweight multi-modal fusion identification model realizes real-time inference on an edge computing device by adopting one or more lightweight technologies, such as depth separable convolution, channel pruning and model quantization.
5. The crop disease, pest and weed identification and pesticide application system based on multi-modal fusion of claim 3, wherein, The precision decision and targeted drug application unit comprises an intelligent decision system, a control system and a targeted drug application execution system, the intelligent decision system is installed on an offline computer and is used for generating a drug application prescription according to disease, pest and weed information of crops, the control system is used for controlling the targeted drug application execution system to spray precision pesticide according to the drug application prescription, and the targeted drug application execution system comprises a mixing tank, a drug conveying pipeline and a nozzle array which are sequentially connected, the nozzle array is installed around the crops to be monitored, the mixing tank is a three-compartment tank and is used for mixing and loading fungicides, insecticides and herbicides respectively, and the nozzle array is composed of a plurality of independently controlled electromagnetic nozzles.
6. A crop disease, pest and weed identification and pesticide application method based on multi-modal fusion, characterized in that, The method comprises the following steps: Step 1: the unmanned aerial vehicle autonomously flies above the crops along a preset route, and a multi-modal data acquisition unit carried on the unmanned aerial vehicle synchronously acquires visible light information, fine spectral information of the crop canopy and geographical position information of the crops during the flight; Step 2: the acquired data is transmitted to the onboard Jetson device in real time, the onboard Jetson device pre-processes and spatio-temporally registers the acquired multi-modal data, and then inputs the spatio-temporally registered image pair into a trained lightweight multi-modal fusion recognition model to perform real-time inference, and outputs disease, pest and weed information of the crops and corresponding geographical position information; Step 3: a drug application prescription is generated according to the disease, pest and weed information of the crops, then the prescription is filled, and finally precision pesticide is sprayed on the area where the crops with diseases, pests and weeds are located according to the geographical position information of the crops with diseases, pests and weeds.
7. The crop disease, pest and weed identification and pesticide application method based on multi-modal fusion according to claim 6, characterized in that, In step 1, the multi-modal data acquisition unit comprises a push-broom hyperspectral camera, a global shutter RGB industrial camera and a synchronization and positioning module, the push-broom hyperspectral camera and the global shutter RGB industrial camera are stably installed below the unmanned aerial vehicle body through a shock-absorbing gimbal, and the synchronization and positioning module is a GNSS / IMU module provided on the unmanned aerial vehicle.
8. The crop disease, pest and weed identification and pesticide application method based on multi-modal fusion according to claim 6, characterized in that, The lightweight multi-modal fusion recognition model is trained by using a crop disease, pest and weed library, the crop disease, pest and weed library contains various crop types and common disease, pest and weed information of the crops, the disease, pest and weed information includes disease areas, pest individuals and weed communities, the disease areas include disease types and disease severities, the pest individuals include pest types and pest quantities, and the weed communities include weed species and weed coverage rates, the crop disease, pest and weed library is a multi-modal image database, including hyperspectral images and visible light images, the disease areas of the crops are pixel-level segmented and labeled on the hyperspectral images, through the labeling, the disease types and the disease severities can be inferred, the pests on the crops are boundary box labeled on the visible light images, and the weeds around the crops are instance segmentation labeled, through the pest boundary box labeling, the pest types and the pest quantities can be inferred, and through the weed labeling, the weed species and the weed coverage rates can be inferred, the input of the lightweight multi-modal fusion recognition model is the hyperspectral images and the visible light images containing the crops, and the output is the disease, pest and weed information of the crops, including the disease areas, the pest individuals and the weed communities.
9. The crop disease, pest and weed identification and pesticide application method based on multi-modal fusion according to claim 8, characterized in that, In step 2, the hyperspectral image after spatio-temporal registration and the visible light image pair are input into the trained lightweight multi-modal fusion recognition model for real-time inference, and the specific process is as follows: Step 2.1, feature extraction: the hyperspectral image is input into the hyperspectral branch feature extraction network, which adopts a lightweight network based on 3D convolution to extract spectral spatial joint features F_hs containing physiological and biochemical information; the visible light image is input into the visible light branch feature extraction network, which adopts a lightweight 2D CNN backbone network to extract color, texture, shape and morphological structure spatial features F_rgb; Step 2.2, adaptive fusion: F_hs and F_rgb are input into the cross-modal attention fusion module to calculate the dynamic weight α related to the current image content, and weighted fusion is performed according to the formula "F_fused=α·F_hs+(1-α)·F_rgb" to obtain the comprehensive feature F_fused; Step 2.3, target recognition and output: F_fused is input into the segmentation head and the detection head at the same time, the segmentation head adopts a semantic segmentation network based on UNet architecture to output pixel-level segmentation masks of disease areas and weed areas, and the detection head adopts a lightweight variant of a single-stage target detection algorithm based on YOLO structure to output the bounding boxes of pest individuals; finally, the model outputs structured recognition results, including the type and area of disease, the type and location of pest, the species and range of weed, and their corresponding image coordinates.
10. The crop disease, pest and weed identification and pesticide application method based on multi-modal fusion according to claim 9, characterized in that, In step 3, the pesticide prescription includes the type and dosage of pesticide, and the control system accurately drives the metering pump and the mixed electromagnetic valve group in the pesticide mixing tank according to the pesticide prescription through the PID control algorithm to realize accurate mixing of the pesticide, and starts the electromagnetic spray head at the corresponding position according to the geographical position information of the diseased crop to realize precise pesticide spraying in a specific area.