Early Diagnosis and Control System for Mango Diseases and Pests Based on Multimodal Sensing and AI

CN122556447APending Publication Date: 2026-08-14BAISE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]此外,芒果园中可见光传感器与高光谱传感器通常独立工作,各自按预设参数采集数据,二者之间缺乏有效的跨模态信息交互和校准机制

Benefits of technology

第一、本系统利用可见光图像中芒果叶片区域的光照强度直方图,结合高光谱传感器的饱和与噪声阈值,实时计算并调整积分时间,使光谱数据在光照波动下保持较高有效像素比例,降低因光照变化引起的误报。边缘计算单元运行轻量级AI推理模型对多模态数据进行初步筛查,仅将疑似病虫害数据提交云端精细诊断,缩短整体响应时间。

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses a mango pest and disease early diagnosis and control system based on multimodal sensing and AI, belonging to the field of smart agriculture and intelligent pest and disease monitoring technology. Addressing the problems of high false alarm rate and response delay in diagnosis under complex lighting conditions using a single sensor, this system simultaneously acquires visible light images, hyperspectral data, and environmental parameters through a multimodal sensing unit. An edge computing unit adjusts the integration time of the hyperspectral sensor in real time based on the light intensity histogram of the leaf area to reduce light interference. A lightweight AI inference model is run for initial pest and disease screening. Only suspected data is submitted to the cloud-based AI decision-making unit to run a CNN-LSTM fusion model for detailed diagnosis and to generate a control plan. The collaborative control execution unit automatically performs spraying and insect trapping operations. This system is mainly used for early detection and timely control of pests and diseases in large-scale mango orchards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart agriculture and intelligent pest and disease monitoring technology, specifically involving an early diagnosis and control system for mango pests and diseases based on multimodal sensing and AI. Background Technology

[0002] Mangoes are an important economic fruit tree in southern my country, and early detection and timely control of diseases and pests are crucial to ensuring mango yield and fruit quality. In large-scale mango cultivation, the initial symptoms of diseases and pests such as anthracnose, angular leaf spot, and thrips often manifest as subtle changes in leaf texture and local spectral abnormalities. These changes are difficult to detect with the naked eye in the early stages of the disease, and by the time the symptoms become obvious, the optimal window for control has been missed.

[0003] Current methods for monitoring mango diseases and pests mainly rely on regular field inspections by plant protection personnel combined with experience-based judgment, or on the use of a single type of sensor for auxiliary detection. When using visible light cameras for image acquisition and analysis, the lighting conditions in orchards are extremely complex, with light intensity, angle, and color temperature constantly changing throughout the day. The visible light characteristics of the same mango leaf differ significantly under direct light, diffused light, and shade. This leads to a high false alarm rate, as visual changes in leaves caused by drastic light fluctuations are easily misdiagnosed as disease or pest symptoms by visible light-based diagnostic systems. When using hyperspectral sensors for spectral diagnosis, hyperspectral data is highly sensitive to light intensity. Under conditions where the sensor integration time is fixed or adjusted only according to a preset program, the leaf reflectance spectrum under strong midday sunlight easily reaches the sensor saturation threshold, causing spectral information distortion and loss. Furthermore, insufficient light energy enters the sensor at dawn, dusk, or when there is heavy cloud cover, and the spectral signal is submerged in sensor dark noise, resulting in an extremely low signal-to-noise ratio, which also fails to provide effective diagnostic evidence. A single sensor cannot simultaneously suppress highlights and preserve shadows in complex lighting conditions, which is the physical reason for the high false alarm rate in diagnosis.

[0004] Regarding diagnostic response speed, existing methods relying on manual inspections are limited by manpower coverage and inspection cycles, often resulting in delays of several days or even longer from the actual occurrence of pests and diseases to their discovery. Even with automated sensor-based data collection solutions, if all collected data needs to be transmitted back to a remote server for centralized processing and analysis, a large amount of redundant data collected under normal, anomaly-free conditions consumes communication bandwidth and computing resources. Truly anomalous samples requiring immediate response are delayed in the transmission and processing queue, making it difficult to reduce the overall response time from data collection to obtaining diagnostic results. A key challenge in shortening diagnostic response delays is how to quickly screen data at the data collection end and submit only valuable, suspected data for in-depth analysis.

[0005] Furthermore, in mango orchards, visible light sensors and hyperspectral sensors typically operate independently, each collecting data according to preset parameters, lacking an effective cross-modal information exchange and calibration mechanism. When ambient light changes, the hyperspectral sensor cannot adaptively adjust its acquisition parameters to adapt to the current lighting conditions. The information on light changes reflected in the visible light images cannot be used to guide the parameter settings of the hyperspectral sensor, making it difficult to guarantee the radiometric consistency of multimodal data under different lighting conditions. This lack of cross-modal coordination further exacerbates the impact of changes in the lighting environment on diagnostic accuracy, and is a fundamental reason why the problem of single sensors being susceptible to ambient light interference still exists in multi-sensor systems. Summary of the Invention

[0006] One object of the present invention is to address at least the aforementioned deficiencies and to provide at least the advantages that will be described later.

[0007] This invention provides a mango pest and disease early diagnosis and control system based on multimodal sensing and AI. It can use the light intensity histogram of mango leaf area in visible light image to calculate and adjust the integration time of hyperspectral sensor in real time, so that the hyperspectral sensor maintains the effective spectral pixel ratio under light change conditions, reducing the impact of ambient light change on hyperspectral data. At the same time, a lightweight AI inference model is run through the edge computing unit to perform preliminary screening of multimodal sensing data, and only suspected pest and disease data is submitted to the cloud for detailed diagnosis, shortening the response time from data acquisition to obtaining diagnostic results.

[0008] This invention provides a mango pest and disease early diagnosis and control system based on multimodal sensing and AI, comprising: The multimodal sensing unit is deployed in the mango orchard. Each multimodal sensing unit includes a visible light camera with a resolution of no less than 2 million pixels, a hyperspectral sensor with a working wavelength in the range of 900-1700nm, an air temperature and humidity sensor, and a soil moisture sensor. The edge computing unit communicates with the multimodal sensing unit, receives visible light images, hyperspectral data, air temperature and humidity data, and soil moisture data collected by the multimodal sensing unit, performs noise reduction and normalization preprocessing on the visible light images and hyperspectral data, runs a lightweight AI inference model to perform preliminary screening of pests and diseases on the preprocessed data, and calculates and adjusts the integration time of the hyperspectral sensor in real time based on the light intensity histogram of the mango leaf area in the visible light image with the goal of maximizing the number of effective spectral pixels. The cloud-based AI decision-making unit connects to the edge computing unit via the network, receives visible light images and hyperspectral data that are initially screened by the edge computing unit as suspected pests and diseases, as well as air temperature and humidity data and soil moisture data at the corresponding collection time, and runs a CNN-LSTM fusion model to perform a detailed diagnosis of the type and severity of pests and diseases, generating a control plan that includes the amount of pesticide to be sprayed, the spraying range, and the start and stop instructions for the insect-attracting lamps. In the CNN-LSTM fusion model, CNN is used to extract spatial features, and LSTM is used to extract temporal features. The collaborative prevention and control execution unit, including intelligent sprayers, insect-attracting lamps, and dispatchable drones, is connected to the cloud-based AI decision-making unit to receive prevention and control plans and automatically execute spraying and insect-attracting operations.

[0009] The mango pest and disease early diagnosis and control system based on multimodal sensing and AI of this invention, during operation, uses a multimodal sensing unit deployed in the mango orchard to simultaneously collect visible light images, hyperspectral data, and orchard environmental parameters of mango leaves through a visible light camera, a hyperspectral sensor, an air temperature and humidity sensor, and a soil moisture sensor. After receiving the multimodal data, the edge computing unit first extracts the mango leaf area from the visible light image and calculates the light intensity histogram of the area. Combining the saturation radiance threshold and dark noise equivalent radiance threshold of the hyperspectral sensor, the optimal integration time is calculated in real time with the goal of maximizing the number of effective spectral pixels in the leaf area, and written to the hyperspectral sensor. This allows the hyperspectral sensor to dynamically adjust its operation when the light intensity fluctuates. The exposure parameters are adjusted to maintain the radiometric consistency of the spectral data acquisition. The edge computing unit simultaneously runs a lightweight AI inference model to perform preliminary screening of pests and diseases on the preprocessed visible light images and hyperspectral data. Data identified as suspected pests and diseases, along with the environmental parameters at the corresponding acquisition time, are uploaded to the cloud AI decision unit. The cloud AI decision unit runs a CNN-LSTM fusion model to perform a detailed diagnosis of the type and severity of pests and diseases on the received data. The CNN branch extracts the spatial features of the visible light images and the spectral features of the hyperspectral data, while the LSTM branch fuses the temporal features from multiple acquisition times. The diagnostic results are then converted into a control plan, including the amount of pesticide sprayed, the spraying range, and the start and stop instructions for the insect-attracting lamps, by the collaborative control execution unit and executed automatically.

[0010] Preferably, the edge computing unit performs denoising and normalization preprocessing on the visible light image and hyperspectral data, specifically including: A hierarchical preprocessing strategy is constructed. First, the fluctuation level of the data collected by the current multimodal sensing unit is evaluated. The fluctuation level is quantified by the mean of the noise standard deviation of each band of the hyperspectral sensor in n consecutive acquisition cycles and the rate of change of pixel values ​​between visible light image frames. When the fluctuation level is lower than the preset threshold, the first processing mode is adopted: the visible light image is subjected to edge-preserving denoising using guided filtering, the hyperspectral data is subjected to denoising method based on minimum noise separation transformation, the denoised hyperspectral data is radiometrically calibrated and corrected according to the current integration time of the hyperspectral sensor, the denoised hyperspectral data is normalized to the [0,1] interval, and the denoised visible light image is normalized to the [0,1] interval after histogram equalization. When the fluctuation level is equal to or higher than the preset threshold, it indicates that there is data quality degradation caused by environmental disturbances. The system then switches to the second processing mode: the visible light image is processed with dehazing enhancement based on dark channel priors and then guided filtering is performed; the hyperspectral data is denoised using a multivariate scattering correction-based method to eliminate spectral baseline drift and scattering noise caused by water vapor and dust scattering on the leaf surface; the denoised hyperspectral data is radiometrically calibrated according to the current integration time of the hyperspectral sensor; and the processed hyperspectral data and visible light image data are then jointly registered in spectral and spatial order before normalization to correct the intermodal spatial shift caused by environmental factors. All hyperspectral data preprocessing is synchronously associated with the optimal integration time parameter of the hyperspectral sensor corresponding to the current acquisition cycle to complete radiometric consistency normalization correction.

[0011] The edge computing unit employs a layered preprocessing strategy for the received visible light images and hyperspectral data. First, the mean of the noise standard deviation of each band within multiple consecutive acquisition cycles of the hyperspectral sensor, along with the rate of change of pixel values ​​between adjacent frames of the visible light image, is used to quantify the fluctuation level of the current data. When the fluctuation level is below a preset threshold, the first processing mode is adopted: the visible light image is denoised using guided filtering with edge preservation, the hyperspectral data is denoised using minimum noise separation transform, and then radiometric calibration correction is performed on the denoised hyperspectral data based on the current effective integration time of the hyperspectral sensor. The calibrated hyperspectral data is normalized to the [0,1] interval, and the denoised visible light image is normalized to the [0,1] interval after histogram equalization. When the fluctuation level is equal to or higher than a preset threshold, the system switches to the second processing mode: For the visible light image, a dehazing enhancement process based on dark channel priors is first applied, followed by guided filtering. For the hyperspectral data, a multivariate scattering correction method is used to eliminate spectral baseline drift and scattering noise caused by water vapor and dust scattering from the leaf surface. Similarly, radiometric calibration correction is performed based on the current integration time. Then, the processed hyperspectral data and visible light image data are jointly registered spectrally and spatially before normalization to correct spatial position shifts in the two modalities caused by environmental factors. In all modes, the preprocessing of hyperspectral data is synchronously associated with the optimal integration time parameter of the current acquisition cycle to complete radiometric consistency normalization correction.

[0012] Preferably, the edge computing unit runs a lightweight AI inference model to perform preliminary screening for pests and diseases on the preprocessed data, specifically including: A dual-branch lightweight feature extraction network and an environment coding module are deployed within the edge computing unit; The preprocessed visible light image is input into the first branch with MobileNetV3 as the skeleton to extract the visible light spatial feature map. The preprocessed hyperspectral data is input into the second branch with a one-dimensional convolutional network as the skeleton to extract the spectral feature vector. At the same time, the environmental coding module maps the air temperature and humidity data and soil moisture data in the current acquisition period into environmental feature vectors. The visible light spatial feature map is concatenated with the spectral feature vector obtained by global average pooling and the environmental feature vector. The concatenation is then input into a classification head consisting of two fully connected layers. The classification head outputs two scores in parallel: a suspected pest and disease score and an environmental stress confidence score. The edge computing unit dynamically adjusts the screening threshold based on the environmental stress confidence score: when the environmental stress confidence score is greater than the preset first threshold, the screening threshold for suspected pests and diseases is increased from the default value T0 to T1, where T1 is greater than T0, in order to suppress false positives in visible light or spectral representations caused by environmental stress from being misjudged as pests and diseases; when the environmental stress confidence score is less than or equal to the first threshold, the default value T0 is maintained as the screening threshold. The suspected pest score is compared with the currently effective screening threshold. If the suspected pest score is greater than the screening threshold, it is determined to be a suspected pest. The suspected pest score, environmental stress confidence score, spectral angular distance, and the corresponding preprocessed visible light image and hyperspectral data are then reported to the cloud. Simultaneously, a normal mango leaf hyperspectral reference library is established at the edge. The spectral angular distance between the current preprocessed hyperspectral data and the corresponding normal samples in the reference library is calculated in the characteristic bands of the 900-1000nm and 1400-1700nm ranges. When the spectral angular distance is higher than the preset second threshold and the environmental stress confidence score is lower than the preset third threshold, it is directly judged as a suspected pest or disease. The suspected pest or disease score, environmental stress confidence score, spectral angular distance, and the corresponding preprocessed visible light image and hyperspectral data are all reported to the cloud, regardless of whether the suspected pest or disease score exceeds the screening threshold.

[0013] A dual-branch lightweight feature extraction network and an environmental coding module are deployed within the edge computing unit. Preprocessed visible light images and hyperspectral data are processed by a MobileNetV3 skeleton and a convolutional network to extract spatial feature maps and spectral feature vectors, respectively. The environmental coding module maps air temperature and humidity data and soil moisture data into environmental feature vectors. After concatenation of the three types of features, the classification head outputs a suspected pest / disease score and an environmental stress confidence score in parallel. When the environmental stress confidence score exceeds a first threshold, the screening threshold is increased from its default value to suppress false positives caused by environmental stress; otherwise, the default threshold is maintained. If the suspected score exceeds the current screening threshold, it is classified as a suspected pest / disease and reported to the cloud. Simultaneously, a hyperspectral reference library of normal mango leaves is maintained at the edge. The spectral angular distance between the current spectrum and normal samples in the feature bands is calculated. When the spectral angular distance exceeds a second threshold and the environmental stress confidence score is lower than a third threshold, it is directly classified as a suspected pest / disease and reported, regardless of whether the suspected score meets the criteria.

[0014] Preferably, the edge computing unit adjusts the integration time of the hyperspectral sensor in real time according to the illumination intensity of the visible light image, specifically including: The edge computing unit uses a lightweight semantic segmentation model deployed on it to extract the mango leaf region from the visible light image, filter out the background and strongly reflective non-leaf regions, and calculate the illumination intensity histogram of all pixels in the mango leaf region. The saturation radiance threshold and dark noise equivalent radiance threshold of the hyperspectral sensor are obtained. Based on the illumination intensity histogram, the integration time that maximizes the proportion of effective pixels in the mango leaf region that is neither saturated nor higher than the dark noise equivalent radiance under the current gain of the hyperspectral sensor is calculated and used as the optimal integration time for the current frame. The edge computing unit also maintains an integration time series window of length k, records the optimal integration time calculated for k consecutive frames, obtains the slope of the light intensity change trend by linear fitting the integration time series, and uses the result of linear fitting to predict the estimated optimal integration time for the next acquisition moment. The estimated optimal integration time is weighted and fused with the optimal integration time of the current frame to obtain the final transmission integration time. The weighting coefficient is dynamically determined based on the goodness of fit of the linear fit. The higher the goodness of fit, the greater the weight of the estimated optimal integration time. The final transmission integration time is written to the hyperspectral sensor before the end of the current acquisition cycle so that it can take effect in the next acquisition cycle.

[0015] The edge computing unit utilizes a lightweight semantic segmentation model deployed on it to extract the mango leaf region from the visible light image, filtering out the background and strongly reflective non-leaf regions, and statistically analyzing the illumination intensity histogram of all pixels within the leaf region. Combining the saturation radiance threshold and the dark noise equivalent radiance threshold of the hyperspectral sensor, it calculates the integration time that maximizes the proportion of effective pixels within the leaf region that are neither saturated nor higher than the dark noise equivalent radiance under the current gain, which is taken as the optimal integration time for the current frame. Simultaneously, a sequence window recording the optimal integration times for multiple consecutive frames is maintained. Linear fitting is performed on the integration time sequence within the window to obtain the linear fitting result, predicting the estimated optimal integration time for the next acquisition moment. The estimated optimal integration time is then weighted and fused with the optimal integration time of the current frame. The weighting coefficient is dynamically determined based on the goodness of fit of the linear fitting; the higher the goodness of fit, the greater the weight of the estimated value. The final integration time is then sent and written to the hyperspectral sensor before the end of the current acquisition cycle for use in the next cycle.

[0016] Preferably, the cloud-based AI decision-making unit runs a CNN-LSTM fusion model to perform detailed diagnosis of the type and severity of pests and diseases, specifically including: The CNN-LSTM fusion model includes a cross-modal multi-head attention fusion module, an environmental condition gating module, and a temporal multi-task output module; The CNN-LSTM fusion model also receives the suspected pest and disease scores, environmental stress confidence scores, and spectral angular distance reported by the edge computing units. It concatenates the suspected pest and disease scores and environmental stress confidence scores into an edge prior feature vector, uses the edge prior feature vector as the initial hidden state bias of the temporal encoder, and concatenates the spectral angular distance as an additional input dimension into the multi-level spectral feature vector. The visible light image reported by the edge computing unit is input into the pre-trained ResNet-18 backbone network to extract multi-scale spatial feature maps. The corresponding hyperspectral data is input into a spectral encoder consisting of three consecutive one-dimensional convolutional blocks to extract multi-level spectral feature vectors. Each convolutional block of the spectral encoder outputs feature representations with different spectral resolutions. The cross-modal multi-head attention fusion module uses the highest-level spatial feature map in the multi-scale spatial feature map as the query and the multi-level spectral feature vector as the key and value to calculate the cross-modal attention weight between spatial location and spectral band, and generate a spatial-spectral joint enhanced feature map, so that the model focuses on the pixel position in the visible light image that corresponds to the spatial location of the spectral anomaly region. The spatial-spectral joint enhanced feature map is then concatenated with the multi-level spectral feature vector at the channel level after global average pooling to obtain the multimodal fusion feature vector at the current time. The environmental condition gating module receives air temperature, humidity and soil moisture data reported by the edge computing unit at the current acquisition time. After mapping by the fully connected layer, it generates a set of environmental gating coefficients. The environmental gating coefficients are used to scale each dimension of the multimodal fusion feature vector element by element to suppress feature responses that are highly correlated with the current environmental conditions but unrelated to pests and diseases. The multimodal fusion feature vector after environmental gating is fed into a time encoder composed of two layers of LSTM. The two layers of LSTM maintain a hidden state across time steps, fuse the feature vector at the current time step with the historical hidden state, and output a time-aware feature vector. The temporal multi-task output module includes a pest and disease type classification branch and a severity regression branch. The classification branch takes the temporal-aware feature vector as input and outputs the probability of each pest and disease type. The severity regression branch takes the concatenation result of the temporal-aware feature vector and the embedding vector of the highest probability type of the classification branch as input and outputs the severity index corresponding to the type. The severity index is a weighted combination of the lesion area ratio and the normalized spectral index.

[0017] The CNN-LSTM fusion model running in the cloud-based AI decision-making unit includes a cross-modal multi-head attention fusion module, an environmental condition gating module, and a temporal multi-task output module. The model receives suspected pest and disease scores, environmental stress confidence scores, and spectral angular distances reported by the edge computing unit. The first two are concatenated into an edge prior feature vector, which serves as the initial hidden state bias for the temporal encoder. Simultaneously, the spectral angular distance is concatenated into the multi-level spectral feature vectors. The reported visible light images are processed by a pre-trained ResNet-18 backbone network to extract multi-scale spatial feature maps, while hyperspectral data is processed by a spectral encoder consisting of three consecutive one-dimensional convolutional blocks to extract multi-level spectral feature vectors. The cross-modal multi-head attention fusion module uses the highest-level spatial feature map as the query and the multi-level spectral feature vectors as the key and value, calculating cross-modal attention weights to generate a spatial-spectral joint enhanced feature map. After global average pooling, this map is concatenated with the multi-level spectral feature vectors at the channel level to obtain the multi-modal fusion feature vector. The environmental condition gating module receives air temperature, humidity, and soil moisture data at the current acquisition time. It generates environmental gating coefficients through a fully connected layer, scaling each dimension of the multimodal fusion feature vector element-wise to suppress environmental responses unrelated to pests and diseases. The gated features are then fed into a two-layer LSTM temporal encoder, which fuses the current-time features with historical hidden states to output a temporal-aware feature vector. The classification branch of the temporal multi-task output module outputs the probability of each pest and disease category, while the severity regression branch takes the concatenation of the temporal-aware feature vector and the embedding vector of the highest-probability category as input, outputting a severity index—a weighted combination of the lesion area ratio and the normalized spectral index.

[0018] Preferably, the spatial-spectral joint enhanced feature map is concatenated with multi-level spectral feature vectors at the channel level after global average pooling, specifically including: Before global average pooling, the spatial-spectral joint enhanced feature map is divided into g equally sized feature map groups along the channel dimension. The spatial attention weighted average within each feature map group is performed to obtain a g-dimensional group spatial convergence vector. The group spatial convergence vector is then concatenated with the global representation vector obtained after global average pooling of the spatial-spectral joint enhanced feature map to form a spatial feature descriptor with multi-granularity spatial information. Normalized spectral vegetation index is calculated for each level in the multi-level spectral feature vector. The normalized spectral vegetation index is the normalized ratio of the reflectance of the feature band sensitive to chlorophyll to the reflectance of the feature band sensitive to cell structure. The normalized spectral vegetation indices of each level are spliced ​​together to obtain the spectral index feature vector. When concatenating the spatial feature descriptor with the multi-level spectral feature vector at the channel level, the spectral index feature vector is also concatenated to form the multimodal fusion feature vector at the current time. Before concatenation, the multi-level spectral feature vector undergoes adaptive filtering through a learnable channel gate. The channel gate takes the spatial feature descriptor as a conditional input and outputs the retention probability of each channel of the multi-level spectral feature vector. Only channels with retention probabilities higher than the preset gate threshold are concatenated with the spatial feature descriptor and the spectral index feature vector.

[0019] Before global average pooling, the spatial-spectral joint enhanced feature map is divided into multiple feature map groups of equal size along the channel dimension. Within each group, a spatial attention-weighted average is performed to obtain a group spatial convergence vector. This convergence vector is then concatenated with the global representation vector obtained from global average pooling to form a spatial feature descriptor with multi-granular spatial information. For each level of the multi-level spectral feature vector, a normalized spectral vegetation index is calculated. This index is the normalized ratio of the reflectance of chlorophyll-sensitive feature bands to the reflectance of cell structure-sensitive feature bands. The indices at each level are concatenated to obtain a spectral index feature vector. Before concatenation, the multi-level spectral feature vector undergoes adaptive filtering using a learnable channel gating system with the spatial feature descriptor as input. The system outputs the retention probability of each channel, concatenating only channels with a preset gating threshold with the spatial feature descriptor and the spectral index feature vector to form the multimodal fusion feature vector for the current time step.

[0020] Preferably, the environmental condition gating module receives air temperature and humidity and soil moisture data reported by the edge computing unit at the current acquisition time, and generates a set of environmental gating coefficients after mapping by a fully connected layer, specifically including: The environmental condition gating module also receives the air temperature and humidity sequence and soil moisture sequence of m consecutive collection times before the current collection time reported by the edge computing unit. Based on the air temperature and humidity sequence, it calculates the duration of leaf wetness and effective accumulated temperature within a preset time. Based on the soil moisture sequence, it calculates the soil moisture deficit index. It then concatenates the duration of leaf wetness, effective accumulated temperature, and soil moisture deficit index with the air temperature and humidity and soil moisture data at the current collection time, and integrates them with the environmental stress confidence score reported by the edge computing unit. The results are then input into the fully connected layer to generate the initial environmental gating coefficient. In the environmental condition gating module, a learnable pest-environment association matrix is ​​pre-set. Each row of the pest-environment association matrix corresponds to the environmental parameter response pattern embedding of a candidate mango pest type under typical occurrence conditions. The initial environmental gating coefficient is used as a query and scaled dot product attention is calculated with the pest-environment association matrix to obtain attention weights that are softly aligned with each candidate pest type. The attention weights are then used to perform weighted summation on each row of the pest-environment association matrix to generate an environmental gating bias vector that is aware of pest conditions. The initial environmental gating coefficient is added element-wise to the environmental gating bias vector of the pest and disease condition perception to obtain the final environmental gating coefficient. The final environmental gating coefficient is then used to scale each dimension of the multimodal fusion feature vector element-wise.

[0021] The environmental condition gating module receives air temperature and humidity sequences and soil moisture sequences from multiple consecutive time points prior to the current acquisition time. Based on this, it calculates the duration of leaf moisture, effective accumulated temperature, and soil moisture deficit index within a preset time period. This data is then concatenated with the current air temperature and humidity and soil moisture data, and further fused with the environmental stress confidence score reported by the edge computing unit. This data is then input into the fully connected layer to generate initial environmental gating coefficients. The module pre-defines a learnable pest-environment association matrix, with each row corresponding to the environmental parameter response pattern embedding of a candidate pest type under typical occurrence conditions. Using the initial environmental gating coefficients as the query, a scaled dot product attention calculation is performed with this matrix to obtain attention weights softly aligned with each candidate pest type. The matrix is ​​then weighted and summed to generate an environmental gating bias vector for pest condition awareness. The initial environmental gating coefficients are then element-wise added to this bias vector to obtain the final environmental gating coefficients, which are used to scale each dimension of the multimodal fusion feature vector element-wise.

[0022] Preferably, the cloud-based AI decision-making unit generates a control plan that includes spraying dosage, spraying range, and start / stop instructions for the insect-attracting lamps, specifically including: After obtaining the types and severity of pests and diseases, the cloud-based AI decision-making unit acquires the IoT device status data of each smart sprayer and insect-attracting lamp in the collaborative prevention and control execution unit. The IoT device status data includes at least the remaining liquid, battery level, current location coordinates, and nozzle type of each smart sprayer, as well as the current working status, light intensity level, and cumulative working time of each insect-attracting lamp. A two-layer bipartite graph is constructed, consisting of an equipment capability layer and a prevention and control requirement layer. The equipment capability layer uses each available intelligent sprayer, insect-attracting lamp, and schedulable drone as nodes, while the prevention and control requirement layer uses the occurrence locations of each pest and disease determined by detailed diagnosis and the required prevention and control actions as nodes. The required prevention and control actions include the type of spraying, the amount of spraying required based on the pest and disease severity index, the start and stop requirements of the insect-attracting lamp, and the requirement of drone supplementary prevention. The edge weights of the two-layer bipartite graph are determined by the weighted sum of the spatial distance between the equipment and the occurrence location, the matching degree of the current status of the equipment, and the urgency of prevention and control. The minimum cost matching algorithm with capacity constraints is run on a bipartite graph to minimize the total scheduling cost. Under the constraints of the upper limit of the spraying capacity of each intelligent sprayer and the effective coverage radius of each insect-attracting lamp, the algorithm solves the target spraying range, travel path and spraying amount allocation of each intelligent sprayer, as well as the start and stop commands and target light intensity levels of each insect-attracting lamp. When multiple pest and disease locations compete for the same intelligent sprayer and the remaining liquid in the sprayer cannot simultaneously meet the spraying needs of all locations, a prevention and control priority index is calculated based on the severity of the pests and diseases and the risk of spread at each location. The liquid is then allocated to the top k locations with the highest prevention and control priority index. For the remaining locations, alternative sprayers are rescheduled or they are marked as areas requiring supplementary prevention and a supplementary prevention reminder is generated and pushed to the management personnel.

[0023] After obtaining the types and severity of pests and diseases, the cloud-based AI decision-making unit acquires IoT device status data for each smart sprayer and insect-attracting lamp, including pesticide residue, battery level, current location coordinates, nozzle type, and the current operating status, light intensity level, and cumulative operating time of the insect-attracting lamp. A two-layer bipartite graph is constructed, with available devices as device capability layer nodes, and the locations of each pest and disease occurrence and their required pesticide dosage, as well as the start / stop requirements of the insect-attracting lamp, as control demand layer nodes. The edge weights are determined by a weighted sum of the spatial distance between the device and the occurrence location, the device status matching degree, and the urgency of control. A minimum cost matching algorithm with capacity constraints is run on this bipartite graph to solve for the target spraying range, travel path, and pesticide dosage allocation for each sprayer, as well as the start / stop commands and target light intensity level for each insect-attracting lamp, while satisfying the constraints of the sprayer capacity limit and the effective coverage radius of the insect-attracting lamp. When multiple locations compete for the same sprayer and the remaining pesticide solution is insufficient, the locations are sorted according to the prevention and control priority index calculated based on the severity of pests and diseases and the risk of spread. Priority is given to high-priority locations, and the remaining locations are rescheduled or marked as areas to be covered and a reminder for additional treatment is generated.

[0024] Preferably, when the cloud-based AI decision-making unit solves for the target spraying range, travel path, and spraying quantity allocation of each intelligent sprayer, it also includes: Receive orchard micro-meteorological data collected in real time by a multimodal sensing unit or an external micro-weather station. The micro-meteorological data includes at least wind speed, wind direction, and rainfall probability. Micrometeorological data is input into a preset spraying operation safety window determination model. When the wind speed exceeds the preset safety threshold or the probability of rainfall is higher than the preset rainfall threshold, the current moment is determined as an unsprayable window, and the affected spraying tasks are suspended. At the same time, based on the window prediction of meteorological forecast data, the earliest executable time and effective spraying time window of the suspended tasks are calculated. If the predicted effective spraying time window exceeds the upper limit of the allowable delay for pest and disease control, alternative control strategies are triggered, including increasing the insect-attracting light intensity level of the corresponding area of ​​the insect-attracting lamp or dispatching drones for emergency supplementary control. When multiple executable spraying tasks overlap in time and space, a distributed temporal path coordination graph is constructed with each sprayer as an agent. In the coordination graph, each agent exchanges spatiotemporal occupancy information and detects path intersection conflicts based on the currently allocated target spraying range and travel path. The spatiotemporal priority negotiation mechanism determines the passage sequence of each sprayer at the intersection point. The spatiotemporal priority negotiation mechanism uses the severity of pests and diseases and the remaining power of the sprayer as joint priority factors and outputs a conflict-free temporal execution plan. The temporal execution plan includes the predetermined passage time and avoidance waiting point for each sprayer on each ridge.

[0025] When the cloud-based AI decision-making unit calculates the spraying range, travel path, and pesticide distribution for each sprayer, it also receives real-time orchard micro-meteorological data, such as wind speed, wind direction, and rainfall probability, collected by multimodal sensing units or external micro-weather stations. This data is then input into the spraying operation safety window determination model. When the wind speed or rainfall probability exceeds a preset safety threshold, the current window is determined to be unsuitable for spraying, and the affected spraying task is suspended. Based on weather forecast data, the earliest executable time and effective spraying window for the suspended task are calculated. If the predicted effective window exceeds the allowable delay limit for pest and disease control, an alternative control strategy is triggered, such as increasing the intensity of insect-attracting lights in the corresponding area or dispatching drones for emergency supplementary control. When multiple executable spraying tasks overlap in time and space, a distributed temporal path coordination graph is constructed with each sprayer as an agent. Each agent exchanges spatiotemporal occupancy information and detects path intersection conflicts. The spatiotemporal priority is negotiated with the severity of pests and diseases and the remaining power of the sprayer as joint priority factors to determine the passage sequence of each sprayer at the intersection point. The output is a conflict-free temporal execution plan that includes the predetermined passage time and avoidance waiting point of each sprayer on each ridge.

[0026] Preferably, it also includes a data self-learning feedback unit, which collects data on pest and disease control rate and pesticide damage after the collaborative prevention and control execution unit executes the prevention and control plan, and feeds it back to the cloud AI decision-making unit; the cloud AI decision-making unit uses the received feedback data to update the CNN-LSTM fusion model.

[0027] Preferably, the data self-learning feedback unit collects data on pest and disease control rates and pesticide damage after the collaborative prevention and control execution unit implements the prevention and control plan, and feeds this data back to the cloud-based AI decision-making unit, specifically including: The data self-learning feedback unit receives visible light images and hyperspectral data re-collected by the multimodal sensing unit within a preset post-control evaluation period after the collaborative control execution unit performs spraying and insect trapping operations, as well as the post-control pest and disease type and severity index output by the cloud AI decision unit after running the CNN-LSTM fusion model on the re-collected data. The severity index of pests and diseases output by the cloud-based AI decision-making unit before prevention and control is denoted as S. pre The severity index of pests and diseases output by the CNN-LSTM fusion model at the same location during the post-control assessment period is denoted as S. post Calculate the pest and disease control rate CR, when S pre When S = 0, CR = 1; when S = 0, CR = 1. pre When >0, CR = max(0, (S pre -S post ) / S pre ); Meanwhile, the visible light images re-acquired after the control measures were input into the pre-trained semantic segmentation model to extract the mango leaf region. Within this region, three pesticide damage characterization indicators—the ratio of yellowing area, the ratio of necrotic spot area, and the leaf curling index—were calculated. The pesticide damage index PI was obtained by weighted summation of the three pesticide damage characterization indicators. The data self-learning feedback unit associates the pest control rate (CR) and pesticide damage index (PI) at each location of pest occurrence with the corresponding control plan's spraying type, spraying amount, and pest type, generating structured feedback data, which is then uploaded to the cloud-based AI decision-making unit. After receiving structured feedback data, the cloud-based AI decision-making unit constructs a mapping dataset from prevention and control plan parameters to prevention and control effects. It then periodically uses this mapping dataset to incrementally fine-tune the CNN-LSTM fusion model. During incremental fine-tuning, a penalty term weighted by the drug damage index PI is introduced into the loss function. This allows the model to generate a negative feedback bias on the combination of diagnosis and prevention and control parameters that are prone to drug damage while optimizing diagnostic accuracy. This penalty term is activated only when PI exceeds a preset drug damage tolerance threshold.

[0028] After the collaborative prevention and control execution unit completes the spraying and insect-attracting operations, the data self-learning feedback unit, within a preset evaluation period, re-collects visible light images and hyperspectral data at the same location. The cloud-based AI decision-making unit then runs the CNN-LSTM fusion model again to output the severity index of the pests and diseases after control. The normalized difference between the severity index from the pre-control detailed diagnosis and the post-control re-diagnosis is calculated to obtain the pest and disease control rate. Simultaneously, the re-collected visible light images after control are input into a pre-trained semantic segmentation model to extract the mango leaf region. Within this region, the ratio of yellowing area, the ratio of necrotic spot area, and the leaf curling index are calculated, and a weighted sum is obtained to obtain the phytotoxicity index. The control rate and phytotoxicity index for each pest and disease location are correlated with the corresponding control plan's spraying type, spraying amount, and pest and disease type to generate structured feedback data, which is then uploaded to the cloud-based AI decision-making unit. Based on this, the cloud-based AI decision-making unit constructs a mapping dataset from prevention and control plan parameters to prevention and control effects, periodically performs incremental fine-tuning of the model, and introduces a penalty term weighted by the pesticide damage index into the loss function. This penalty term is activated only when the pesticide damage index exceeds a preset tolerance threshold, so that the model generates a negative feedback bias for parameter combinations that are prone to causing pesticide damage.

[0029] Preferably, the cloud-based AI decision-making unit uses the received feedback data to update the CNN-LSTM fusion model, specifically including: The cloud-based AI decision-making unit serves as the central server for federated learning, while the edge computing units deployed in each mango orchard serve as local clients for federated learning. The central server maintains a global CNN-LSTM fusion model, and each local client maintains its own local copy of the CNN-LSTM fusion model. When the accumulated structured feedback data sample size of an edge computing unit in an orchard reaches the preset federated update trigger threshold, the edge computing unit sends a federated update ready request to the cloud AI decision unit. The cloud AI decision unit starts this round of federated learning. The edge computing unit, as a participating client in this round of federated learning, uses the locally accumulated structured feedback data to perform several rounds of local fine-tuning on the local CNN-LSTM fusion model copy, calculates the difference in model parameters before and after local fine-tuning as the local model update amount, and uploads the local model update amount to the cloud AI decision unit. After receiving the local model update data from all participating clients in this round, the cloud-based AI decision-making unit does not directly perform weighted average aggregation. Instead, it first calculates the cosine similarity between each pair of local model update data uploaded by each client to construct a model update direction similarity matrix between clients. Based on this similarity matrix, the participating clients are divided into several client clusters with the same update direction using spectral clustering. Clients within each client cluster exhibit similar model optimization trends in this federated update. For each client cluster, the cloud AI decision unit uses the number of structured feedback data samples from each client within the cluster as weight to perform a weighted average of the local model update amounts of each client within the cluster, thus obtaining the cluster aggregate update amount for that client cluster. Then, the cluster aggregate update amounts of each client cluster are weighted and summed using the proportion of the total sample size of each cluster as weight to generate the global model update amount. Finally, the global model update amount is applied to the global CNN-LSTM fusion model to complete this round of global model update, and the updated global model parameters are distributed to each edge computing unit to synchronously update each local CNN-LSTM fusion model copy.

[0030] The cloud-based AI decision-making unit acts as the central server for federated learning, while the edge computing units in each orchard serve as local clients. The central server maintains a global CNN-LSTM fusion model, and each client maintains a local copy of the model. When the accumulated structured feedback data sample size of a client reaches a preset threshold, it uses local data to fine-tune the local model copy for several rounds. The difference in model parameters before and after fine-tuning is calculated as the local model update amount and uploaded. After receiving the update amounts from all participating clients in this round, the central server first calculates the cosine similarity between each pair to construct an update direction similarity matrix. Then, it uses spectral clustering to divide the clients into several client clusters based on the consistency of their update directions. For each cluster, a weighted average is performed using the sample size as the weight to obtain the cluster aggregated update amount. Finally, the aggregated update amounts of each cluster are weighted and summed according to the proportion of the total sample size of each cluster to generate the global model update amount. This global model update amount is applied to complete the update for this round, and the updated global model parameters are distributed to each client to synchronize their local copies.

[0031] The present invention has at least the following beneficial effects: First, this system utilizes the light intensity histogram of mango leaf regions in visible light images, combined with the saturation and noise thresholds of the hyperspectral sensor, to calculate and adjust the integration time in real time. This ensures that the spectral data maintains a high proportion of effective pixels under light fluctuations, reducing false alarms caused by changes in light intensity. The edge computing unit runs a lightweight AI inference model to perform preliminary screening of multimodal data, submitting only suspected pest and disease data to the cloud for detailed diagnosis, thus shortening the overall response time.

[0032] Secondly, the edge computing unit automatically selects a preprocessing strategy based on the degree of data fluctuation. When the environment is stable, it employs edge-preserving denoising to retain subtle features. When the environment deteriorates, it introduces dehazing, scattering correction, and inter-modal spatial registration to ensure that radiometrically consistent and spatially aligned multimodal data is provided to the cloud. The edge-side screening network integrates visible light, hyperspectral, and environmental features, dynamically adjusts the screening threshold based on environmental stress confidence, and performs forced supplementary reporting by combining the spectral angular distance of the normal spectral library. This balances screening sensitivity and resistance to environmental interference under limited computing power.

[0033] Third, the integration time adjustment is achieved by using the leaf region illumination histogram and sensor thresholds as constraints to solve for the optimal integration time. Time-series fitting is then used to predict illumination change trends, enabling a feedforward response to future illumination and ensuring radiometric consistency across multiple frames of spectral data. A cloud-based CNN-LSTM fusion model establishes a correspondence between spatial and spectral features using cross-modal attention, enhancing the ability to locate early, minute lesions. Environmental condition gating introduces multi-timeframe environmental accumulation indicators and a pest-environment correlation matrix to differentially scale the fused features, suppressing environmental stress interference. Multi-granularity spatial pooling and channel gating filter retain minute anomalous signals while filtering out irrelevant spectral channels, improving the information density of the fused features.

[0034] Fourth, the prevention and control plan generation model the real-time status of equipment and physical constraints as a bipartite graph matching problem with capacity constraints, and arbitrates and schedules according to prevention and control priorities during resource competition. The micro-meteorological safety window determination enables spraying tasks to be actively suspended under adverse weather conditions and converted to insect trapping enhancement or drone supplementary prevention. Multiple sprayers coordinate through distributed time-series paths to eliminate cross-row conflicts.

[0035] Fifth, the data self-learning feedback unit quantifies the control rate by comparing diagnostic results at the same location before and after control measures. It extracts pesticide damage characteristics from visible light images to synthesize a pesticide damage index, and introduces a pesticide damage gating penalty term during model updates, applying negative feedback constraints to parameter combinations prone to pesticide damage. In federated learning updates, clustering is performed based on the similarity of model update directions, followed by intra-cluster aggregation and inter-cluster fusion. This allows a few orchard signals with abnormal updates to participate in global model updates as independent clusters, mitigating performance degradation caused by cross-domain data distribution shifts.

[0036] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Detailed Implementation

[0037] The present invention will be further described in detail below with reference to embodiments, so that those skilled in the art can implement it based on the description.

[0038] In a 1,000-mu (approximately 67 hectares) Tainong No. 1 mango orchard in Baise, Guangxi, anthracnose and thrips are the two most serious pests and diseases that cause damage before and after the rainy season each year. Previously, the orchard used a single hyperspectral diagnostic system, deploying one hyperspectral sensor per 15 mu (approximately 1 hectare) in the orchard. Each sensor continuously collected mango canopy spectral data at fixed integration times, and all data was transmitted back to a remote server for centralized analysis via a 4G network. In actual use, it was found that the system's diagnostic accuracy fluctuated drastically with weather changes. The false alarm rate was as high as 23.5% during cloudy days, and the overall diagnostic accuracy was only about 81.5%. Furthermore, the average time from data acquisition to obtaining diagnostic results exceeded 30 minutes. The main reason was that under strong midday sunlight, the leaf reflectance spectrum frequently reached the sensor's saturation threshold, resulting in severe spectral distortion. In the early morning, evening, and when under heavy cloud cover, insufficient light energy entered the sensor, and the spectral signal was overwhelmed by dark noise, leading to the loss of effective diagnostic information. Simultaneously, a large amount of redundant spectral data collected under normal lighting conditions occupied communication bandwidth and server processing resources, causing real abnormal samples requiring diagnosis to accumulate and delay in the transmission and calculation queues.

[0039] To address the aforementioned issues, this embodiment deploys the mango pest and disease early diagnosis and control system based on multimodal sensing and AI of the present invention. One multimodal sensing unit is installed per 10 mu (approximately 1.65 acres) in the orchard. Each unit integrates an 8-megapixel visible light camera, a hyperspectral sensor operating in the 900nm-1700nm wavelength range, an air temperature and humidity sensor, and a soil moisture sensor. Each sensing unit is connected to an edge computing box deployed in the field via a network cable. The edge computing box has a built-in quad-core ARM processor and a neural network acceleration unit, running a lightweight semantic segmentation model and a MobileNetV3 screening network. After the system is powered on, the visible light camera acquires visible light images of the mango canopy at a frequency of one frame every 30 seconds. The edge computing box uses the semantic segmentation model to extract the mango leaf area for each frame, filtering out soil background and strongly reflective non-leaf areas caused by direct sunlight, and statistically analyzing the grayscale histogram of all pixels within the leaf area. The edge computing box reads the saturation radiance threshold and dark noise equivalent radiance threshold specified in the hyperspectral sensor manual. Constrained by the current gain level, it iterates through the selectable integration time range, calculating the proportion of effective pixels with light intensity between the dark noise equivalent radiance and saturation radiance within the leaf area at each integration time. The integration time that maximizes this proportion is selected as the optimal integration time for the current frame, and this optimal integration time is written to the hyperspectral sensor register before the end of the current acquisition cycle. This allows the hyperspectral sensor to be exposed at the new integration time in the next acquisition cycle. The acquired hyperspectral data is normalized and stored after minimum noise separation transformation denoising and radiometric calibration. At this point, the system defaults to the first processing mode. Simultaneously, the edge computing box inputs the preprocessed visible light image and hyperspectral data into the MobileNetV3 screening network. The network outputs pest and disease suspicion scores and environmental stress confidence scores in parallel. When the environmental stress confidence score exceeds a preset first threshold, the screening threshold automatically increases from the default value. Only samples with suspicion scores exceeding the current effective threshold are packaged and uploaded to the cloud server along with temperature, humidity, and soil moisture data.

[0040] Taking a typical operation from 10:00 AM to 11:00 AM on April 15, 2025 as an example, the weather changed from cloudy to sunny, and the light intensity rose sharply from about 8,000 lux to over 70,000 lux within 30 minutes before falling back to about 20,000 lux due to cloud cover. The system continuously acquired 60 frames of visible light images corresponding to 60 acquisition cycles. The optimal integration time calculated by the edge computing box after statistically analyzing the illumination histogram of the blade area gradually shortened from 9.2 milliseconds at the darkest time to 0.7 milliseconds at the brightest time, and then extended to 3.5 milliseconds as the light intensity decreased. Throughout the process, the pixel saturation rate in the blade area remained below 0.3%, and the proportion of pixels below the dark noise threshold was always controlled within 2.1%. During this period, the MobileNetV3 screening network identified 12 suspected pest and disease samples from 60 sets of multimodal data and uploaded them to the cloud. The cloud-based CNN-LSTM fusion model performed a detailed diagnosis on these 12 samples, confirming that 11 of them showed early symptoms of anthracnose, while 1 sample was a false positive due to normal leaves. The average interval between the diagnosis completion time and the corresponding data acquisition time was 4.2 minutes. In comparison, the working records of a single hyperspectral diagnostic system originally deployed in the same area during the same period were retrieved. With a fixed integration time of 3 milliseconds, the pixel saturation rate reached 31% in 15 consecutive frames of spectral data acquired during a period of rapid light enhancement, and the proportion of pixels with a signal-to-noise ratio below 5 dB in 8 consecutive frames of spectral data acquired during a period of light reduction reached 27%. The average delay in outputting the diagnostic results after all data processing was 35 minutes. The false positive rate was 22% during cloudy weather, and 7 missed reports occurred during the midday strong light period due to spectral distortion.

[0041] At the same mango plantation in Baise, Guangxi, the system operates stably after dynamic adjustments based on the completion time of integration. However, during the mango flowering and fruit setting period from March to May each year, dew frequently forms on the leaf surfaces in the early morning and evening. At midday, during the hot and dry period, agricultural operations and natural winds stir up large amounts of dust that adhere to the leaf surfaces. In actual operation, the edge computing box found that morning water vapor condensation on the leaf surfaces caused an overall rise in the baseline of the hyperspectral reflectance spectrum received by the hyperspectral sensor across all wavelengths. Afternoon dust caused a uniform hazy blur in the visible light image. Simultaneously, the Mie scattering effect of dust particles caused a wavelength-dependent gradual shift in the near-infrared spectrum. If a fixed denoising and normalization preprocessing procedure is used—that is, Gaussian filtering is uniformly applied to the visible light image followed by histogram equalization, and standard normal transformation is uniformly applied to the hyperspectral data followed by normalization—a large number of false positives appear in the edge screening network. Specifically, during the morning dew period, the rise in spectral baseline was misidentified as an abnormal increase in near-infrared reflectance caused by pest and disease stress; during the afternoon dust period, the blurring of texture caused by image fogging was misjudged as the early powdery covering characteristics of powdery mildew, resulting in a false alarm rate of about 18% for edge screening during this period, with a large amount of normal data being incorrectly reported to the cloud.

[0042] To address the aforementioned data quality degradation issue, the edge computing box implements a layered preprocessing strategy. First, the edge computing box maintains a sliding window in memory with a length of 10 acquisition cycles. Within each cycle, it calculates the standard deviation and averages the noise in each band of the hyperspectral sensor within the 900nm to 1700nm range. Simultaneously, it calculates the average rate of change of corresponding pixel grayscale values ​​between adjacent frames of the visible light image. The two are then weighted and summed to obtain the data fluctuation index for the current moment. On a typical afternoon of cloudy skies followed by sudden dust storms, the system detected that this fluctuation index climbed from approximately 0.12 to over the preset threshold of 0.35 within 15 minutes. The edge computing box determined that the current data quality had degraded due to environmental disturbances and automatically switched from the first processing mode to the second processing mode. After switching, the visible light image is first processed using a dark channel prior dehazing algorithm. The atmospheric light value and transmittance map are estimated based on the minimum value of the dark channel in a local area of ​​the image, and the clear leaf image after dehazing is retrieved. Then, guided filtering is applied to the dehazing result to preserve edges and reduce noise. For the hyperspectral data, a multivariate scattering correction method is used. All sample spectra are linearly regressed against the reference spectrum. The regression intercept is subtracted from each spectrum and then divided by the regression slope to eliminate the overall spectral baseline drift and scattering noise caused by water vapor and dust scattering on the leaf surface. In the radiometric calibration stage, the absolute radiance value of the hyperspectral data is converted according to the currently effective integration time. Then, the processed hyperspectral data and the visible light image are jointly registered in spectral and spatial order before normalization. By calculating the positional correspondence of the leaf vein intersection points in the two modal images, affine transformation is used to correct the pixel-level spatial offset between modalities caused by water vapor refraction and dust scattering.

[0043] To verify the actual effect of the hierarchical preprocessing strategy, within a 30-minute period during the same dusty weather, the preprocessing results from both the first and second processing modes output by the edge computing box of this system were screened in parallel for comparison. Under the first processing mode, the edge screening network output 24 samples with suspected pest and disease scores exceeding the threshold. After review by the cloud-based CNN-LSTM fusion model, 19 of these were confirmed as false alarms caused by dust interference, a false alarm rate of approximately 79%. Under the second processing mode, the screening network output 6 samples exceeding the threshold. Cloud-based review confirmed that all 6 were genuine early anthrax lesion samples, with no false alarms. Simultaneously, a retrospective analysis of the spectral angular distance of the 6 genuine lesion samples was performed. The preprocessed spectrum output by the second processing mode had an average spectral angular distance of 0.21 radians from the normal sample database in the characteristic bands, significantly greater than the 0.08 radian threshold. In contrast, the first processing mode, due to the failure to eliminate dust scattering effects, resulted in spectral feature compression, with an average spectral angular distance of only 0.06 radians, failing to trigger effective diagnosis. Radiometric consistency was assessed by performing a radiometric consistency evaluation on 30 consecutive cycles of preprocessed spectral data before and after the switch. The inter-frame variation coefficient of radiance in each band under the second processing mode decreased from 11% in the first processing mode to 3.2%, indicating that the spectral-spatial joint registration and radiometric consistency normalization correction effectively eliminated the radiometric inconsistency problem of multi-frame data caused by environmental factors.

[0044] During the high-temperature and drought season from July to September each year in the mango planting base in Baise, Guangxi, daytime temperatures consistently exceed 35 degrees Celsius, and soil moisture content drops below 40% of field capacity. Mango leaves exhibit widespread water loss, curling, and yellowing due to water stress. After the system completed adaptive adjustment of the integral time and data stratification preprocessing as described, it ran stably. However, the conventional screening model running within the edge computing box encountered serious false alarms when processing this type of data. Specifically, the leaf spectra caused by drought stress showed a decrease in reflectance in the near-infrared band similar to early-stage anthracnose. The visible light texture characteristics of curled leaves were highly similar to leaf deformities caused by thrips damage in the shallow response of the convolutional neural network. This led the screening model to upload a large number of leaves from drought-stressed areas as suspected pest and disease samples to the cloud with high confidence for several consecutive days, with the highest number of false alarms reaching 47 per day. The cloud-based diagnostic resources were overwhelmed by a large amount of invalid data, and the overall screening accuracy of the system dropped to approximately 71%.

[0045] To address the difficulty in distinguishing between environmental stress and pest / disease stress, the edge computing box implements the dual-branch screening and environmental stress adaptive adjustment mechanism described in claim 3. The edge computing box is pre-programmed with a trained dual-branch lightweight network and a normal mango leaf hyperspectral reference library. The first branch uses MobileNetV3 as its backbone, taking a preprocessed three-channel visible light image as input. After spatial feature map extraction via depthwise separable convolution, it is compressed into a 128-dimensional visible light feature vector using global average pooling. The second branch uses four one-dimensional convolutional blocks as its backbone, each containing 64 convolutional kernels of length 5 and a max-pooling layer. It takes preprocessed hyperspectral data as input and outputs a 256-dimensional spectral feature vector. The environmental encoding module is a three-layer fully connected network, taking four environmental parameters—air temperature, air humidity, soil moisture content, and soil temperature—as input for the current acquisition period and outputting a 64-dimensional environmental feature vector. The three feature vectors are concatenated and fed into a classification head consisting of two fully connected layers. The first layer contains 128 neurons and uses the ReLU activation function. The second layer outputs two values, which are mapped by the Sigmoid function to a suspected pest and disease score and an environmental stress confidence score between 0 and 1, respectively.

[0046] On a typical hot and dry afternoon, a multimodal sensing unit acquired data from a mango canopy. The visible light image clearly showed upward curling of leaf edges and yellowing between veins. Hyperspectral data revealed reflectance anomalies in the 980nm and 1480nm bands. After the edge computing unit fed the data into a dual-branch network, the classification head output a suspected pest / disease score of 0.71 and an environmental stress confidence score of 0.83. Since the environmental stress confidence score of 0.83 exceeded the preset first threshold of 0.6, the edge computing unit automatically increased the screening threshold from the default value of 0.5 to 0.8. The suspected pest / disease score of 0.71 did not reach the increased threshold, and the sample was not classified as a suspected pest / disease, avoiding false positives. Simultaneously, the edge computing unit maintained a normal mango leaf hyperspectral reference library in memory containing 200 spectra of normal leaves acquired under confirmed pest / disease-free and suitable environmental conditions. Each spectrum covered 512 bands from 900nm to 1700nm. The edge computing box compares the current sample's spectra in the 900nm-1000nm and 1400nm-1700nm ranges with each normal spectrum in the reference library band by band, calculating the cosine of the angle to obtain an average spectral angular distance of 0.19 radians. This value exceeds the preset second threshold of 0.15 radians, while the environmental stress confidence score of 0.83 is higher than the preset third threshold of 0.5. Therefore, the spectral angular distance forced reporting mechanism is not triggered, and the sample is ultimately determined to be a pure environmental stress response.

[0047] In contrast, the batch of data was input into a conventional single-branch screening network without environmental stress adaptive adjustment. This network only processed visible light images using MobileNetV3, outputting a single pest / disease confidence score of 0.68. Based on a fixed threshold of 0.5, it was classified as a suspected pest / disease and reported to the cloud. After a detailed diagnosis of the sample by the cloud-based CNN-LSTM fusion model, considering features such as the yellowing area in the time-series images not expanding over three consecutive acquisition cycles and the linear correlation between changes in hyperspectral 1400nm band reflectance and decreasing soil moisture content, the sample was ultimately determined to be due to drought stress rather than pests / diseases. However, cloud computing resources were already consumed. Statistics showed that during a seven-day period of high temperature and drought, the edge processing unit of this system output 31 suspected samples, of which 28 were confirmed as actual pests / diseases by the cloud, achieving a screening accuracy of approximately 90%. In contrast, the control network without environmental stress adaptive adjustment output 112 suspected samples during the same period, with only 34 confirmed as actual pests / diseases by the cloud, achieving a screening accuracy of approximately 30%. The remaining samples were false alarms caused by drought stress.

[0048] In the daily operation of the mango planting base in Baise, Guangxi, after the system completes the integration time adjustment, layered preprocessing, and dual-branch screening, the edge computing box's adaptability to drastic light fluctuations is significantly better than the fixed integration time scheme. However, when there are complex lighting conditions in the orchard with both direct sunlight spots and dense shadows, adjusting the integration time based solely on the overall average light intensity of the visible light image or simple regional light statistics is still insufficient. Specifically, at noon in summer, sunlight penetrates the canopy gaps and forms circular bright spots with a diameter of about 5 to 15 centimeters on some leaves. The radiance of the spot area reaches 2 to 3 times the sensor saturation threshold, while the radiance of the shadow area in the same frame, which is blocked by the upper leaves, is only slightly higher than the equivalent radiance of dark noise. When using the overall average light intensity to calculate the integration time, the proportion of saturated pixels in the spot area exceeds 15%, and the proportion of pixels with a signal-to-noise ratio below 5 dB in the shadow area reaches 21%, with saturation and underexposure coexisting in the same frame of spectral data. In addition, cumulus clouds develop rapidly in the afternoons of summer in the south. The movement of the cloud layer causes the canopy light intensity to fluctuate drastically from thousands to tens of thousands of lux within a few seconds to tens of seconds. The integration time calculated and issued in the previous collection cycle is seriously mismatched with the actual light intensity before and after the cloud cover.

[0049] To address the aforementioned issues, the edge computing box implements the fine-tuning mechanism for integral time. The lightweight semantic segmentation model deployed on the edge computing box employs a U-Net architecture, with the encoder using MobileNetV2 as its backbone network. After pre-training on a general image segmentation dataset, it undergoes transfer training using 2000 labeled images of a mango orchard scene. The labeled categories include mango leaves, branches, soil, sky, and reflective areas. After each visible light image enters the edge computing box, the semantic segmentation model completes inference within 15 milliseconds, outputting a pixel-level classification result of the same size as the input image. The edge computing box extracts the pixel region classified as a mango leaf, iterates through the grayscale values ​​of all pixels within this region, and generates a light intensity histogram by counting the pixel frequencies in 256 intervals from 0 to 255. The edge computing box reads the saturation radiance threshold and dark noise equivalent radiance threshold stored in the hyperspectral sensor datasheet, and converts these two thresholds into corresponding pixel grayscale value intervals based on the sensor's current gain level. With the sensor currently at its gain, the integration time can be selected from 0.1 ms to 20 ms in 0.1 ms increments. For each selectable integration time value, the edge computing box maps each gray level in the illumination histogram to an equivalent radiance value based on the linear relationship between radiance and integration time. It counts the proportion of pixels falling above the dark noise threshold and below the saturation threshold to the total number of leaf pixels and selects the integration time that maximizes this proportion as the optimal integration time for the current frame.

[0050] The edge computing unit maintains an 8-frame integration time series window in memory, recording the optimal integration time values ​​for 8 consecutive frames. With each new frame of data, the edge computing unit performs a least-squares linear fit on the timestamps and integration time values ​​within the window, obtaining the slope and intercept. The slope represents the average rate of change of integration time over time, and the goodness-of-fit R² value represents the linearity of the illumination change trend. The current time is added to an offset of one acquisition cycle and substituted into the fitted line to obtain the estimated optimal integration time for the next acquisition moment. If the goodness-of-fit is higher than 0.85, it indicates a clear and stable illumination change trend; the weight of the estimated optimal integration time is 0.7, and the weight of the optimal integration time for the current frame is 0.3. If the goodness-of-fit is between 0.5 and 0.85, the weights of both are 0.5. If the goodness-of-fit is lower than 0.5, it indicates that the illumination change has no obvious trend or is in a state of random fluctuation; the weight of the estimated value is 0.2, and the weight of the optimal integration time for the current frame is 0.8. The final integral time is obtained after weighted summation and written to the hyperspectral sensor register via serial communication before the end of the current acquisition cycle.

[0051] A comparative verification was conducted during a period of active cumulus cloud development, from 2:00 PM to 2:30 PM on a certain afternoon in August 2025. During this period, the canopy illumination intensity fluctuated drastically between approximately 6,000 lux and 95,000 lux, with cloud drift cycles of approximately 2 to 5 minutes. Our system was run synchronously with a comparative scheme that used overall average illumination to calculate the integration time without time-series prediction. In the comparative scheme, at 2:12 PM, when a dense cumulus cloud suddenly blocked sunlight, the actual illumination dropped sharply from approximately 80,000 lux to approximately 9,000 lux. However, the integration time remained at 1.1 milliseconds, calculated based on strong light conditions in the previous cycle. This resulted in only 62% of the effective pixels above dark noise in the shadowed areas of the hyperspectral data frame. In the same integration time series window, the first 8 frames of data in this system showed a clear downward trend, with a negative linear fitting slope, a goodness of fit of 0.91, and an estimated optimal integration time of 6.8 milliseconds. After weighted fusion with the current frame's optimal integration time of 7.2 milliseconds, the final integration time was 7.0 milliseconds, achieving an effective pixel ratio of 97%. At 14:18, after the clouds moved away, the illumination suddenly increased. Due to a response lag, the comparison scheme acquired a frame of spectral data with a saturated pixel ratio of 23%. This system shortened the integration time in advance through time series prediction, controlling the saturated pixel ratio to 0.7%. Statistics show that within this 30-minute period, this system acquired 1800 frames of hyperspectral data. The effective pixel ratio in the leaf area consistently remained above 94%, and the average inter-frame coefficient of variation for radiance across different bands was 3.6%. In contrast, the effective pixel ratio of the comparison scheme dropped to a minimum of 61% during the same period, and the average inter-frame coefficient of variation for radiance was 12%. The cloud-based CNN-LSTM fusion model's extraction of temporal features from the comparison scheme's data was significantly interfered with.

[0052] During the hot and dry season from July to September each year in the mango planting base in Baise, Guangxi, mango plants often suffer from water stress and multiple diseases and pests such as anthracnose and angular leaf spot. Conventional CNN-LSTM diagnostic models deployed in the cloud have revealed two shortcomings when processing such data. First, the model simply splices together features extracted independently from visible light images and hyperspectral data without establishing a correspondence between visible light spatial texture and hyperspectral band response. When the diameter of early anthracnose lesions is only 1 to 3 mm and accounts for less than 0.5% of the leaf surface, the spatial features are submerged by global average pooling, and the spectral features cannot focus on the corresponding pixels of the lesions due to the lack of spatial guidance. Secondly, the model failed to incorporate air temperature and humidity and soil moisture as constraints into the reasoning. It misidentified the decrease in near-infrared reflectance of leaves caused by drought stress as a characteristic of anthracnose, and misidentified leaf margin necrosis caused by high temperature burn as angular leaf spot. Under the combined conditions of high temperature above 35 degrees Celsius and soil moisture content below 40%, the misidentification rate of disease types reached about 26%, and the average deviation of the severity index exceeded 0.3, resulting in a systematic overestimation of the amount and scope of spraying in subsequent control plans.

[0053] To address the aforementioned issues, the cloud-based AI decision unit runs the CNN-LSTM fusion model described in this system for refined diagnosis. Upon model startup, it loads the ResNet-18 backbone network weights pre-trained on a general image dataset, along with all parameters from a spectral encoder, a cross-modal multi-head attention fusion module, an environmental condition gating module, and a two-layer LSTM temporal encoder trained using 5000 multimodal mango disease and pest samples. When the edge computing unit reports a set of suspected disease and pest data, the data packet includes preprocessed visible light images, hyperspectral data, air temperature and humidity values ​​at the current acquisition time, and soil moisture values, as well as the suspected disease and pest score (0.71), environmental stress confidence score (0.83), and spectral angular distance (0.19 radians) output from the edge screening stage. The model concatenates the suspected disease and pest score and the environmental stress confidence score into a two-dimensional edge prior feature vector, assigning it to the initial hidden state and cell state of the two-layer LSTM, so that the temporal encoder carries the initial judgment bias of the edge end regarding the current sample at the start of inference. At the same time, the spectral angular distance is added as an extra dimension to the end of the multi-level spectral feature vector output by the spectral encoder.

[0054] Visible light images, after size normalization, are input into the ResNet-18 backbone network. The network's four residual blocks output four layers of feature maps with spatial resolution halved sequentially and channel counts doubled sequentially. The highest-level feature map, with a resolution of 1 / 16 of the input and 256 channels, is selected as the multi-scale spatial feature map. Hyperspectral data is input into the spectral encoder. Three consecutive one-dimensional convolutional blocks have kernel sizes of 7, 5, and 3, with a stride of 1, and each layer outputs 64, 128, and 256 channels, respectively, resulting in three feature vectors at different spectral resolution levels, denoted as low-level, mid-level, and high-level spectral features. A cross-modal multi-head attention fusion module uses eight attention heads. The spatial locations of the highest-level spatial feature map are flattened one by one as query vectors, and the high-level spectral features are used as key-value vectors. Attention weights between each spatial location and each spectral band are calculated, generating a spatial-spectral joint enhanced feature map of the same size as the input spatial feature map. In this model, the response values ​​of the pixel regions corresponding to early anthrax lesions are significantly enhanced.

[0055] The spatial-spectral joint enhanced feature map is concatenated with low-level, mid-level, and high-level spectral feature vectors along the channel dimension after global average pooling to obtain the multimodal fusion feature vector at the current time. The environmental condition gating module receives data such as air temperature (36.2 degrees Celsius), air humidity (45%), and soil moisture content (18%) at the current acquisition time. A set of environmental gating coefficients with the same dimension as the multimodal fusion feature vector is generated through a fully connected layer. Each dimension of the fusion feature vector is element-wise multiplied by the corresponding gating coefficient, suppressing the temperature-related dimension and moderately enhancing the humidity-related dimension. The gated fusion feature vector is fed into a two-layer LSTM, and the output temporal-aware feature vector is updated by combining edge prior bias and the current hidden state. The classification branch of the temporal multi-task output module takes the temporal-aware feature vector as input, and outputs anthrax probability of 0.87, angular leaf spot probability of 0.09, powdery mildew probability of 0.03, and normal probability of 0.01 after passing through a fully connected layer and Softmax normalization. The severity regression branch concatenates the temporal-aware feature vector with the learnable embedding vector of anthrax category and inputs it into the fully connected layer, outputting a severity index of 0.34. This index represents the result of a weighted combination of lesion area ratio and normalized spectral index, corresponding to mild to moderate infection.

[0056] The same set of data was input into a comparative model for diagnosis. The comparative model used ResNet-18 to extract visible light features and three-layer one-dimensional convolution to extract spectral features, which were then directly concatenated without cross-modal attention or environmental gating. Classification and regression results were output via a single-layer LSTM. The comparative model output an anthracnose probability of 0.52 and a angular leaf spot probability of 0.41, misclassifying the disease as angular leaf spot. The severity index output was 0.68, significantly higher than the actual value. Statistical analysis was performed on 37 suspected samples reported by the marginal end on the same hot and dry day. The diagnostic results of this system were subsequently verified by field sampling and microscopic examination by plant protection personnel. The system accurately identified 34 disease types, with an accuracy rate of approximately 92%. The average absolute deviation of the severity index from the measured lesion area ratio and the spectral index was 0.06. The comparative model accurately identified 23 disease types on the same samples, with an accuracy rate of approximately 62%, and the average absolute deviation of the severity index was 0.31. Based on accurate disease type identification, the severity regression branch of this system receives the embedding vector of the highest probability category of the classification branch as a conditional input, and the output severity index quantification is more in line with the actual development level of the disease.

[0057] In the daily operation of the mango planting base in Baise, Guangxi, after the system completes multimodal data collection, preprocessing, edge screening, and cloud-based precision diagnosis according to the aforementioned scheme, the cloud-based CNN-LSTM fusion model has significantly improved the accuracy of identifying common diseases such as anthracnose and angular leaf spot. However, it still misses some samples when processing very early-stage micro-disease samples. The diameter of the lesions in these samples is usually only 1 to 2 millimeters, accounting for less than 0.3% of the area in the entire mango leaf image, and is almost invisible to the naked eye. When the spatial-spectral joint enhanced feature map is directly compressed into a single vector by global average pooling, the weak abnormal response of the lesion region is submerged by the normal tissue pixel response, which is tens of thousands of times greater. The activation value of the disease-related dimension in the pooled feature vector is only less than 0.05 higher than that of the normal region. This slight difference is further diluted when it is subsequently spliced ​​and fused with multi-level spectral feature vectors. Meanwhile, in the multi-level spectral feature vectors output by the spectral encoder, a large number of redundant band channels unrelated to the current stage of disease development are rigidly spliced ​​into the fusion features. This forces the temporal encoder to learn discrimination patterns from hundreds of dimensions of irrelevant noise, significantly increasing the risk of overfitting when the sample size of early micro-diseases is already limited. In a field test conducted at the base in August 2025 during the initial stage of anthrax, the conventional global pooling splicing scheme had a detection rate of only about 61% for samples with lesion area less than 0.5%, and nearly 40% of the missed samples developed into moderate to severe infections within the following two weeks.

[0058] To address the aforementioned issues, the cloud-based AI decision unit employs the multi-granularity spatial enhancement and spatially conditionalized spectral filtering scheme described in this system when performing multimodal feature fusion. When the spatial-spectral joint enhanced feature map is output from the cross-modal multi-head attention fusion module, its size is 256 channels multiplied by 14 pixels multiplied by 14 pixels. Before being fed into global average pooling, these 256 channels are uniformly divided into 8 groups along the channel dimension, with each group containing 32 channels. For the 32 channel feature maps within each group, a Softmax normalized weight is calculated for each spatial location across the 32 channels. Then, a weighted average is performed on each spatial location using this weight, resulting in a 14x14 single-channel spatial attention map. The average of all spatial locations in this map is then taken to obtain a scalar value; a total of 8-dimensional group spatial convergence vectors are obtained from the 8 groups. Simultaneously, global average pooling is directly performed on the original 256-channel feature map to obtain a 256-dimensional global representation vector. By concatenating the 8-dimensional group spatial convergence vector with the 256-dimensional global representation vector, a 264-dimensional spatial feature descriptor is obtained. The 8-dimensional group convergence vector explicitly preserves the small-scale high-response patterns that may be concentrated in micro-lesions.

[0059] The low-level, mid-level, and high-level spectral feature vectors output by the spectral encoder, with 64, 128, and 256 channels respectively, were used to calculate normalized spectral vegetation indices for each level. During calculation, the average reflectance value in the 900-1000 nm range was used as the baseline value for chlorophyll sensitivity, and the average reflectance value in the 1420-1500 nm range was used as the baseline value for cell structure sensitivity. The difference between the two values ​​was subtracted, and the result was divided by the sum of the two values ​​to obtain the index value corresponding to that spectral level. The three index values ​​were concatenated to obtain a 3D spectral index feature vector, which was then used as an explicit physiological prior for pigment and cell structure changes in subsequent fusion.

[0060] Before channel-level stitching, the multi-level spectral feature vectors undergo adaptive filtering via learnable channel gating. Channel gating is implemented using a two-layer fully connected network, taking a 264-dimensional spatial feature descriptor as input and outputting a retention probability vector corresponding to the total 448 dimensions of the multi-level spectral feature vector, with each element ranging from 0 to 1. A gating threshold of 0.5 is set, retaining only channels with a probability higher than 0.5 for stitching, discarding the rest. When the input sample is a very early anthrax sample, the spatial convergence vector in the spatial feature descriptor detects localized minor anomalous patterns. Based on this, the channel gating outputs a retention probability of 0.87 to 0.93 for spectral channels related to phenolic compound accumulation in the near-infrared band, while outputting a retention probability of only 0.12 to 0.31 for a large number of redundant channels related to slight changes in chlorophyll in the visible light band. Ultimately, only about 38% of the spectral channels are retained. The spatial feature descriptor, the filtered spectral channels, and the 3D spectral index feature vector are concatenated to form a concise and high-information-density multimodal fusion feature vector for the current moment, which is then fed into the subsequent environmental condition gating and timing encoder.

[0061] Diagnostic comparisons were made on the same batch of early anthracnose samples under conditions of enabling and disabling multi-grained pooling and channel gating. A total of 45 samples were collected and individually confirmed by field plant protection experts using a 20x handheld magnifying glass. Among them, 23 samples were at the very early stage with lesion area less than 0.5%, and 22 samples were at the early stage with lesion area between 0.5% and 2%. The diagnostic results with multi-grained pooling and channel gating disabled showed that 14 very early-stage samples were correctly identified, a detection rate of approximately 61%. The diagnostic results with this solution enabled showed that 20 very early-stage samples were correctly identified, a detection rate of approximately 87%, an improvement of 26 percentage points. For the 22 early-stage samples, the detection rates of the two solutions were approximately 91% and 95%, respectively, with little difference, indicating that multi-grained pooling plays a crucial role mainly when the lesion area is very small. Analysis of the spectral channels after gating revealed a 78% agreement rate with the characteristic bands related to phenol oxidation and cell wall degradation recorded in the literature on the pathogenesis of mango anthracnose, indicating that the channels screened by the gating mechanism have a reasonable plant pathological basis.

[0062] In the long-term operation of the mango planting base in Baise, Guangxi, the system, following the aforementioned scheme, has been able to effectively distinguish between environmental stresses such as drought and pest and disease stresses. However, diagnostic confusion still exists when dealing with anthracnose and angular leaf spot. Anthracnose spores require continuous leaf moisture for more than 12 hours to germinate, and usually occur in concentrated periods of high humidity after several consecutive days of heavy dew or rain. Angular leaf spot, on the other hand, prefers a weather transition from high temperatures to sudden rainfall, with the pathogen invading when the leaf stomata open. In mid-June 2025, the base experienced a special weather event: for the previous five consecutive days, the soil moisture content remained between 22% and 25%, and the daytime high temperature was 33 to 35 degrees Celsius. On the morning of the sixth day, dense fog lasted for eight hours, with leaf moisture lasting for nine hours. At 10:00 AM, the temperature suddenly rose to 36 degrees Celsius, and the air humidity plummeted from 95% to 42%. In the following 48 hours, early symptoms of anthracnose and angular leaf spot appeared simultaneously in multiple areas of the orchard. After the edge computing box reports suspected samples, the cloud-based simplified environmental condition gating module generates a gating coefficient based solely on the current instantaneous air temperature and humidity. For samples collected during the transition from dense fog to a sudden temperature rise, the instantaneous air humidity of 42% was judged as a dry condition. The gating coefficient incorrectly suppressed the 980 nm and 1450 nm band responses that are sensitive to anthrax characteristics, misclassifying early anthrax lesions as angular leaf spot. Nine out of 27 samples in this batch were misclassified, and the accuracy rate of distinguishing between the two diseases was only about 67%.

[0063] To address the aforementioned issues, the environmental condition gating module of the cloud-based AI decision-making unit executes the cumulative environmental index extraction and pest and disease condition perception modulation scheme described in this system. The environmental condition gating module maintains a sliding time window with a length of 72 acquisition cycles in memory, calculating at 10-minute intervals, covering a 12-hour time span. When the edge computing box reports the sample collected at 10:00 AM on the 6th day, the environmental condition gating module retrieves 72 sets of air temperature and humidity sequences and soil moisture sequences from the previous 71 times and the current time within the window. Based on the air temperature and humidity sequences, the module counts the number of consecutive acquisition cycles where the relative humidity is consistently above 85%, multiplying this by the acquisition cycle duration to obtain a leaf surface moisture duration of 8.5 hours; it then accumulates the portion of air temperature exceeding 15 degrees Celsius cycle by cycle and calculates the time integral to obtain an effective accumulated temperature of 186 degrees Celsius per day. Based on the soil moisture sequence, the average soil moisture content of the previous 72 cycles at the current time is taken as 24%, compared with the field capacity of 35%, to calculate a soil moisture deficit index of 0.31. The above parameters—leaf surface wetting duration of 8.5 hours, effective accumulated temperature of 186 degrees Celsius per day, and soil moisture deficit index of 0.31—are concatenated with the current air temperature of 36 degrees Celsius, air humidity of 42%, and soil moisture content of 22% to form a 6-dimensional environmental parameter vector. This vector is then fused with the environmental stress confidence score of 0.62 reported by the edge computing box and input into the fully connected layer to generate the initial environmental gating coefficients.

[0064] The module's pre-defined pest-environment association matrix has 6 rows and 128 columns. The 6 rows correspond to 6 candidate pest / disease types: anthracnose, angular leaf spot, powdery mildew, thrips damage, leaf gall midge damage, and physiological stress. Each row's 128-dimensional vector is the environmental parameter response pattern embedding for that pest / disease type under typical occurrence conditions, learned during the training phase from 6000 historical data points labeled with pest / disease type and environmental parameters. The initial environmental gating coefficients are scaled and dot-product attention-based with each row of the pest-environment association matrix. The dot product of the initial coefficients with the anthracnose response pattern embedding is 0.81, and with the angular leaf spot response pattern embedding is 0.43. After Softmax normalization, the attention weights for anthracnose are 0.64 and angular leaf spot are 0.21, with the remaining weights distributed among the other 4 types. The rows of the association matrix are then weighted and summed according to their corresponding attention weights to generate a pest / disease condition-aware environmental gating bias vector. In this bias vector, the dimension related to the anthrax characteristic band response receives a positive bias of approximately 0.3, while the dimension related to keratosis receives a bias of only approximately 0.1. The initial environmental gating coefficients are added element-wise to the bias vector to obtain the final environmental gating coefficients.

[0065] After element-wise scaling of the multimodal fusion feature vectors of 27 samples from the same transition time point using the final environmental gating coefficient, the data was fed into the temporal encoder and classification branch. The classification branch output that 11 samples had an anthracnose probability higher than 0.5, 13 samples had an angular leaf spot probability higher than 0.5, and 3 samples were normal. Laboratory isolation and culture verification using diseased leaf samples collected by field plant protection personnel revealed 12 positive samples for anthracnose, 12 positive samples for angular leaf spot, and 3 normal samples. This system misclassified only one anthracnose sample as angular leaf spot, achieving a differentiation accuracy of approximately 93%. In comparison, the simplified environmental condition gating module, which uses only instantaneous environmental parameters to generate the gating coefficient, achieved a differentiation accuracy of approximately 67%. Further analysis of the contributions of the two cumulative indicators, duration of foliar wetting and effective accumulated temperature, revealed that for samples with a duration of foliar wetting exceeding 7 hours and an effective accumulated temperature between 150 and 220 degrees Celsius per day, the response intensity of the anthracnose-related channel in the final environmental gating coefficient increased by an average of about 45%. This modulation direction is highly consistent with the anthracnose occurrence conditions recorded in plant pathology literature, indicating that the attention mechanism of the pest-environment association matrix has indeed learned an environmental response pattern with actual pathological significance.

[0066] In the daily operation of the mango planting base in Baise, Guangxi, the system completes the entire process from multimodal perception and edge screening to cloud-based precise diagnosis according to the aforementioned scheme. The cloud-based AI decision-making unit then outputs the types and severity indices of pests and diseases at each location where they occur. In late July 2025, three pest and disease sites appeared simultaneously in the southeast area of ​​the base: Site A, located in plot No. 12, showed moderate anthracnose infection with a severity index of 0.52 and a clear trend of lesion spread; Site B, located in plot No. 15, showed mild thrips damage with a severity index of 0.28; and Site C, located in plot No. 18, showed severe angular leaf spot infection with a severity index of 0.71. This site was located upwind of the orchard, posing a high risk of spread to adjacent downwind plots. At the time, the base had two intelligent sprayers in operation. Sprayer No. 1 was located near field No. 10, with 18 liters of pesticide remaining, 85% battery power, and equipped with a standard fan-shaped nozzle. Sprayer No. 2 was located near field No. 20, with only 6 liters of pesticide remaining, 42% battery power, and equipped with a conical atomizing nozzle. The effective spraying capacity was capped at 20 liters and 8 liters respectively. There were also four insect-attracting lamps installed in fields No. 11, 14, 17, and 20.

[0067] Following the conventional method of generating control plans, the cloud platform generates operational instructions solely based on pest and disease diagnosis results and preset spraying amounts per acre, without considering the real-time status of the equipment or physical constraints. This method, after providing spraying plans for three pest and disease sites, assigned points A and C to the closer sprayer #1, requiring a total of 22 liters of pesticide. However, sprayer #1 only carried 18 liters. After completing spraying at point A, sprayer #1 ran out of pesticide, and point C only received approximately 4 liters before being forced to stop operations, resulting in ineffective coverage of the severely affected angular leaf spot area. Simultaneously, point B was assigned to sprayer #2, but at this time, sprayer #2's battery was only at 42%, and it ran out of power en route to field #15, becoming stranded between the rows. Ultimately, only point A was successfully controlled; insufficient coverage at point C allowed angular leaf spot to spread to three adjacent fields within a week.

[0068] To address the aforementioned issues, the cloud-based AI decision-making unit, after obtaining the diagnostic results, executes the prevention and control plan generation mechanism described in this system, which involves equipment capability perception and dynamic priority arbitration. First, the cloud-based unit initiates a status query to the data acquisition terminals of each device in the collaborative prevention and control execution unit via the MQTT IoT protocol, obtaining real-time status data for sprayers No. 1 and No. 2, as well as the current working status, light intensity level, and cumulative working time of the four insect-attracting lamps. Using the two available sprayers and four insect-attracting lamps as equipment capability layer nodes, and the three pest occurrence locations (A, B, and C) and their corresponding prevention and control action requirements as prevention and control requirement layer nodes, a two-layer bipartite graph is constructed. The prevention and control actions at point A include a requirement of 11 liters of anthracnose-specific fungicide and the coordinated start / stop requirement for insect-attracting lamps within a 15-meter coverage radius; the requirements at point B include 6 liters of thrips control agent; and the requirements at point C include 15 liters of angular leaf spot-specific fungicide. In calculating the edge weights, the spatial distance between the equipment and the location of the occurrence is measured using normalized Euclidean distance from GPS coordinates. For equipment status matching, the matching score for the fan-shaped nozzle of sprayer No. 1 with anthracnose and angular leaf spot foliar spraying is 0.9, and with thrips shoot spraying is 0.5; the matching score for the conical atomizing nozzle of sprayer No. 2 with thrips shoot spraying is 0.9. The total weight of each edge is determined by weighting the spatial distance weight, status matching score weight, and urgency of control weight at a ratio of 0.3, 0.3, and 0.4, respectively.

[0069] A minimum cost matching algorithm with capacity constraints is run on a bipartite graph. The constraints are: sprayer #1 has a maximum spraying capacity of 20 liters, sprayer #2 has 8 liters, and each insect-attracting lamp has an effective coverage radius of 15 meters. After iterative solving, the algorithm outputs the matching results: sprayer #1 is assigned to point A with a spraying capacity of 11 liters, and the target spraying area is a circular area with a radius of 15 meters centered on point A; sprayer #2 is assigned to point B with a spraying capacity of 6 liters; the insect-attracting lamp in field #17 is turned on and adjusted to a medium light intensity to cover the area at point A. At this time, the 15-liter spraying capacity demand at point C exceeds the respective capacity limits of the remaining 9 liters for sprayer #1 and the remaining 2 liters for sprayer #2, resulting in a conflict where multiple points compete for the same equipment and the capacity is insufficient. The cloud-based system calculates a prevention and control priority index for the three locations based on the severity of pests and diseases and their spread risk. Location C, with a angular leaf spot severity index of 0.71 and a high spread risk due to its upwind location, scores 0.87. Location A, with moderate anthrax but a clear spread trend, scores 0.64. Location B, with mild thrips and slow spread, scores 0.31. After prioritizing, the remaining 9 liters of pesticide from sprayer #1 are allocated to location C, which is marked as a partially covered area requiring supplementary treatment. A supplementary treatment reminder is automatically generated and pushed to the administrator's mobile terminal, indicating that a backup sprayer needs to be dispatched or that a portable sprayer should be used within 2 hours to cover the remaining 6 liters of pesticide at location C. For location B, sprayer #2 continues its normal operation according to the original allocation plan. Sprayer #2, with 42% battery, can complete a round trip based on the task path estimation and retains a 15% safety margin, requiring no rescheduling.

[0070] After receiving the supplementary prevention reminder, the management personnel dispatched the backup sprayer from the warehouse to complete the supplementary prevention of the remaining area at point C within 40 minutes. Spraying was completed at all three pest and disease sites within 2 hours of diagnosis. A week later, a follow-up visit to each site and surrounding area showed that angular leaf spot disease upwind of point C had not spread to adjacent fields, the anthracnose severity index at point A had decreased from 0.52 to 0.08, and the thrips population density at point B had decreased from 3.2 per leaf to 0.3. Compared to the conventional approach where sprayer #2 was delayed due to insufficient power, resulting in missed prevention at point B and insufficient coverage at point C leading to disease spread, this approach, through equipment capacity awareness and priority arbitration, reliably ensured the protection of key targets under limited control resources.

[0071] At the mango planting base in Baise, Guangxi, after the system completed equipment capability perception and priority arbitration according to the aforementioned plan, the cloud-based AI decision-making unit issued the target spraying range, travel path, and spraying amount allocation plan for each sprayer to the collaborative prevention and control execution unit. During a concentrated outbreak of pests and diseases in mid-August 2025, the cloud-based system dispatched sprayer No. 1 to field No. 12 to carry out anthracnose control, and sprayer No. 2 to field No. 15 to carry out thrips control. The planned travel paths of the two machines spatially overlapped at the intersection of the ridges between fields No. 12 and No. 15, and the estimated travel time was less than 40 seconds apart. The conventional spraying operation method at the base, which lacked micro-meteorological monitoring equipment, exposed two prominent problems in this incident. First, the two machines did not coordinate their timing at the intersection. The No. 2 sprayer arrived at the intersection first and continued to move forward, while the No. 1 sprayer arrived later and braked suddenly to wait. As a result, the wheels of the two machines repeatedly ran over an area of ​​about 3 square meters near the intersection, causing the soil around the roots of 4 mango trees to become compacted, and the lateral roots of one of the trees were exposed and broken. Secondly, starting at 1:00 PM that day, the wind speed gradually increased from 2.1 meters per second, reaching a gust of 6.8 meters per second during the operation of both sprayers at 1:35 PM, exceeding the safe spraying wind speed limit of 5.0 meters per second. However, the operation was not stopped, causing the anthrax fungicide droplets sprayed by sprayer No. 1 to drift to the adjacent No. 11 field. This field was in the pre-harvest safety interval period, and the residual amount of drifted pesticide was tested to be 0.18 mg / kg. Although it did not exceed the standard, it posed a quality risk. The thrips control agent sprayed by sprayer No. 2 was deflected by crosswinds, and the actual coverage area was about 4.5 meters away from the target. The amount of pesticide deposited in the core area of ​​point B was only about 55% of the target value, and the control effect was significantly reduced.

[0072] To address the aforementioned issues, the cloud-based AI decision-making unit, while determining the target spraying range, travel path, and spray volume allocation for each sprayer, simultaneously executes the micro-meteorological safety window determination and distributed temporal path coordination scheme described in this system. An integrated micro-weather station mounted on the multimodal sensing unit collects wind speed, wind direction, and rainfall probability data every 5 minutes, transmitting this data to the cloud-based AI decision-making unit via an edge computing box. The cloud maintains a spraying operation safety window determination model, with a preset safe wind speed threshold of 5.0 meters per second and a rainfall probability threshold of 60%. In the aforementioned mid-August event, wind speed data collected before 13:00 fluctuated between 1.8 and 3.5 meters per second, and the rainfall probability fluctuated between 15% and 25%, indicating a sprayable window. The cloud issued operation instructions for both sprayers as planned. At 13:05, the micro-weather station reported a sudden increase in wind speed to 5.6 meters per second and a rainfall probability of 72%. The cloud-based safety window determination model immediately marked the current moment as a non-sprayable window. The cloud sends a task suspension command to sprayers No. 1 and No. 2 simultaneously via the 4G network. After receiving the command, both machines stop spraying and move to the nearest waiting point to shut down and wait.

[0073] Simultaneously, the cloud retrieved hourly forecast data from the China Meteorological Administration for the township where the orchard was located, showing that the wind speed was expected to drop below 3.5 meters per second after 2:30 PM, with the probability of rainfall decreasing to 45%. The earliest executable time for suspending the task, calculated by the cloud, was 2:30 PM, 85 minutes from the current time. In fruit tree pathology, the allowable delay limits for anthracnose and thrips control are 120 minutes and 180 minutes respectively; 85 minutes is within the allowable range, and there is no need to trigger alternative control strategies. At 1:10 PM, the wind speed further increased to 7.2 meters per second, and the probability of rainfall rose to 85%. The cloud-updated forecast data showed that the wind speed drop was postponed to 3:30 PM. At this point, the earliest executable time for suspending the task was 145 minutes from the completion of the diagnosis, exceeding the 120-minute allowable delay limit for anthracnose control. The cloud immediately triggered an alternative control strategy. After querying the equipment capability data, it increased the target light intensity of the insect-attracting lamps for field No. 12 from medium to high to partially replace the combined control effect of fungicides on thrips. It also pushed an emergency drone supplementary control request to the management terminal. The agricultural drones equipped at the base took off at 13:28 and completed emergency anthracnose spraying coverage of field No. 12 at 13:45. At 14:20, the actual wind speed dropped to 3.2 meters per second, and the probability of rainfall dropped to 38%. The cloud re-determined a spraying window and sent resumption instructions to sprayers No. 1 and No. 2.

[0074] Before resuming operations, the cloud detected that the remaining tasks of the two sprayers still overlapped in time and space. The cloud initiated a distributed temporal path coordination mechanism, constructing a temporal path coordination map with sprayer No. 1 and sprayer No. 2 as agents. Each agent exchanged the currently assigned remaining target spraying range and updated travel path through the cloud. The intersection conflict point was detected at the intersection of field ridges at GPS coordinates 23°24′16″N, 106°37′41″E. The remaining power of the two sprayers was 56% for No. 1 and 28% for No. 2. The currently assigned pest and disease severity was an anthracnose severity index of 0.52 for No. 1 and a thrips damage severity index of 0.28 for No. 2. After weighted calculation using severity as the first priority factor and remaining power as the second priority factor, the passage priority score of sprayer No. 1 was higher than that of sprayer No. 2. The consultation mechanism determined that sprayer No. 1 would pass through the intersection first at 14:31, while sprayer No. 2 would wait at a yielding point 4.5 meters south of the intersection. After sprayer No. 1 passed, sprayer No. 2 would pass through the intersection at 14:33. The final sequential execution plan specified that sprayer No. 1's scheduled passage time in field No. 12 was from 14:31 to 14:37, and sprayer No. 2's scheduled passage time in the same field No. 2 was from 14:34 to 14:40. The distance between the two sprayers at the intersection was greater than 15 meters, with no risk of conflict. After both sprayers completed their operations according to the plan, the pesticide deposition in the target areas of fields No. 12 and No. 15 both reached more than 92% of the target value. No drift pesticide residue was detected in adjacent fields during the pre-harvest safety interval. There was no plant crushing damage in the intersection area, and the two sprayers did not experience any mutual waiting delays throughout the entire process.

[0075] At a mango plantation in Baise, Guangxi, after the system completed a closed-loop operation from perception, screening, diagnosis to control execution according to the aforementioned plan, the collaborative control execution unit carried out daily spraying and insect trapping operations, effectively controlling pests and diseases. However, after about three months of continuous operation, technicians discovered a regular anomaly in the historical data stored in the cloud: in the control plans generated by the system, the recommended spraying dosage of fungicide for angular leaf spot continuously increased in multiple control cycles, gradually climbing from an initial average of 1.2 liters per mu to 2.4 liters per mu. Reviewing the diagnostic records of the corresponding fields revealed that the average control rate remained above 92%, and the control effect was stable. However, some fields developed yellowing and fine necrotic spots on the leaf edges 7 to 10 days after spraying. These symptoms were highly similar to the leaf edge necrosis characteristics of angular leaf spot itself. In subsequent diagnosis, the cloud-based CNN-LSTM fusion model misjudged this type of pesticide-induced leaf edge necrosis as a recurrence of angular leaf spot, thus outputting a control plan with a higher spraying dosage. The higher dosage further aggravated the pesticide damage, forming a self-reinforcing cycle of misdiagnosis. Conventional data feedback methods only count the binary result of whether pests and diseases have subsided after control as a supervisory signal for model updates, without quantitatively assessing the secondary pesticide damage that the control program itself may cause. In parameter updates, the model continuously strengthened the positive correlation between high dosage and high control rate. The amount of pesticide used increased by about 45% cumulatively within three months, while the proportion of pesticide-damaged area expanded from less than 1% initially to about 8%.

[0076] To address the aforementioned issues, the cloud-based AI decision-making unit activates the data self-learning feedback and pesticide damage gating penalty mechanism described in this system. After each collaborative control execution unit completes spraying and insect trapping operations, the system sets the post-control evaluation period to 72 hours. At the 72nd hour, the multimodal sensing unit re-collects visible light images and hyperspectral data from the same location. The edge computing box processes this data according to the aforementioned preprocessing and screening procedures before uploading it to the cloud. The cloud-based AI decision-making unit then runs the CNN-LSTM fusion model on this re-collected data again, outputting a pest and disease severity index after control. Taking anthracnose control in field No. 18 in August 2025 as an example, the severity index output by the cloud-based precise diagnosis before control was 0.68. 72 hours later, the severity index output by the re-diagnosis at the same location was 0.09. Because the severity index before control was greater than 0, according to the formula CR= max(0,(S pre -S post ) / S pre The normalized difference between the two was calculated to be 0.87, which means that the pest and disease control rate of this control was 87%, and the control effect was good and far exceeded the preset qualified line of 60%.

[0077] Meanwhile, the cloud-based AI decision-making unit inputs the re-acquired visible light images after the control measures into a pre-trained semantic segmentation model to extract mango leaf regions. Within each segmented leaf pixel region, color and morphological features are analyzed pixel-by-pixel. The yellowing area is identified by the set of pixels whose green channel values ​​are lower than a preset ratio of red and blue channel values. The necrotic spot area is identified by the set of pixels whose grayscale values ​​are lower than a preset dark threshold and are distributed as discrete dots. The leaf curling index is calculated by the degree of deviation between the curvature of the leaf contour edge extracted by semantic segmentation and the standard contour of a normal leaf. The weights of these three indicators are set to 0.4, 0.4, and 0.2, respectively, and the weighted sum is used to obtain the pesticide damage index. In this sample field, the yellowing area ratio is 0.03, the necrotic spot area ratio is 0.01, and the leaf curling index is 0.05. The weighted pesticide damage index is 0.026, which is lower than the preset pesticide damage tolerance threshold of 0.35, indicating that the control measures did not cause significant pesticide damage.

[0078] The data self-learning feedback unit correlated the control rate of 87% and the phytotoxicity index of 0.026 with the following parameters in the control plan: the spray type was difenoconazole, the spraying rate was 1.4 liters per acre, and the pest / disease type was anthracnose. These control plan parameters were reported and recorded by the collaborative control execution unit during operation and uploaded to the cloud-based AI decision-making unit, generating a structured feedback data entry which was then uploaded to the cloud-based AI decision-making unit's feedback database. After one month of accumulation, the feedback database contained approximately 480 structured feedback data entries covering different disease types, pesticide types, and spraying rates. The cloud-based AI decision-making unit periodically used this data to incrementally fine-tune the CNN-LSTM fusion model, introducing a penalty term weighted by the phytotoxicity index into the original multi-task loss function during fine-tuning. The activation condition for the penalty term is that the drug damage index exceeds the tolerance threshold of 0.35. When the drug damage index corresponding to a training sample is lower than 0.35, the weight of the penalty term is zero, and the sample only participates in gradient update with the goal of diagnostic accuracy. When the drug damage index is higher than 0.35, the penalty term is applied to the diagnosis-prevention parameter combination corresponding to the sample with the drug damage index as the coefficient, so that the model gradient is updated in the direction of suppressing such combinations.

[0079] From the approximately 480 feedback data points, 31 samples (approximately 6.5%) with a pesticide damage index exceeding 0.35 were identified as high-dose pesticide operations. These samples corresponded to control measures involving spraying amounts exceeding 2.0 liters per acre or specific pesticides applied in combination with mango young leaves. After introducing a pesticide damage penalty and performing 10 rounds of incremental fine-tuning, the model before and after the fine-tuning was compared and evaluated using the same test set. Before the fine-tuning, the model's recommendation probability for high-dose solutions was approximately 72%, which decreased to approximately 38% after the fine-tuning. The model's probability of misdiagnosing pesticide damage symptoms as disease recurrence was approximately 28%, which decreased to approximately 9% after the fine-tuning. In the following two months, the average spraying amount in the automatically generated control measures decreased from 1.9 liters per acre to 1.5 liters per acre, a reduction of approximately 21%, while the integrated pest management rate remained above 90%. The proportion of pesticide-damaged area gradually decreased from approximately 8% to approximately 1.5%, effectively breaking the previous self-reinforcing cycle of misdiagnosis.

[0080] While the system was operating stably at the Baise mango planting base in Guangxi, it was also deployed and operated independently at the Jinhuang mango orchard in Sanya, Hainan, and the Tainong mango orchard in Yuanjiang, Yunnan. The three orchards span approximately 1100 kilometers, belonging to three distinct climate zones: the Youjiang River Valley dry-hot climate, the tropical maritime monsoon climate, and the dry-hot river valley climate. The main varieties cultivated are Tainong No. 1, Jinhuang mango, and Tainong No. 1, respectively. Soil types vary significantly, ranging from red soil in Baise to coastal sandy loam in Sanya and dry red soil in Yuanjiang. After approximately three months of independent accumulation of structured feedback data in each orchard, the Baise base accumulated approximately 480 data points, the Sanya base approximately 310, and the Yuanjiang base approximately 260, all triggering the federated update threshold. The cloud-based AI decision-making unit initiated the federated learning update process, with the edge computing boxes in each orchard registering as local clients with the central server to participate in this round of federated updates.

[0081] Traditional federated learning aggregation methods, after each client uploads its local model update, directly perform a weighted average using the data sample size of each client as the weight to generate the global model update. This method implicitly assumes that the local data distribution of each client is consistent with the global data distribution, but this assumption clearly does not hold true in the mango pest and disease diagnosis scenario. The Sanya base is located in the tropics, with high temperature and humidity year-round. Anthracnose and powdery mildew are the dominant diseases year-round, while the thrips outbreak period is about 45 days different from that of Baise and Yuanjiang. Due to the dry-hot valley foehn effect, the Yuanjiang base experiences a much higher frequency of angular leaf spot disease than the other two bases, often accompanied by physiological leaf curling caused by extreme dryness. When the Sanya base's local model was updated in the previous round of local fine-tuning due to a concentrated thrips outbreak, causing the model parameters to be updated towards enhancing thrips-related feature responses, the cosine similarity between its update direction and the update vectors of the Baise and Yuanjiang bases, which primarily optimized for anthracnose and angular leaf spot, was only 0.21 and 0.19, respectively. If the traditional federated average is used, the update volume of the Sanya base is diluted in the weighted average by the larger sample size of the Baise and Yuanjiang bases. The local optimization effect of the Sanya base on thrips diagnosis is almost wiped out in the global model, resulting in catastrophic amnesia.

[0082] To address the aforementioned issues, the cloud-based AI decision unit employs the spectral clustering aggregation scheme based on update direction similarity described in this system when performing federated aggregation. This round of federated updates involved three clients: Baise, Sanya, and Yuanjiang. The cloud received local model updates from these clients. The cloud first calculated the cosine similarity between each pair of the three update vectors, constructing a 3x3 model update direction similarity matrix between the clients. The results showed that the cosine similarity between the update vectors of Baise and Yuanjiang was 0.78, indicating that both exhibited similar model optimization trends, primarily focusing on enhancing the responses to anthrax and keratosis features. The similarity between Sanya and Baise was 0.21, and between Sanya and Yuanjiang was 0.19, indicating that Sanya's update direction deviated significantly from the former two.

[0083] Based on this similarity matrix, the cloud platform dynamically clusters participating clients using spectral clustering. Spectral clustering treats the similarity matrix as an adjacency matrix of a graph, calculates the eigenvectors of the Laplacian matrix, and clearly divides the three clients into two clusters in the feature space: Baise and Yuanjiang are assigned to the first cluster, with consistent update directions and convergent optimization goals; Sanya, due to its unique optimization direction, forms a separate second cluster. The cloud platform weights the local model update amounts of the 480 samples from Baise and the 260 samples from Yuanjiang within the first cluster, using their respective sample sizes as weights, to obtain the cluster aggregate update amount for the first cluster. For the 310 samples from Sanya in the second cluster, the cluster aggregate update amount is simply the local model update amount for Sanya itself. Subsequently, the cloud performs a weighted summation of the cluster update amounts of the two clusters, with each cluster having a percentage of the total sample size. The first cluster has 740 samples, accounting for approximately 70.5%, while the second cluster has 310 samples, accounting for approximately 29.5%. This generates a global model update amount, which is then applied to the global CNN-LSTM fusion model to complete this round of global model update. The updated global model parameters are then distributed to the edge computing boxes in the three orchards via an encrypted channel, synchronously updating each local model copy.

[0084] To verify the actual effectiveness of the clustering aggregation scheme, in the same round of federated updates, the cloud simultaneously ran the traditional federated averaging scheme to output another set of global model parameters for comparison. After the update, the global models under the two aggregation schemes were tested using 112 samples from a local test set (not previously used by the Sanya base for training). The test set included 38 anthrax samples, 21 angular leaf spot samples, 18 powdery mildew samples, and 35 thrips damage samples. The global model under the traditional federated averaging scheme achieved an accuracy of approximately 63% in identifying thrips damage samples, a decrease of 26 percentage points compared to approximately 89% after local fine-tuning at the Sanya base, indicating significant catastrophic forgetting. In contrast, the global model under the clustering aggregation scheme achieved an accuracy of approximately 87% in identifying thrips samples in Sanya, only a decrease of 2 percentage points compared to the locally fine-tuned model, demonstrating that diagnostic knowledge related to thrips features was effectively preserved in the second cluster's independent updates from Sanya. On anthrax and keratosis samples shared by Baise and Yuanjiang, the identification accuracy of the two aggregation schemes was basically the same: the cluster aggregation scheme was approximately 93% and 91%, respectively, while the traditional federated averaging scheme was approximately 92% and 90%, respectively. The cluster aggregation scheme, while compatible with the unique data distribution in Sanya, did not sacrifice the common diagnostic capabilities of most clients. Using a total of 458 samples from the cross-domain test set across the three bases, the overall diagnostic accuracy of the global model under the cluster aggregation scheme was approximately 91%, while the traditional federated averaging scheme was approximately 84%. The variance of cross-domain diagnostic performance decreased from approximately 3.8% for the traditional scheme to approximately 1.5%, indicating that cluster aggregation effectively mitigated the impact of cross-domain data distribution shifts on the consistency of the global model.

[0085] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be applied to various fields suitable for the present invention. Other modifications can be readily made by those skilled in the art; therefore, the present invention is not limited to the specific details without departing from the general concept defined by the claims and their equivalents.

Claims

1. A mango pest and disease early diagnosis and control system based on multimodal sensing and AI, characterized in that, include: The multimodal sensing unit is deployed in the mango orchard. Each multimodal sensing unit includes a visible light camera with a resolution of no less than 2 million pixels, a hyperspectral sensor with a working wavelength in the range of 900-1700nm, an air temperature and humidity sensor, and a soil moisture sensor. The edge computing unit communicates with the multimodal sensing unit, receives visible light images, hyperspectral data, air temperature and humidity data, and soil moisture data collected by the multimodal sensing unit, performs noise reduction and normalization preprocessing on the visible light images and hyperspectral data, runs a lightweight AI inference model to perform preliminary screening of pests and diseases on the preprocessed data, and calculates and adjusts the integration time of the hyperspectral sensor in real time based on the light intensity histogram of the mango leaf area in the visible light image with the goal of maximizing the number of effective spectral pixels. The cloud-based AI decision-making unit connects to the edge computing unit via the network, receives visible light images and hyperspectral data that are initially screened by the edge computing unit as suspected pests and diseases, as well as air temperature and humidity data and soil moisture data at the corresponding collection time, and runs a CNN-LSTM fusion model to perform a detailed diagnosis of the type and severity of pests and diseases, generating a control plan that includes the amount of pesticide to be sprayed, the spraying range, and the start and stop instructions for the insect-attracting lamps. In the CNN-LSTM fusion model, CNN is used to extract spatial features, and LSTM is used to extract temporal features. The collaborative prevention and control execution unit, including intelligent sprayers, insect-attracting lamps, and dispatchable drones, is connected to the cloud-based AI decision-making unit to receive prevention and control plans and automatically execute spraying and insect-attracting operations.

2. The mango disease and pest early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, The edge computing unit performs denoising and normalization preprocessing on visible light images and hyperspectral data, specifically including: A hierarchical preprocessing strategy is constructed. First, the fluctuation level of the data collected by the current multimodal sensing unit is evaluated. The fluctuation level is quantified by the mean of the noise standard deviation of each band of the hyperspectral sensor in n consecutive acquisition cycles and the rate of change of pixel values ​​between visible light image frames. When the fluctuation level is lower than the preset threshold, the first processing mode is adopted: the visible light image is subjected to edge-preserving denoising using guided filtering, the hyperspectral data is subjected to denoising method based on minimum noise separation transformation, the denoised hyperspectral data is radiometrically calibrated and corrected according to the current integration time of the hyperspectral sensor, the denoised hyperspectral data is normalized to the [0,1] interval, and the denoised visible light image is normalized to the [0,1] interval after histogram equalization. When the fluctuation level is equal to or higher than the preset threshold, it indicates that there is data quality degradation caused by environmental disturbances. The system then switches to the second processing mode: the visible light image is processed with dehazing enhancement based on dark channel priors and then guided filtering is performed; the hyperspectral data is denoised using a multivariate scattering correction-based method to eliminate spectral baseline drift and scattering noise caused by water vapor and dust scattering on the leaf surface; the denoised hyperspectral data is radiometrically calibrated according to the current integration time of the hyperspectral sensor; and the processed hyperspectral data and visible light image data are then jointly registered in spectral and spatial order before normalization to correct the intermodal spatial shift caused by environmental factors. All hyperspectral data preprocessing is synchronously associated with the optimal integration time parameter of the hyperspectral sensor corresponding to the current acquisition cycle to complete radiometric consistency normalization correction.

3. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, The edge computing unit runs a lightweight AI inference model to perform preliminary screening for pests and diseases on the preprocessed data, specifically including: A dual-branch lightweight feature extraction network and an environment coding module are deployed within the edge computing unit; The preprocessed visible light image is input into the first branch with MobileNetV3 as the skeleton to extract the visible light spatial feature map. The preprocessed hyperspectral data is input into the second branch with a one-dimensional convolutional network as the skeleton to extract the spectral feature vector. At the same time, the environmental coding module maps the air temperature and humidity data and soil moisture data in the current acquisition period into environmental feature vectors. The visible light spatial feature map is concatenated with the spectral feature vector obtained by global average pooling and the environmental feature vector. The concatenation is then input into a classification head consisting of two fully connected layers. The classification head outputs two scores in parallel: a suspected pest and disease score and an environmental stress confidence score. The edge computing unit dynamically adjusts the screening threshold based on the environmental stress confidence score: when the environmental stress confidence score is greater than the preset first threshold, the screening threshold for suspected pests and diseases is increased from the default value T0 to T1, where T1 is greater than T0, in order to suppress false positives in visible light or spectral representations caused by environmental stress from being misjudged as pests and diseases; when the environmental stress confidence score is less than or equal to the first threshold, the default value T0 is maintained as the screening threshold. The suspected pest score is compared with the currently effective screening threshold. If the suspected pest score is greater than the screening threshold, it is determined to be a suspected pest. The suspected pest score, environmental stress confidence score, spectral angular distance, and the corresponding preprocessed visible light image and hyperspectral data are then reported to the cloud. Simultaneously, a normal mango leaf hyperspectral reference library is established at the edge. The spectral angular distance between the current preprocessed hyperspectral data and the corresponding normal samples in the reference library is calculated in the characteristic bands of the 900-1000nm and 1400-1700nm ranges. When the spectral angular distance is higher than the preset second threshold and the environmental stress confidence score is lower than the preset third threshold, it is directly judged as a suspected pest or disease. The suspected pest or disease score, environmental stress confidence score, spectral angular distance, and the corresponding preprocessed visible light image and hyperspectral data are all reported to the cloud, regardless of whether the suspected pest or disease score exceeds the screening threshold.

4. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, The edge computing unit adjusts the integration time of the hyperspectral sensor in real time based on the illumination intensity of the visible light image, specifically including: The edge computing unit uses a lightweight semantic segmentation model deployed on it to extract the mango leaf region from the visible light image, filter out the background and strongly reflective non-leaf regions, and calculate the illumination intensity histogram of all pixels in the mango leaf region. The saturation radiance threshold and dark noise equivalent radiance threshold of the hyperspectral sensor are obtained. Based on the illumination intensity histogram, the integration time that maximizes the proportion of effective pixels in the mango leaf region that is neither saturated nor higher than the dark noise equivalent radiance under the current gain of the hyperspectral sensor is calculated and used as the optimal integration time for the current frame. The edge computing unit also maintains an integration time series window of length k, records the optimal integration time calculated for k consecutive frames, obtains the slope of the light intensity change trend by linear fitting the integration time series, and uses the result of linear fitting to predict the estimated optimal integration time for the next acquisition moment. The estimated optimal integration time is weighted and fused with the optimal integration time of the current frame to obtain the final transmission integration time. The weighting coefficient is dynamically determined based on the goodness of fit of the linear fit. The higher the goodness of fit, the greater the weight of the estimated optimal integration time. The final transmission integration time is written to the hyperspectral sensor before the end of the current acquisition cycle so that it can take effect in the next acquisition cycle.

5. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, The cloud-based AI decision-making unit runs a CNN-LSTM fusion model to perform detailed diagnosis of pest and disease types and severity, specifically including: The CNN-LSTM fusion model includes a cross-modal multi-head attention fusion module, an environmental condition gating module, and a temporal multi-task output module; The CNN-LSTM fusion model also receives the suspected pest and disease scores, environmental stress confidence scores, and spectral angular distance reported by the edge computing units. It concatenates the suspected pest and disease scores and environmental stress confidence scores into an edge prior feature vector, uses the edge prior feature vector as the initial hidden state bias of the temporal encoder, and concatenates the spectral angular distance as an additional input dimension into the multi-level spectral feature vector. The visible light image reported by the edge computing unit is input into the pre-trained ResNet-18 backbone network to extract multi-scale spatial feature maps. The corresponding hyperspectral data is input into a spectral encoder consisting of three consecutive one-dimensional convolutional blocks to extract multi-level spectral feature vectors. Each convolutional block of the spectral encoder outputs feature representations with different spectral resolutions. The cross-modal multi-head attention fusion module uses the highest-level spatial feature map in the multi-scale spatial feature map as the query and the multi-level spectral feature vector as the key and value to calculate the cross-modal attention weight between spatial location and spectral band, and generate a spatial-spectral joint enhanced feature map, so that the model focuses on the pixel position in the visible light image that corresponds to the spatial location of the spectral anomaly region. The spatial-spectral joint enhanced feature map is then concatenated with the multi-level spectral feature vector at the channel level after global average pooling to obtain the multimodal fusion feature vector at the current time. The environmental condition gating module receives air temperature, humidity and soil moisture data reported by the edge computing unit at the current acquisition time. After mapping by the fully connected layer, it generates a set of environmental gating coefficients. The environmental gating coefficients are used to scale each dimension of the multimodal fusion feature vector element by element to suppress feature responses that are highly correlated with the current environmental conditions but unrelated to pests and diseases. The multimodal fusion feature vector after environmental gating is fed into a time encoder composed of two layers of LSTM. The two layers of LSTM maintain a hidden state across time steps, fuse the feature vector at the current time step with the historical hidden state, and output a time-aware feature vector. The temporal multi-task output module includes a pest and disease type classification branch and a severity regression branch. The classification branch takes the temporal-aware feature vector as input and outputs the probability of each pest and disease type. The severity regression branch takes the concatenation result of the temporal-aware feature vector and the embedding vector of the highest probability type of the classification branch as input and outputs the severity index corresponding to the type. The severity index is a weighted combination of the lesion area ratio and the normalized spectral index.

6. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 5, characterized in that, The spatial-spectral joint enhanced feature map is then concatenated with multi-level spectral feature vectors at the channel level after global average pooling. Specifically, this includes: Before global average pooling, the spatial-spectral joint enhanced feature map is divided into g equally sized feature map groups along the channel dimension. The spatial attention weighted average within each feature map group is performed to obtain a g-dimensional group spatial convergence vector. The group spatial convergence vector is then concatenated with the global representation vector obtained after global average pooling of the spatial-spectral joint enhanced feature map to form a spatial feature descriptor with multi-granularity spatial information. Normalized spectral vegetation index is calculated for each level in the multi-level spectral feature vector. The normalized spectral vegetation index is the normalized ratio of the reflectance of the feature band sensitive to chlorophyll to the reflectance of the feature band sensitive to cell structure. The normalized spectral vegetation indices of each level are spliced ​​together to obtain the spectral index feature vector. When concatenating the spatial feature descriptor with the multi-level spectral feature vector at the channel level, the spectral index feature vector is also concatenated to form the multimodal fusion feature vector at the current time. Before concatenation, the multi-level spectral feature vector undergoes adaptive filtering through a learnable channel gate. The channel gate takes the spatial feature descriptor as a conditional input and outputs the retention probability of each channel of the multi-level spectral feature vector. Only channels with retention probabilities higher than the preset gate threshold are concatenated with the spatial feature descriptor and the spectral index feature vector.

7. The mango disease and pest early diagnosis and control system based on multimodal sensing and AI according to claim 5, characterized in that, The environmental condition gating module receives air temperature, humidity, and soil moisture data reported by the edge computing unit at the current acquisition time. After mapping by the fully connected layer, it generates a set of environmental gating coefficients, specifically including: The environmental condition gating module also receives the air temperature and humidity sequence and soil moisture sequence of m consecutive collection times before the current collection time reported by the edge computing unit. Based on the air temperature and humidity sequence, it calculates the duration of leaf wetness and effective accumulated temperature within a preset time. Based on the soil moisture sequence, it calculates the soil moisture deficit index. It then concatenates the duration of leaf wetness, effective accumulated temperature, and soil moisture deficit index with the air temperature and humidity and soil moisture data at the current collection time, and integrates them with the environmental stress confidence score reported by the edge computing unit. The results are then input into the fully connected layer to generate the initial environmental gating coefficient. In the environmental condition gating module, a learnable pest-environment association matrix is ​​pre-set. Each row of the pest-environment association matrix corresponds to the environmental parameter response pattern embedding of a candidate mango pest type under typical occurrence conditions. The initial environmental gating coefficient is used as a query and scaled dot product attention is calculated with the pest-environment association matrix to obtain attention weights that are softly aligned with each candidate pest type. The attention weights are then used to perform weighted summation on each row of the pest-environment association matrix to generate an environmental gating bias vector that is aware of pest conditions. The initial environmental gating coefficient is added element-wise to the environmental gating bias vector of the pest and disease condition perception to obtain the final environmental gating coefficient. The final environmental gating coefficient is then used to scale each dimension of the multimodal fusion feature vector element-wise.

8. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, The cloud-based AI decision-making unit generates a control plan that includes spraying dosage, spraying range, and start / stop instructions for the insect-attracting lamps, specifically including: After obtaining the types and severity of pests and diseases, the cloud-based AI decision-making unit acquires the IoT device status data of each smart sprayer and insect-attracting lamp in the collaborative prevention and control execution unit. The IoT device status data includes at least the remaining liquid, battery level, current location coordinates, and nozzle type of each smart sprayer, as well as the current working status, light intensity level, and cumulative working time of each insect-attracting lamp. A two-layer bipartite graph is constructed, consisting of an equipment capability layer and a prevention and control requirement layer. The equipment capability layer uses each available intelligent sprayer, insect-attracting lamp, and schedulable drone as nodes, while the prevention and control requirement layer uses the occurrence locations of each pest and disease determined by detailed diagnosis and the required prevention and control actions as nodes. The required prevention and control actions include the type of spraying, the amount of spraying required based on the pest and disease severity index, the start and stop requirements of the insect-attracting lamp, and the requirement of drone supplementary prevention. The edge weights of the two-layer bipartite graph are determined by the weighted sum of the spatial distance between the equipment and the occurrence location, the matching degree of the current status of the equipment, and the urgency of prevention and control. The minimum cost matching algorithm with capacity constraints is run on a bipartite graph to minimize the total scheduling cost. Under the constraints of the upper limit of the spraying capacity of each intelligent sprayer and the effective coverage radius of each insect-attracting lamp, the algorithm solves the target spraying range, travel path and spraying amount allocation of each intelligent sprayer, as well as the start and stop commands and target light intensity levels of each insect-attracting lamp. When multiple pest and disease locations compete for the same intelligent sprayer and the remaining liquid in the sprayer cannot simultaneously meet the spraying needs of all locations, a prevention and control priority index is calculated based on the severity of the pests and diseases and the risk of spread at each location. The liquid is then allocated to the top k locations with the highest prevention and control priority index. For the remaining locations, alternative sprayers are rescheduled or they are marked as areas requiring supplementary prevention and a supplementary prevention reminder is generated and pushed to the management personnel.

9. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 8, characterized in that, When the cloud-based AI decision-making unit calculates the target spraying range, travel path, and spraying quantity allocation for each intelligent sprayer, it also includes: Receive orchard micro-meteorological data collected in real time by a multimodal sensing unit or an external micro-weather station. The micro-meteorological data includes at least wind speed, wind direction, and rainfall probability. Micrometeorological data is input into a preset spraying operation safety window determination model. When the wind speed exceeds the preset safety threshold or the probability of rainfall is higher than the preset rainfall threshold, the current moment is determined as an unsprayable window, and the affected spraying tasks are suspended. At the same time, based on the window prediction of meteorological forecast data, the earliest executable time and effective spraying time window of the suspended tasks are calculated. If the predicted effective spraying time window exceeds the upper limit of the allowable delay for pest and disease control, alternative control strategies are triggered, including increasing the insect-attracting light intensity level of the corresponding area of ​​the insect-attracting lamp or dispatching drones for emergency supplementary control. When multiple executable spraying tasks overlap in time and space, a distributed temporal path coordination graph is constructed with each sprayer as an agent. In the coordination graph, each agent exchanges spatiotemporal occupancy information and detects path intersection conflicts based on the currently allocated target spraying range and travel path. The spatiotemporal priority negotiation mechanism determines the passage sequence of each sprayer at the intersection point. The spatiotemporal priority negotiation mechanism uses the severity of pests and diseases and the remaining power of the sprayer as joint priority factors and outputs a conflict-free temporal execution plan. The temporal execution plan includes the predetermined passage time and avoidance waiting point for each sprayer on each ridge.

10. The mango pest and disease early diagnosis and control system based on multimodal sensing and AI according to claim 1, characterized in that, It also includes a data self-learning feedback unit, which collects data on pest and disease control rate and pesticide damage after the collaborative prevention and control execution unit executes the prevention and control plan, and feeds it back to the cloud AI decision-making unit; The cloud-based AI decision-making unit uses the received feedback data to update the CNN-LSTM fusion model.