Crop disease and pest early identification and positioning method based on visual-spectrum-environment multi-modal feature fusion

CN122618478BActive Publication Date: 2026-09-18SHANDONG DEEP BLUE ZHIPU DIGITAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611097528.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-18
Estimated Expiration
2046-07-23

AI Technical Summary

Technical Problem

[0003]针对上述问题,本发明提出基于视觉-光谱-环境多模态特征融合的作物病虫害早期识别与定位方法,该基于视觉-光谱-环境多模态特征融合的作物病虫害早期识别与定位方法通过三个并行分支分别提取视觉、光谱、环境三类模态特征,全面覆盖作物宏观表型、微观生理生化及环境上下文信息,同时引入跨模态自适应对齐模块,通过互信息损失与门控注意力权重,消除异构数据之间的语义鸿沟与尺度差异,实现多模态特征的高效融合,提升了特征表达的全面性与准确性,为病虫害早期识别与定位提供了可靠的特征支撑,解决了现有技术单模态特征失真、信息不全面的问题

Benefits of technology

1、本发明通过三个并行分支分别提取视觉、光谱、环境三类模态特征,全面覆盖作物宏观表型、微观生理生化及环境上下文信息,同时引入跨模态自适应对齐模块,通过互信息损失与门控注意力权重,消除异构数据之间的语义鸿沟与尺度差异,实现多模态特征的高效融合,提升了特征表达的全面性与准确性,为病虫害早期识别与定位提供了可靠的特征支撑,解决了现有技术单模态特征失真、信息不全面的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618478B_ABST
    Figure CN122618478B_ABST
Patent Text Reader

Abstract

The application provides a crop disease and pest early identification and positioning method based on visual-spectrum-environment multi-modal feature fusion, relates to the field of agricultural intelligent detection technology, and comprises the following steps: collecting three types of modal data of crop high-resolution RGB images, multispectral reflectivity data and real-time environment parameters; constructing a multi-branch parallel feature extraction network to extract features of the three types of modal data respectively to obtain visual features, spectral features and environment context features; and embedding space alignment and fusion of the three types of modal features are carried out through a cross-modal adaptive alignment module; the application extracts three types of modal features of vision, spectrum and environment through three parallel branches, comprehensively covers crop macro phenotype, micro physiological and biochemical and environment context information, simultaneously introduces a cross-modal adaptive alignment module, eliminates semantic gap and scale difference between heterogeneous data through mutual information loss and gate attention weight, and realizes efficient fusion of multi-modal features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural intelligent detection technology, and in particular to a method for early identification and location of crop diseases and pests based on the fusion of visual-spectral-environmental multimodal features. Background Technology

[0002] With the rapid development of agricultural modernization, the early and accurate identification and location of crop diseases and pests has become crucial for ensuring agricultural yields, reducing pesticide overuse, and achieving green agriculture. Currently, crop disease and pest identification and location technologies are mainly divided into two categories: traditional manual identification and AI-based automated identification. Traditional manual identification relies on growers' experience, resulting in low efficiency, strong subjectivity, and an inability to achieve large-scale field monitoring and early warning. AI-based automated identification technologies often use single-modal data (such as visual images) for modeling, extracting crop phenotypic features through models such as convolutional neural networks to achieve disease and pest identification and location. Some technologies combine multispectral data to improve identification accuracy. Meanwhile, convenient identification tools based on mobile phone photography have emerged, achieving preliminary disease and pest identification through image database matching, and have already been applied in some crop cultivation. In addition, existing technologies also include threshold detection methods for spectral data used for crop stress state assessment, and YOLO series networks for target detection used for disease and pest area location. However, existing technologies have limitations in single-modal feature extraction. Visual features are easily distorted by environmental factors such as field lighting and shading, while spectral features lack environmental context support and cannot fully reflect crop growth status and pest and disease stress. Traditional spectral threshold detection uses a static baseline, which cannot adapt to individual crop differences and growth cycle changes, making it difficult to achieve accurate early warning of pest and disease incubation periods. Early symptoms of pests and diseases are "latent," and subtle physiological changes are difficult to capture by conventional feature extraction methods, resulting in low early identification accuracy. Identification and localization are mostly independent modules, which have a disconnect, resulting in low localization accuracy and inability to meet the needs of early pest and disease localization in small areas. There is a lack of effective anti-interference mechanisms in complex field environments, and the model has poor robustness in scenarios such as changes in light, temperature and humidity fluctuations, and differences in crop growth, resulting in unstable identification and localization accuracy. Therefore, this invention proposes a crop pest and disease early identification and localization method based on visual-spectral-environmental multimodal feature fusion to solve the problems existing in the prior art. Summary of the Invention

[0003] To address the aforementioned issues, this invention proposes a method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion. This method extracts visual, spectral, and environmental modal features through three parallel branches, comprehensively covering macroscopic phenotypes, microscopic physiological and biochemical information, and environmental context. Simultaneously, a cross-modal adaptive alignment module is introduced, using mutual information loss and gated attention weights to eliminate semantic gaps and scale differences between heterogeneous data, achieving efficient fusion of multimodal features. This improves the comprehensiveness and accuracy of feature representation, providing reliable feature support for early identification and localization of diseases and pests, and solving the problems of single-modal feature distortion and incomplete information in existing technologies.

[0004] To achieve the objectives of this invention, the invention is implemented through the following technical solution: a method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion, comprising the following steps: S1: Acquire three types of modal data: high-resolution RGB images of crops, multispectral reflectance data, and real-time environmental parameters; S2: Construct a multi-branch parallel feature extraction network to extract features from three types of modal data, obtaining visual features, spectral features, and environmental context features; S3: The three types of modal features are embedded spatially aligned and fused through a cross-modal adaptive alignment module to obtain multimodal fused features; S4: Employ temporal differential attention mechanism to enhance early stress features of multimodal fusion features, amplifying small early physiological changes in crops; S5: Construct a latent period early warning model based on spectral band shift and physiological baseline to determine the latent period of crop diseases and pests; S6: Input the enhanced fusion features into the integrated recognition-localization network, and simultaneously output the pest and disease category and precise location information; S7: Through an adaptive anti-interference optimization mechanism for complex field environments, the model's recognition and positioning accuracy is ensured to remain stable in complex field environments.

[0005] Further improvements are made in S2, where the multi-branch parallel feature extraction network includes three parallel branches: the visual feature branch uses an improved lightweight convolutional neural network to extract macroscopic phenotypic features such as crop texture, shape, and color from high-resolution RGB images. The lightweight convolutional neural network is either EfficientNet or MobileNetV4, which is improved by adjusting the size and number of convolutional kernels and introducing depthwise separable convolution and attention mechanisms to reduce parameter redundancy while improving feature extraction accuracy; the spectral feature branch uses a one-dimensional convolutional neural network or Transformer model to process green light, red light, red edge, and near-infrared key band reflectance data acquired by multispectral sensors, extracting microscopic spectral features related to crop chlorophyll content, water content, and nitrogen levels, focusing on capturing narrowband features sensitive to early stress; the environmental feature branch uses a fully connected network to process real-time environmental parameters such as temperature, humidity, light intensity, and soil conductivity, and extracts environmental context features by performing nonlinear mapping of environmental parameters through hidden layer activation functions, with the ReLU function being used as the activation function.

[0006] A further improvement lies in the following: In S3, the implementation process of the cross-modal adaptive alignment module is as follows: First, the mutual information loss between visual features and spectral features is calculated to measure the correlation between the two types of modal features. The calculation formula is as follows: Where I(X,Y) represents the mutual information between visual feature X and spectral feature Y, H(X) represents the information entropy of visual feature X, H(Y) represents the information entropy of spectral feature Y, and H(X,Y) represents the joint information entropy of visual feature X and spectral feature Y; then, the gating attention weights are calculated using environmental features, and the visual and spectral features are weighted and adjusted. The weight calculation formula is as follows: Where W is the gating attention weight vector, The sigmoid activation function is used, We is the environmental feature weight matrix, E is the environmental context feature vector, and be is the bias term. Finally, the weighted visual features, spectral features, and environmental features are concatenated and fused to obtain the multimodal fusion feature. ,in This is a multimodal fusion feature vector.

[0007] A further improvement lies in the following: During the calculation of mutual information loss, the specific formula for calculating information entropy H(X) is as follows: Where n is the feature dimension of visual feature X, p(x) i Let be the probability distribution of the i-th dimension of the visual feature; the formula for calculating the joint information entropy H(X,Y) is: Where m is the feature dimension of the spectral feature Y, p(x i ,y j) represents the joint probability distribution of the i-th dimension of visual features and the j-th dimension of spectral features.

[0008] A further improvement lies in the following: In S4, the implementation process of the temporal differential attention mechanism is as follows: Multiple frames of multimodal data from the same plant or region are continuously collected, denoted as t0, t1, t2, ..., tn, where ti represents the multimodal data at time i; the multimodal fusion feature difference between adjacent time points is calculated using the following formula: ,in, Let Ft be the difference signal of the fused features at adjacent time points, Ft be the multimodal fused feature at time t, and Ft-1 be the multimodal fused feature at time t-1. A multi-head attention mechanism is used to perform weighted fusion of the difference features and the absolute features at the current time point. The fusion formula is as follows: ,in, To enhance early stress characteristics, This represents a multi-head attention calculation function that uses the current time-time features as the query vector and the difference features as the key and value vectors. It achieves feature enhancement through attention weight allocation, focusing on non-natural fluctuation feature channels in the time dimension.

[0009] Further improvements include: the multi-head attention mechanism includes 4-8 attention heads, each of which independently calculates attention weights. The outputs of each attention head are normalized and concatenated, and then the final enhanced features are obtained through a fully connected layer. The acquisition frequency of time-series multi-frame data is 1 frame every 2-4 hours, ensuring that small physiological changes in the early stages of crops can be captured, while avoiding data redundancy.

[0010] Further improvements are made in S5, where the latent period early warning model based on spectral band shift and physiological baseline adopts a dynamic construction method for crop health spectral baselines. The specific implementation steps are as follows: 72 hours of continuous hyperspectral data are collected during the crop's healthy period; individualized health spectral baselines are established using statistical methods, covering a 95% confidence interval; the baseline calculation formula is: ,in wavelength The healthy spectral baseline range, wavelength The mean of hyperspectral reflectance, wavelength The standard deviation of hyperspectral reflectance is used, with 1.96 being the coefficient corresponding to the 95% confidence interval; key stress bands are identified: 780–795nm, 676–745nm, and 550–570nm; early warning criteria are set: when the spectral reflectance of a pixel in a key stress band exceeds the baseline confidence interval for 24 consecutive hours, and the environmental stress index increases, it is determined to be a latent disease. The formula for calculating the environmental stress index is: Where ESI is the environmental stress index, T is the difference between temperature and the suitable range, H is the difference between humidity and the suitable range, L is the difference between light intensity and the suitable range, and S is the difference between soil electrical conductivity and the suitable range. These are the weighting coefficients for each environmental parameter, and their sum is 1.

[0011] Further improvements are made in the following aspects: In S6, the integrated identification-localization network architecture simultaneously inputs the enhanced multimodal fusion features into the identification branch and the localization branch, achieving simultaneous inference and output of results. The identification branch adopts an improved classifier composed of fully connected layers and Softmax layers. Combined with the fused feature vector, it optimizes the classification loss function to achieve accurate identification of pest and disease categories, with a focus on adapting to the classification of early-stage minor pests and diseases. The localization branch adopts an improved YOLO network, introducing a feature pyramid module to enhance the localization ability of small areas and minor pests and diseases, outputting the precise coordinates, area, and bounding box of the pest and disease area. The improved YOLO network optimizes the label allocation strategy and detection head structure, adopts a TAL dynamic matching strategy to achieve balanced classification and regression, and adopts a simplified ET-Head to improve detection accuracy. A global contextual attention mechanism is introduced into the network to collaboratively optimize the identification results and localization results, thereby improving localization accuracy.

[0012] Further improvements are made in the following aspects: The collaborative optimization process of the integrated recognition-localization network is as follows: the category confidence output by the recognition branch is used as an important basis for the global context attention weight, and the feature map of the localization branch is adjusted by weighting. The higher the category confidence, the greater the localization weight of the corresponding region. At the same time, the bounding box information output by the localization branch is fed back to the recognition branch to correct the feature extraction focus of the recognition branch, so as to achieve bidirectional optimization of recognition and localization results and keep the localization error within 5%.

[0013] Further improvements are made in S7, which includes an adaptive anti-interference optimization mechanism for complex field environments, encompassing both feature-level and model-level optimizations. At the feature level, an environmental adaptive adjustment module is introduced during multimodal fusion. Based on real-time collected environmental parameters, the weights and extraction strategies of each modality's features are dynamically adjusted. In strong or weak light environments, the weights of spectral features are enhanced to compensate for visual feature distortion; in high temperature and high humidity environments, the weights of environmental features are enhanced to assist in correcting spectral feature biases. At the model level, a dynamic incremental training module is constructed. Using pest and disease data measured in different field environments, the model is incrementally trained. Mini-batch gradient descent is used to update model parameters, continuously optimizing the model's adaptability to complex environments. Simultaneously, a robust loss function is introduced to reduce the impact of environmental interference on model inference. The robust loss function is: , where L robust For robust loss; L base It is the sum of the base loss, cross-entropy loss, and coordinate loss; The penalty coefficient is... F is the mean absolute error function. true For the true feature vector, F pred To predict the feature vector.

[0014] The beneficial effects of this invention are as follows: 1. This invention extracts visual, spectral, and environmental modal features through three parallel branches, comprehensively covering macroscopic phenotypes, microscopic physiological and biochemical information, and environmental contextual information of crops. At the same time, it introduces a cross-modal adaptive alignment module, which eliminates semantic gaps and scale differences between heterogeneous data through mutual information loss and gated attention weights, achieving efficient fusion of multimodal features, improving the comprehensiveness and accuracy of feature expression, providing reliable feature support for early identification and localization of pests and diseases, and solving the problems of single-modal feature distortion and incomplete information in existing technologies.

[0015] 2. This invention establishes individualized health spectral baselines through a dynamic construction method of crop health spectral baselines. Combined with the identification of key stress bands and analysis of environmental stress index, it achieves accurate determination of the incubation period of pests and diseases. Compared with traditional methods, it can detect latent diseases 3-7 days earlier, effectively solving the problem that existing technologies cannot accurately predict the incubation period of pests and diseases, and buying valuable time for early prevention and control.

[0016] 3. This invention amplifies subtle physiological changes in crops in their early stages by calculating the differential signal of multimodal fusion features at adjacent time points. Combined with a multi-head attention mechanism, it focuses on non-natural fluctuation feature channels in the time dimension, enabling reliable identification of subclinical symptoms. This significantly improves the accuracy of identifying early minor pests and diseases, and makes up for the shortcomings of existing technologies in identifying early latent symptoms.

[0017] 4. This invention achieves "one-time reasoning and simultaneous output" by simultaneously inputting multimodal fusion features into the recognition and localization branches. By combining an improved classifier and an improved YOLO network and introducing a global contextual attention mechanism to achieve collaborative optimization, the localization error is controlled within 5%. This enables precise localization of small areas of early pests and diseases, providing accurate location information for precise field control and improving control efficiency and targeting.

[0018] 5. This invention effectively resists interference from complex field environments such as changes in light intensity, temperature and humidity fluctuations, and shading by using feature-level environmental adaptive adjustment and dynamic incremental training and the introduction of robust loss functions at the model level. This allows the model to maintain an accuracy rate of over 90% and a stable positioning accuracy under different environmental conditions, solving the problem of unstable positioning accuracy in complex field environments in existing technologies and meeting the needs of practical field applications. Attached Figure Description

[0019] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0020] To enhance understanding of the present invention, the present invention will be further described in detail below with reference to embodiments. These embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.

[0021] Example 1 according to Figure 1 As shown in the figure, this embodiment proposes a method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion, which is used to achieve multi-branch parallel feature extraction and cross-modal adaptive alignment. The specific steps are as follows: Data Acquisition: High-resolution RGB images of crops (1920×1080 pixels) were acquired using a high-definition camera at a rate of one frame every 3 hours. Reflectance data of crops in four key bands—green (520–560 nm), red (620–680 nm), red edge (700–740 nm), and near-infrared (780–850 nm)—were acquired using a multispectral sensor at a sampling interval of one hour, with a reflectance data accuracy of 0.001. Environmental sensors were used to acquire field temperature (measurement range 0–50℃, accuracy ±0.1℃), humidity (measurement range 0–100%RH, accuracy ±1%RH), light intensity (measurement range 0–100000 lx, accuracy ±10 lx), and soil conductivity (measurement range 0–20 mS / cm, accuracy ±0.01 mS / cm), at a rate of one frame every hour. The subjects of the study were rice plants in paddy fields, collected 10–50 days after rice sowing, covering three stages: the healthy period, the incubation period, and the early stage of disease. A total of 10,000 RGB images, 12,000 sets of multispectral data, and 12,000 sets of environmental parameters were collected.

[0022] Visual feature extraction branch implementation: An improved EfficientNet-B0 is used as the visual feature extraction network. Improvements include: adjusting the kernel size of the first convolutional layer from 3×3 to 5×5 to enhance the ability to capture crop edge textures; adding a BN layer and ReLU activation function after each convolutional layer to reduce overfitting; and introducing a channel attention mechanism (SE-Net) to assign weights to the extracted feature channels, enhancing the expression of useful features. The network input is an RGB image, and the output is a 256-dimensional visual feature vector. Through multiple iterations of training, the visual feature extraction accuracy reaches over 95%.

[0023] Spectral Feature Branch Implementation: A one-dimensional convolutional neural network (1D-CNN) is used as the spectral feature extraction network. The network structure includes three convolutional layers, two pooling layers, and one fully connected layer. The first convolutional layer uses 16 1×3 convolutional kernels with a stride of 1 and the ReLU activation function. The second convolutional layer uses 32 1×5 convolutional kernels with a stride of 1 and the ReLU activation function. The third convolutional layer uses 64 1×7 convolutional kernels with a stride of 1 and the ReLU activation function. The pooling layer uses max pooling with a kernel size of 1×2 and a stride of 2. The fully connected layer outputs a 128-dimensional spectral feature vector. The input is reflectance data from four key bands. Through training and optimization, the spectral features accurately reflect the changes in chlorophyll, water, and nitrogen in crops, with the feature extraction error controlled within 5%.

[0024] Environmental feature branch implementation: A fully connected network is used, containing three hidden layers. The input layer is a 4-dimensional environmental parameter vector (temperature, humidity, light intensity, soil conductivity). The first hidden layer has 64 neurons with ReLU activation function; the second hidden layer has 32 neurons with ReLU activation function; and the third hidden layer has 16 neurons with ReLU activation function. The output layer is a 16-dimensional environmental context feature vector. By normalizing the environmental parameters (mapping the parameters to the [0,1] interval), the network training efficiency and feature extraction accuracy are improved, and the correlation between environmental features and crop growth status reaches over 88%.

[0025] Cross-modal adaptive alignment and fusion: First, the mutual information loss between visual features (X) and spectral features (Y) is calculated, where X is a 256-dimensional vector and Y is a 128-dimensional vector. By statistically analyzing the probability distribution of the feature dimensions, H(X) = 8.2, H(Y) = 7.5, and H(X,Y) = 13.8 are obtained. Substituting these values ​​into the mutual information formula I(X,Y) = 8.2 + 7.5 - 13.8 = 1.9, the mutual information loss is minimized through backpropagation, increasing the correlation between the two modal features to over 90%. Then, the gated attention weights W are calculated, where We is a 128×16 weight matrix, E is a 16-dimensional environmental feature vector, and be is a 128-dimensional bias term. W is obtained as a 128-dimensional weight vector through the Sigmoid activation function. Finally, the weighted visual features (W·X), spectral features ((1-W)·Y), and environmental features (E) are concatenated to obtain a 384-dimensional multimodal fusion feature F. fusion This enables cross-modal alignment and fusion.

[0026] Example 2 according to Figure 1As shown in the figure, this embodiment proposes an early identification and localization method for crop diseases and pests based on the fusion of visual-spectral-environmental multimodal features, which is used to achieve accurate early warning of the incubation period of crop diseases and pests. Taking rice blast disease early warning in paddy fields as an example, the specific steps are as follows: Construction of a healthy spectral baseline: Hyperspectral data were continuously collected for 72 hours 18 days after rice sowing (healthy period), with a 1-hour interval between collections, resulting in 72 sets of hyperspectral data covering the wavelength range of 500–1000 nm. Emphasis was placed on three key stress bands: 780–795 nm (water imbalance), 676–745 nm (chlorophyll degradation), and 550–570 nm (carotenoid changes). Statistical analysis was performed on the reflectance data at each wavelength, calculating the mean μ(λ) and standard deviation σ(λ) for each wavelength. For example, at 680 nm, μ(λ) = 0.32 and σ(λ) = 0.03. Substituting these values ​​into the baseline formula B(λ) = 0.32 ± 1.96 × 0.03, the healthy baseline range for this wavelength was [0.2612, 0.3788]. This process was repeated to construct an individualized healthy spectral baseline covering all key bands, with a 95% confidence interval.

[0027] Monitoring of key stress bands: Real-time acquisition of hyperspectral data of rice, focusing on extracting reflectance data of three key stress bands. Reflectance changes of each pixel are monitored using a sliding window (24-hour window size) to determine whether they exceed the healthy baseline confidence interval. For example, if the reflectance of a pixel at 780 nm wavelength is 0.42, 0.43, 0.41, ..., 0.44 for 24 consecutive hours, all exceeding the healthy baseline range [0.35, 0.40] for that wavelength, then it is preliminarily determined that the pixel is under water imbalance stress.

[0028] Environmental stress index calculation: Real-time collection of field environmental parameters to determine the suitable environmental range for rice growth: temperature 25–30℃, humidity 70–85%RH, light intensity 30000–80000 lx, and soil electrical conductivity 1.5–3.5 mS / cm. The deviation of each environmental parameter from the suitable range is calculated. For example, at a certain moment, if the temperature is 32℃, the deviation is T = 32 - 30 = 2℃; if the humidity is 65%RH, the deviation is H = 70 - 65 = 5%RH; if the light intensity is 25000 lx, the deviation is L = 30000 - 25000 = 5000 lx; ​​and if the soil electrical conductivity is 1.2 mS / cm, the deviation is S = 1.5 - 1.2 = 0.3 mS / cm. Set the weighting coefficients α=0.3, β=0.2, γ=0.3, and δ=0.2, and substitute them into the Environmental Stress Index formula ESI=0.3×2 + 0.2×5 + 0.3×5000 / 10000 + 0.2×0.3=0.6+1.0+0.15+0.06=1.81. When ESI is greater than 1.5, the environmental stress index is considered to have increased.

[0029] Latency period determination: When a pixel meets two conditions: first, the spectral reflectance of the key stress band exceeds the confidence interval of the healthy baseline for 24 consecutive hours; second, the environmental stress index (ESI) is ≥1.5 (rising), then the rice plant corresponding to that pixel is determined to be in the latent period of rice blast. Through field measurements in rice fields, a total of 500 rice plants were monitored, of which 456 plants were actually in the latent period of rice blast. This model detected 416 plants, achieving a detection rate of 91.2%, which is 37% higher than the traditional static threshold detection method (detection rate 54.2%), and detects latent diseases 3–7 days earlier.

[0030] Example 3 according to Figure 1 As shown in the figure, this embodiment proposes a method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion, which is used to enhance the latent features of early stress in crops. Taking the enhancement of early (latent) features of rice blast disease as an example, the specific steps are as follows: Temporal multimodal data acquisition: Multi-frame multimodal data of the same rice plant were continuously acquired at a frequency of 1 frame every 3 hours, for a total of 10 frames, denoted as t0 to t9. Each frame of data contains RGB image, multispectral reflectance data, and environmental parameters. Multimodal fusion features F0 to F9 of each frame of data were extracted using the method in Example 1. Each fusion feature is a 384-dimensional vector.

[0031] Temporal difference feature calculation: Calculate the difference signal ΔF of the fused features of adjacent time points, for example, ΔF1=F1-F0, ΔF2=F2-F1, ..., ΔF9=F9-F8, where ΔF is a 384-dimensional vector. The difference signal can amplify small physiological changes in rice in the early stage, such as the incubation period of rice blast disease, when rice chlorophyll is slightly degraded, resulting in small changes in spectral features. After difference calculation, this change is amplified, making it easier for the model to capture.

[0032] Multi-head attention weighted fusion: A multi-head attention mechanism with 6 attention heads is used to perform weighted fusion of the current fusion feature Ft and the corresponding difference feature ΔFt. Ft is used as the query vector (Q), and ΔFt is used as the key vector (K) and value vector (V). The attention weight of each attention head is calculated using the following formula: Where dk is the dimension of the key vector (384-dimensional), the attention weights are normalized using the softmax function, focusing on channels with significant changes in the differential features. The outputs of the six attention heads are normalized (using the RMSnorm normalization method), and then concatenated through a fully connected layer to obtain the enhanced early stress feature F. enhanced The dimensions remain 384.

[0033] Feature enhancement effect verification: By comparing the features before and after enhancement, the enhanced features can clearly capture the subtle physiological changes during the latent period of rice. The correlation between the features and early stress of rice blast increased from 72% to 89%. When the enhanced features were input into the recognition model, the accuracy of early rice blast recognition increased from 68% to over 85%, effectively solving the problem of difficulty in recognizing early latent features. At the same time, the noise reduction characteristics of differential attention were used to eliminate the interference of environmental noise on the features, and the stability of feature extraction was improved by more than 30%.

[0034] Example 4 according to Figure 1 As shown in the figure, this embodiment proposes a method for early identification and localization of crop diseases and pests based on the fusion of visual-spectral-environmental multimodal features, which is used to achieve integrated accurate identification and localization of diseases and pests. Taking the early identification and localization of rice blast disease as an example, the specific steps are as follows: Network Architecture: The integrated recognition-localization network takes the multimodal fusion features obtained in Example 1 (enhanced by Example 3) as input and includes two parallel branches: a recognition branch and a localization branch, sharing a feature extraction backbone. The recognition branch uses an improved classifier consisting of two fully connected layers and one softmax layer. The first fully connected layer maps the 384-dimensional fusion features to 128 dimensions, with ReLU as the activation function; the second fully connected layer maps to 10 dimensions (corresponding to 10 common rice diseases and pests), and the softmax layer outputs the confidence scores for each category. The localization branch uses an improved YOLOv8 network, introducing a Feature Pyramid (FPN) module to map the fusion features to feature maps of different scales (8×8, 16×16, 32×32), enhancing the localization capability for small areas and minor diseases and pests; optimizing the label allocation strategy by adopting a TAL dynamic matching strategy to solve the problem of imbalanced classification and regression; simplifying the ET-Head structure by removing time-consuming task interaction feature modules to improve inference speed while ensuring localization accuracy.

[0035] Global contextual attention collaborative optimization: The category confidence score output by the recognition branch is used as the global contextual attention weight to adjust the feature map of the localization branch. For example, when the confidence score of the recognition branch outputting rice blast is 0.92 (high confidence), the weight of the localization feature map of the corresponding region is increased, focusing on optimizing the bounding box prediction of that region. At the same time, the bounding box information (coordinates, area) output by the localization branch is fed back to the recognition branch to correct the feature extraction focus of the recognition branch, focusing on the regional features within the bounding box, reducing background interference, and achieving bidirectional collaborative optimization of recognition and localization results.

[0036] Network Training and Validation: Training was conducted using measured data from paddy fields. The training set contained 8000 enhanced multimodal fusion features, and the validation set contained 2000 features. The batch size was 32, and the training run consisted of 100 epochs. The Adam optimizer was used with a learning rate of 0.001. After training, validation tests were performed on 1000 rice plants, including 200 plants with early rice blast (lesion area less than 0.5 cm²). 2 The model achieves an accuracy rate of 88%, with a localization error controlled within 4.2%. The overlap ratio (IOU) between the output bounding box and the actual lesion area is above 0.85, realizing "one-time inference, simultaneous output of recognition and localization results." This solves the problems of disconnect between recognition and localization and low localization accuracy in small areas in the early stages of existing technologies. Furthermore, the model's inference speed reaches 30 frames per second, meeting the needs of real-time field monitoring.

[0037] Example 5 according to Figure 1As shown in the figure, this embodiment proposes a method for early identification and localization of crop diseases and pests based on the fusion of visual-spectral-environmental multimodal features, which is used to improve the robustness of the model in complex field environments. The specific steps are as follows: Feature-level anti-interference optimization: An environmental adaptive adjustment module is introduced during multimodal fusion to collect environmental parameters such as field light intensity, temperature, and humidity in real time, and dynamically adjust the weights of each modal feature based on these parameters. Weight adjustment rules are set as follows: When light intensity is <10000 lx (weak light) or >90000 lx (strong light), the weight of spectral features is adjusted from 0.3 to 0.5, the weight of visual features is adjusted from 0.5 to 0.3, and the weight of environmental features remains unchanged at 0.2 to compensate for the distortion of visual features under weak / strong light conditions; when temperature is >35℃ or humidity is <50%RH (high temperature and low humidity), the weight of environmental features is adjusted from 0.2 to 0.3, the weight of spectral features is adjusted to 0.4, and the weight of visual features is adjusted to 0.3 to help correct the deviation of spectral features caused by high temperature and low humidity; under normal conditions, the weight allocation is 0.5 for visual features, 0.3 for spectral features, and 0.2 for environmental features.

[0038] Model-level anti-interference optimization: A dynamic incremental training module was constructed to collect pest and disease data under different environments, including scenarios with low light, strong light, high temperature, high humidity, and shading. A total of 5000 sets of incremental training data were collected. Mini-batch gradient descent (batch size of 16, learning rate of 0.0001) was used to incrementally train the model, with each incremental training round consisting of 20 rounds, continuously optimizing the model parameters to adapt the model to different field environments. Simultaneously, a robust loss function was introduced, where the base loss L_base is the sum of the cross-entropy loss (recognition branch) and the coordinate loss (localization branch), with a penalty coefficient... =0.5, the MAE function calculates the mean absolute error between the true features and the predicted features, minimizes the robustness loss through backpropagation, and reduces the impact of environmental interference on model inference.

[0039] Anti-interference effect verification: The model was tested under different field conditions, including low light (light intensity 5000 lx), strong light (light intensity 95000 lx), high temperature and humidity (temperature 38℃, humidity 88%RH), and shading (leaf covering 50% of the lesion area). 200 sets of data were tested in each scenario. The test results showed that the model's recognition accuracy remained above 90% under various complex environments, specifically 91.5% in low light, 92.3% in strong light, 90.8% in high temperature and humidity, and 90.2% in shading. The positioning accuracy fluctuation was less than 2%. Compared to models without anti-interference mechanisms (recognition accuracy 75–82%), the robustness and adaptability were significantly improved, meeting the needs of practical field applications.

[0040] Validation data: This invention has been validated through field tests in multiple experimental fields, including rice paddies, wheat fields, and corn fields. It has achieved excellent performance against common crop diseases and pests such as rice blast, wheat stripe rust, and corn leaf blight. Specific data are summarized below: the detection rate of rice blast latent period is 91.2%, a 37% improvement compared to traditional static threshold methods; the detection rate of wheat stripe rust latent period is 89.7%, a 32% improvement compared to traditional methods; the detection rate of corn leaf blight latent period is 90.5%, a 34% improvement compared to traditional methods. The detection rate of all diseases and pests latent periods is 3–7 days earlier, providing ample time for early control. The accuracy rate for identifying early-stage diseases and pests (subclinical symptoms) is 88.3%, the accuracy rate for mid-stage diseases and pests is 95.7%, and the overall accuracy rate is 92.0%. Even under complex environments (weak light, strong light, high temperature and humidity, shading), the accuracy rate remains above 90%, a 13.5% improvement compared to single-modal identification methods (accuracy rate 78.5%). Early-stage pest and disease localization errors were controlled within 4.2% (average 3.8%), and mid-stage pest and disease localization errors were controlled within 3.0%. The average bounding box overlap ratio (IOU) was 0.86, a 19.4% improvement compared to the traditional YOLO network (IOU 0.72). The integrated recognition-localization inference speed reached 30 frames / second, meeting the needs of real-time field monitoring. Within a light intensity range of 5000–95000 lx, temperature of 15–38℃, and humidity of 40–90%RH, the model's recognition accuracy fluctuated by less than 2%, and its localization accuracy fluctuated by less than 1.5%. It exhibited strong adaptability to interference factors such as leaf shading and differences in crop growth, with anti-interference performance improved by more than 35% compared to existing technologies. The model training time is reduced by 28% compared to traditional multimodal models. It has strong hardware adaptability and can be deployed on various platforms such as embedded devices, drones, and mobile terminals. In field testing, the monitoring time for a single plot (10 mu) is reduced by 85% compared to manual monitoring, and the monitoring cost is reduced by 70%. It can realize early and accurate monitoring and control of pests and diseases in large-scale fields.

[0041] This method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion extracts visual, spectral, and environmental modal features through three parallel branches, comprehensively covering macroscopic phenotypes, microscopic physiological and biochemical information, and environmental context. Simultaneously, a cross-modal adaptive alignment module is introduced, using mutual information loss and gated attention weights to eliminate semantic gaps and scale differences between heterogeneous data, achieving efficient fusion of multimodal features and improving the comprehensiveness and accuracy of feature representation. This provides reliable feature support for early identification and localization of diseases and pests, solving the problems of single-modal feature distortion and incomplete information in existing technologies. Furthermore, this invention establishes individualized health spectral baselines through a dynamic construction method for crop health spectral baselines. Combined with key stress band identification and environmental stress index analysis, it achieves accurate determination of the incubation period of diseases and pests, detecting latent diseases 3–7 days earlier than traditional methods. This effectively solves the problem of existing technologies being unable to accurately predict the incubation period of diseases and pests, gaining valuable time for early prevention and control. This invention amplifies subtle early physiological changes in crops by calculating the differential signal of multimodal fusion features at adjacent time points. Combined with a multi-head attention mechanism, it focuses on non-natural fluctuation feature channels in the time dimension, achieving reliable identification of subclinical symptoms and significantly improving the accuracy of identifying early, minor pests and diseases. This overcomes the shortcomings of existing technologies in identifying early, latent symptoms. By simultaneously inputting multimodal fusion features into the identification and localization branches, this invention achieves "one-time inference, simultaneous output." Combining an improved classifier and an improved YOLO network, and introducing a global contextual attention mechanism for collaborative optimization, the localization error is controlled within 5%. This enables precise localization of small areas of early pests and diseases, providing accurate location information for precise field control and improving control efficiency and targeting. This invention effectively resists interference from complex field environments such as changes in light intensity, temperature and humidity fluctuations, and shading by using feature-level environmental adaptive adjustment and dynamic incremental training and the introduction of robust loss functions at the model level. This allows the model to maintain an accuracy rate of over 90% and a stable positioning accuracy under different environmental conditions, solving the problem of unstable positioning accuracy in complex field environments in existing technologies and meeting the needs of practical field applications.

[0042] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A crop disease and pest early identification and positioning method based on visual-spectrum-environment multi-modal feature fusion, characterized in that, Includes the following steps: S1: Acquire three types of modal data: high-resolution RGB images of crops, multispectral reflectance data, and real-time environmental parameters; S2: Construct a multi-branch parallel feature extraction network to extract features from three types of modal data, obtaining visual features, spectral features, and environmental context features; S3: The three types of modal features are embedded spatially aligned and fused through a cross-modal adaptive alignment module to obtain multimodal fused features; S4: Employ temporal differential attention mechanism to enhance early stress features of multimodal fusion features, amplifying small early physiological changes in crops; S5: Construct a latent period early warning model based on spectral band shift and physiological baseline to determine the latent period of crop diseases and pests; S6: Input the enhanced fusion features into the integrated recognition-localization network, and simultaneously output the pest and disease category and precise location information; S7: Through an adaptive anti-interference optimization mechanism for complex field environments, the model's recognition and positioning accuracy is ensured to remain stable in complex field environments; In S2, the multi-branch parallel feature extraction network includes three parallel branches: the visual feature branch adopts an improved lightweight convolutional neural network to extract macroscopic phenotypic features such as crop texture, shape, and color from high-resolution RGB images. The lightweight convolutional neural network is EfficientNet or MobileNetV4, which is improved by adjusting the size and number of convolutional kernels, introducing depthwise separable convolution and attention mechanisms, reducing parameter redundancy while improving feature extraction accuracy. The spectral feature branch uses a one-dimensional convolutional neural network or Transformer model to process reflectance data of key bands in green, red, red edge, and near-infrared light acquired by multispectral sensors. It extracts microscopic spectral features related to crop chlorophyll content, water content, and nitrogen levels, focusing on capturing narrowband features sensitive to early stress. The environmental feature branch uses a fully connected network to process real-time environmental parameters such as temperature, humidity, light intensity, and soil conductivity. It performs nonlinear mapping of environmental parameters through hidden layer activation functions to extract environmental context features. The ReLU function is used as the activation function. In S3, the implementation process of the cross-modal adaptive alignment module is as follows: First, the mutual information loss between visual features and spectral features is calculated to measure the correlation between the two types of modal features. The calculation formula is as follows: Where I(X,Y) represents the mutual information between visual feature X and spectral feature Y, H(X) represents the information entropy of visual feature X, H(Y) represents the information entropy of spectral feature Y, and H(X,Y) represents the joint information entropy of visual feature X and spectral feature Y; then, the gating attention weights are calculated using environmental features, and the visual and spectral features are weighted and adjusted. The weight calculation formula is as follows: Where W is the gating attention weight vector, The sigmoid activation function is used, We is the environmental feature weight matrix, E is the environmental context feature vector, and be is the bias term. Finally, the weighted visual features, spectral features, and environmental features are concatenated and fused to obtain the multimodal fusion feature. ,in This is a multimodal fusion feature vector; In S4, the implementation process of the temporal differential attention mechanism is as follows: Multiple frames of multimodal data from the same plant or region are continuously collected, denoted as t0, t1, t2, ..., tn, where ti represents the multimodal data at time i; the multimodal fusion feature difference between adjacent time points is calculated using the following formula: ,in, Let Ft be the difference signal of the fused features at adjacent time points, Ft be the multimodal fused feature at time t, and Ft-1 be the multimodal fused feature at time t-1. A multi-head attention mechanism is used to perform weighted fusion of the difference features and the absolute features at the current time point. The fusion formula is as follows: ,in, To enhance early stress characteristics, This represents a multi-head attention calculation function that uses the current time-time features as the query vector and the difference features as the key and value vectors. It achieves feature enhancement through attention weight allocation, focusing on non-natural fluctuation feature channels in the time dimension. In S5, the latent period early warning model based on spectral band shift and physiological baseline adopts a dynamic construction method for crop health spectral baseline. The specific implementation steps are as follows: continuously collect 72 hours of hyperspectral data during the crop health period, establish an individualized health spectral baseline using statistical methods, covering a 95% confidence interval, and the baseline calculation formula is as follows: ,in wavelength The healthy spectral baseline range, wavelength The mean of hyperspectral reflectance, wavelength The standard deviation of hyperspectral reflectance is used, with 1.96 being the coefficient corresponding to the 95% confidence interval; key stress bands are identified: 780–795nm, 676–745nm, and 550–570nm; early warning criteria are set: when the spectral reflectance of a pixel in a key stress band exceeds the baseline confidence interval for 24 consecutive hours, and the environmental stress index increases, it is determined to be a latent disease. The formula for calculating the environmental stress index is: Where ESI is the environmental stress index, T is the difference between temperature and the suitable range, H is the difference between humidity and the suitable range, L is the difference between light intensity and the suitable range, and S is the difference between soil electrical conductivity and the suitable range. These are the weighting coefficients for each environmental parameter, and their sum is 1. In S7, the adaptive anti-interference optimization mechanism for complex field environments includes optimizations at the feature level and the model level: At the feature level, an environmental adaptive adjustment module is introduced during multimodal fusion. Based on real-time collected environmental parameters, the weights and extraction strategies of each modality feature are dynamically adjusted. Under strong or weak light conditions, the weights of spectral features are enhanced to compensate for visual feature distortion, and under high temperature and high humidity conditions, the weights of environmental features are enhanced to help correct spectral feature bias. At the model level, a dynamic incremental training module is constructed. Using pest and disease data measured in different field environments, the model is incrementally trained. The model parameters are updated using a mini-batch gradient descent method to continuously optimize the model's adaptability to complex environments. Simultaneously, a robust loss function is introduced to reduce the impact of environmental interference on model inference. The robust loss function is: , where L robust For robust loss; L base It is the sum of the base loss, cross-entropy loss, and coordinate loss; The penalty coefficient is... F is the mean absolute error function. true For the true feature vector, F pred To predict the feature vector.

2. The method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion according to claim 1, characterized in that: In the calculation of mutual information loss, the specific formula for calculating information entropy H(X) is as follows: Where n is the feature dimension of visual feature X, p(x) i Let be the probability distribution of the i-th dimension of the visual feature; the formula for calculating the joint information entropy H(X,Y) is: Where m is the feature dimension of the spectral feature Y, p(x i ,y j ) represents the joint probability distribution of the i-th dimension of visual features and the j-th dimension of spectral features.

3. The method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion according to claim 1, characterized in that: The multi-head attention mechanism includes 4-8 attention heads, each of which independently calculates attention weights. The outputs of each attention head are normalized and concatenated, and then passed through a fully connected layer to obtain the final enhanced features. The time-series multi-frame data is acquired at a frequency of one frame every 2-4 hours to ensure that small physiological changes in the early stages of crops can be captured, while avoiding data redundancy.

4. The method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion according to claim 1, characterized in that: In S6, the integrated identification-localization network architecture simultaneously inputs the enhanced multimodal fusion features into the identification branch and the localization branch, achieving simultaneous inference and output of results. The identification branch employs an improved classifier composed of fully connected layers and Softmax layers. Combined with the fused feature vectors, it optimizes the classification loss function to achieve accurate identification of pest and disease categories, with a focus on classifying early-stage, minor pests and diseases. The localization branch uses an improved YOLO network, introducing a feature pyramid module to enhance the localization capability for small areas and minor pests and diseases, outputting the precise coordinates, area, and bounding box of the pest and disease region. The improved YOLO network optimizes the label allocation strategy and detection head structure, employs a TAL dynamic matching strategy to achieve balanced classification and regression, and uses a simplified ET-Head to improve detection accuracy. A global contextual attention mechanism is introduced into the network to collaboratively optimize the identification and localization results, improving localization accuracy.

5. The method for early identification and localization of crop diseases and pests based on visual-spectral-environmental multimodal feature fusion according to claim 4, characterized in that: The collaborative optimization process of the integrated recognition-localization network is as follows: the category confidence output by the recognition branch is used as an important basis for the global context attention weight, and the feature map of the localization branch is adjusted by weighting. The higher the category confidence, the greater the localization weight of the corresponding region. At the same time, the bounding box information output by the localization branch is fed back to the recognition branch to correct the feature extraction focus of the recognition branch, so as to achieve bidirectional optimization of recognition and localization results and keep the localization error within 5%.

Citation Information

Patent Citations

  • Crop disease identification method based on multi-modal fusion and knowledge distillation

    CN121542654A

  • Intelligent identification and precise prevention and control system for urban vegetable garden diseases and insect pests based on computer vision and multi-modal data fusion

    CN121686039A