Codonopsis pilosula pest recognition system based on image recognition
The Codonopsis pilosula pest identification system, which utilizes multispectral imaging and environmental awareness, dynamically adjusts image enhancement and cross-modal feature fusion. By combining dynamic federated learning and user feedback, it generates a risk heat map, solving the accuracy problem of Codonopsis pilosula pest identification in low-light and high-humidity environments, and achieving efficient identification and precise control.
Patent Information
- Application Number
- CN202511503641.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-02
AI Technical Summary
The existing Codonopsis pilosula pest identification system has low accuracy in low light and high humidity environments, leading to missed detections and spread of pests. This causes farmers to miss the opportunity for prevention and control, increase the application of chemical pesticides, and cause economic losses.
The system uses a multispectral imaging module and an environmental perception suite to collect data in real time. It performs preliminary processing through edge computing, dynamically adjusts the image enhancement strategy, reconstructs the super-resolution features of hidden pests, generates joint classification probabilities through a cross-modal feature fusion network, generates risk heat maps by combining a dynamic federated learning mechanism and meteorological data, and automatically generates adversarial example iterative models based on user feedback.
It improves the robustness of pest identification in Codonopsis pilosula, reduces missed detections, guides precise pesticide application, increases the identification rate, forms a closed-loop optimization mechanism, reduces chemical pesticide application, and lowers economic losses.
Smart Images

Figure CN121259445A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pest identification technology, and in particular to a Codonopsis pilosula pest identification system based on image recognition. Background Technology
[0002] Codonopsis pilosula, a traditional and precious Chinese medicinal herb, is known for its effects of tonifying the spleen and replenishing qi, nourishing blood and promoting body fluid production, and its market demand continues to grow. Its high economic value and wide range of uses in medicine, health products, and exports make it an important crop for increasing farmers' income. At the same time, pest problems have become a key bottleneck restricting the industry's development, making efficient identification of pests affecting Codonopsis pilosula particularly important in pest control.
[0003] However, over time, the robustness of traditional Codonopsis pilosula pest identification systems has become increasingly apparent. For example, in a Codonopsis pilosula planting base, after three consecutive days of rain during the rainy season, water droplets adhered to the leaf surface, blurring the camera lens. Whiteflies hid on the back of the leaves. Due to the rain obscuring the lens and the blurred image, the system could not capture a clear image. The accuracy of identification dropped significantly in low light and high humidity environments, and a large number of whiteflies were missed. Ultimately, the whiteflies spread to the entire field within 72 hours, resulting in a yellowing rate of over 40% of the leaves and a 25% reduction in yield. At the same time, because the system did not issue a warning, farmers missed the opportunity for control and were forced to apply chemical pesticides twice more, causing serious economic and property losses. Therefore, an image recognition-based Codonopsis pilosula pest identification system is proposed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing an image recognition-based pest identification system for Codonopsis pilosula.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: A system for identifying pests of Codonopsis pilosula based on image recognition, comprising: Perception Layer: Deploys multispectral imaging modules (visible light / near infrared / hyperspectral cameras) and environmental perception kits (temperature, humidity, light, and wind speed sensors) to collect images of Codonopsis pilosula plants and environmental data in real time; performs preliminary compression and noise filtering on the images through edge computing nodes, synchronously records sensor data and marks timestamps; and transmits the processed multimodal data stream to the preprocessing layer. Preprocessing layer: Dynamically adjust image enhancement strategies based on environmental perception data: improve contrast under low light conditions and use blind deconvolution to deblur under high wind speeds; perform super-resolution reconstruction on images of hidden pests to generate multi-scale feature pyramids; fuse environmental parameters (such as temperature and humidity) with image features to output standardized data adapted to complex field scenarios for analysis by the model layer. Model layer: Image features, environmental features, and biological features (such as the presence of natural enemy insects) are integrated through a cross-modal feature fusion network to generate joint classification probabilities; a dynamic federated learning mechanism is adopted, in which each regional node independently trains a local model, and the aggregation parameters are periodically encrypted and the global model is updated; the model aggregation cycle is dynamically adjusted according to the frequency of environmental changes (such as the rainy season) to improve cross-regional generalization ability; Application layer: Real-time risk heat maps are generated by combining meteorological data with pest and disease spread models to guide precise pesticide application; users submit feedback on misjudgments through the APP, which automatically generates adversarial samples and triggers model iteration; monthly performance reports are output to optimize key parameters (such as the proportion of adversarial samples); finally, a decision report containing identification results, prevention and control suggestions and ecological impact assessment is output, forming a closed loop of "identification-feedback-evolution".
[0006] The above technical solution further includes: Furthermore, the deployment includes a multispectral imaging module with visible light, near-infrared, and hyperspectral cameras, and an environmental sensing kit including temperature, humidity, light intensity, and wind speed sensors, comprising the following steps: Hardware installation: An industrial-grade CCD camera (4096×2160 resolution, 10fps) is installed on a central pillar (3 meters high) in the planting area, covering a diameter of 50 meters with no blind spots. An InGaAs sensor with a wavelength range of 700-1000nm is installed alongside the visible light camera, using a beam splitter for same-view imaging. A pushbroom hyperspectral imager (400-1000nm spectral range, 2.5nm spectral resolution) is installed on a movable track, scanning the entire field once a week (at a speed of 0.1m / s) to generate a hyperspectral data cube. An SHT35-DIS temperature and humidity sensor is installed at the height of the plant canopy (50cm above the ground), with two sensors deployed per acre to avoid direct sunlight. A BH1750FVI light sensor is installed alongside the temperature and humidity sensor to measure canopy light intensity. A three-cup anemometer is installed in an open area of the field (2 meters high), away from obstructions. Configure synchronization device parameters: The visible light camera is set to automatic exposure and automatic white balance, with digital noise reduction turned off to preserve original image details; the near-infrared camera has a fixed exposure time and its gain is adjusted to a signal-to-noise ratio greater than 40dB to ensure clear features on the back of the leaves; the hyperspectral camera is configured with pushbroom speed and exposure time to generate a data cube; the temperature and humidity sensor is set to a sampling frequency of 1 time / minute, a data upload cycle of 5 minutes, and real-time alarms for abnormal values (such as temperature >35℃); the light sensor is configured with dynamic range adjustment to automatically compensate for changes in light intensity (such as switching between cloudy and sunny days); the wind speed sensor is set to an averaging mode (10-minute average value) with threshold triggering (activating the image deblurring algorithm when the wind speed is greater than 5m / s). All devices synchronize timestamps via the NTP protocol to ensure a one-to-one correspondence between images and environmental data; multispectral camera and sensor data are published in topics via ROS (Robot Operating System); Protection calibration: The camera and sensors are enclosed in an IP67-rated enclosure and equipped with an automatic heating and defogging function (activated when the temperature is <10℃). A sunshade is installed on the hyperspectral camera track to prevent direct sunlight from causing the sensor to overheat. The multispectral camera is radiometrically calibrated monthly using a standard reflector to ensure consistent spectral data. The environmental sensor is compared with a standard instrument quarterly, and automatic calibration is triggered when the deviation exceeds 5%.
[0007] Data transmission and storage: Visible / near-infrared images are transmitted in real time to edge computing nodes via a 5G network. The hyperspectral data cube is stored locally and then uploaded to the cloud daily. Environmental data is transmitted to the gateway (with a coverage distance of 1km) via the LoRaWAN protocol and then forwarded to the cloud via a 4G network. Edge nodes are equipped with SSDs to store high-resolution images and raw environmental data from the past 7 days as a cloud backup. The hyperspectral data cube is compressed (H.265 encoding, compression ratio 10:1) and stored on a NAS device with a retention period of 30 days.
[0008] Furthermore, the method of dynamically adjusting the image enhancement strategy based on environmental perception data includes the following steps: Dynamic strategy selection logic: In low-light scenarios (light intensity less than 50 lux), the Retinex algorithm is triggered to enhance contrast and improve the visibility of small pests (such as whiteflies) on the underside of leaves. In high-wind-speed scenarios (wind speed greater than 5 m / s), a blind deconvolution algorithm (kernel size 11×11, number of iterations 50) is applied to remove motion blur and restore the outline features of fast-moving thrips pests. Special treatment for hidden pests: An SRCNN model (3-layer convolution, filter size 9×9 / 1×1 / 5×5) was applied to the images of the underside of the leaves to increase the resolution to 2000 dpi, clearly presenting the fine features such as the antennae of aphids; the reconstructed images were downsampled (scale factor 0.5 / 1.0 / 2.0), and HOG features (cell size 8×8, number of orientations 9) were extracted at each scale and stitched together to form a multi-scale feature vector; Fusion of environmental parameters and image features: Environmental parameters (temperature, humidity, wind speed) are encoded into 128D vectors (LSTM time series encoding), and concatenated with image features (HOG+SIFT, dimension 2048D) to form 2197D features; dimensionality reduction is achieved through a fully connected layer (size 1024), and standardized data (HDF5 format) adapted to complex field scenarios is output for analysis by the model layer; Strategy optimization: Farmers use the app to mark misjudgment cases (such as misjudging ladybugs as aphids), which automatically generates adversarial samples. These adversarial samples are mixed with the original data (in a ratio of 1:5), triggering fine-tuning of the Retinex / blind deconvolution algorithm parameters. The strategy is updated within 72 hours.
[0009] Furthermore, the super-resolution reconstruction of the images of concealed pests to generate a multi-scale feature pyramid includes the following steps: Filtering images of hidden pests: The environmental perception suite detects low light or high wind speed, or users manually mark "areas that need enhancement"; it quickly locates hidden areas on the back of leaves and in stem crevices using a shallow CNN model, and marks the locations of potential pests (such as areas where whiteflies congregate). Super-resolution reconstruction (SRCNN model): Extract a hidden area image patch (64×64 pixels). If the original resolution is less than 500dpi, automatically trigger reconstruction. Algorithm flow: First convolutional layer: 9×9 filter extracts low-frequency features, outputs 64 channels, activation function ReLU; The second convolutional layer uses a 1×1 filter to perform non-linear transformation of feature mapping, with 32 output channels. The third convolutional layer: a 5×5 filter reconstructs high-frequency details, outputting a super-resolution image with the same size as the input; Calculate the PSNR (Peak Signal-to-Noise Ratio) of the reconstructed image and the original high-resolution image, with a target value greater than 30 dB; Generate a multi-scale feature pyramid: The reconstructed images were downsampled at three levels (scale factors 0.5 / 1.0 / 2.0) to cover the entire life cycle features of the pest. Scale 0.5: Captures the wing veins and markings of adult insects (such as thrips wing veins); Scale 1.0: Identify the body wall features of larvae (nymphs) (such as the waxy layer of whitefly nymphs); Scale 2.0: Detects egg arrangement patterns (e.g., aphid egg masses). Feature extraction: HOG characteristics: cell size 8×8, number of orientations 9, describing the distribution of local gradient orientations; SIFT features: keypoint detection (Hessian matrix threshold 0.04), generating 128-dimensional descriptors; Feature fusion: The three-scale features (HOG+SIFT) are concatenated into a 2197D vector, and then reduced to 1024D by PCA to form a multi-scale feature pyramid.
[0010] Integration with environmental parameters: The feature pyramid is concatenated with temperature, humidity, and wind speed data (timestamp error <10ms) to generate a joint feature vector; the data is then normalized using the BatchNorm layer to output standardized data suitable for model layer analysis.
[0011] Furthermore, the integration of image features, environmental features, and biological features through a cross-modal feature fusion network includes the following steps: Multimodal feature extraction preprocessing: The super-resolution reconstructed image was divided into 8×8 cell sizes, and gradient histograms in 9 directions were extracted to generate a 2048-dimensional feature vector. Keypoints were detected using the Hessian matrix to generate a 128-dimensional descriptor, which was then concatenated with the HOG features to form a 2197D vector. Temperature, humidity, and wind speed data (sampling interval of 1 minute) were input into a two-layer LSTM network (hidden layer size 64) to extract the 128D feature vector at the last time step. A pest morphology feature library (such as aphid antenna length and spider mite body color R value) was constructed based on expert knowledge, and a 512D feature vector was generated through KNN matching. Dynamic weighting: The timestamps of image features and environmental features are synchronized using the NTP protocol to ensure that the features correspond to the field conditions at the same time. The locations (pixel coordinates) of pests in the image features are mapped to the coordinate system of the environmental sensor, and local temperature, humidity, and wind speed data are associated. The modal weights (0.6 for image, 0.3 for environment, and 0.1 for organisms) are calculated using Softmax to suppress noisy modalities (such as invalid images under strong winds). The SE module is applied to each channel of the feature vector to improve the response values of key features (such as the wing veins of the whitefly). Cross-modal feature fusion network structure: Image, environmental, and biological features are input into a three-layer fully connected network (size 2048→1024→64), with ReLU activation function. The dimensionality-reduced features (1024+64+256=1344D) are concatenated into a fused feature, which is then input into a two-layer fully connected network (1344→512→256), with ReLU activation function. The loss function includes a classification loss with cross-entropy loss (labeled as pest type) and a contrast loss to increase the distance between features of different categories (margin=2.0) to improve inter-class separability. Generate joint classification probability: The fused features are input into the Softmax layer to generate N-dimensional classification probabilities (N is the number of pest species, such as aphids, spider mites, and whiteflies); the probability distribution is adjusted by temperature scaling to avoid overfitting.
[0012] Furthermore, the dynamic federated learning mechanism includes the following steps: First, local models are trained. Each edge node independently trains a lightweight CNN model based on local data (image features, environmental parameters, and biological features), using the SGD optimizer and a loss function of cross-entropy + contrastive loss (weight 0.5). Then, the aggregation period is dynamically adjusted: the aggregation period (from 1 hour to 24 hours) is automatically adjusted according to the frequency of environmental changes (such as wind speed fluctuations greater than 2 m / s or sudden changes in temperature and humidity). The more drastic the environmental changes, the higher the aggregation frequency, ensuring that the model can quickly adapt to new scenarios. Next, encrypted parameters are aggregated. Each node uploads its model parameters to the cloud after homomorphic encryption using Paillier. The cloud calculates a weighted average (weight = local data volume / total data volume) to generate global model parameters and distribute them. The entire process does not expose the original data, protecting farmers' privacy. Finally, heterogeneous adaptation is performed. Models from different regions are integrated using a federated averaging algorithm, allowing each node to retain some personalized parameters (such as convolutional kernels for specific pests), improving the generalization ability of the global model in the identification of cross-regional pests such as aphids and spider mites.
[0013] Furthermore, the process of generating a real-time risk heat map by combining meteorological data with a pest and disease spread model includes the following steps: Multi-source data fusion: Meteorological data (temperature, humidity, rainfall, wind speed) is acquired in real time through the national meteorological station API and spatially interpolated with field environmental data (such as canopy humidity) collected by the sensing layer to generate a meteorological field with a 5m×5m grid; the pest and disease spread model adopts the improved SEIR model; Risk calculation: The model uses the current distribution of pests and diseases (detected via a multispectral imaging module) as initial conditions to simulate the spread path over the next 24 hours; combining meteorological field data with the spread simulation results, a weighted formula is used: Where M is the risk value, P is the diffusion coverage rate, and D is the environmental suitability; environmental suitability is determined by temperature, humidity, and rainfall (e.g., suitability increases by 30% when canopy humidity is greater than 80%). Generate a visual heatmap: The risk values are mapped to a geographic coordinate system, and a pseudo-color heat map (red for high risk and blue for low risk) is generated using OpenCV and overlaid on the field map. The heat map is then pushed to the farmer's APP in real time through edge nodes, indicating high-risk areas and recommending prevention and control measures.
[0014] Furthermore, the user submits feedback on misjudgments via the APP, which automatically generates adversarial examples and triggers model iteration, including the following steps: Handling user misjudgment feedback and generating adversarial examples: Farmers upload misclassified images via an app, such as ladybugs being misclassified as aphids, and select the misclassification type (missed / misclassified / location deviation). The app automatically labels the timestamp and geographical location. After the feedback image is reconstructed by a preprocessing layer using super-resolution, it is input into the WGAN-GP model to generate adversarial examples: the perturbation intensity ε=0.1 (L∞ norm constraint) ensures that the visual difference between the adversarial examples and the original images is less than 5% (the threshold that humans can distinguish). The generated samples are mixed with the original labels in a 1:5 ratio to build an incremental training dataset, providing data support for model iteration. Federal update verification: Edge nodes fine-tune their local models using incremental datasets, employing cross-entropy plus contrastive loss as the loss function. The fine-tuned model parameters are uploaded to the cloud via Paillier homomorphic encryption and federated with parameters from other regions, with the weights being the ratio of local data volume to total data volume. The updated global model is then distributed to each node, and the accuracy improvement is verified on a test set containing adversarial examples (e.g., whitefly identification rate increases from 85% to 93%). If the improvement is less than 3%, the perturbation strength is adjusted to ε=0.15 to trigger a second iteration, until the loss value is less than 0.1, forming a closed-loop optimization mechanism of "feedback-generation-iteration-verification".
[0015] The present invention has the following beneficial effects: In this invention, multi-spectral imaging modules and environmental perception kits are deployed to collect and preliminarily process multi-modal data in real time. Based on environmental data, images are dynamically enhanced to reconstruct super-resolution features of hidden pests and fuse environmental parameters. Through a cross-modal fusion network, images, environmental and biological features are integrated, and model parameters are aggregated across regions using a dynamic federated learning mechanism to adapt to environmental changes and adjust the update cycle. Finally, a risk heatmap is generated to guide pesticide application. Adversarial sample iterative models are automatically generated based on user feedback, and a decision report containing prevention and control suggestions and ecological assessments is output, forming a "recognition-feedback-evolution" closed loop, which effectively solves the problem of insufficient robustness in the current identification of Codonopsis pilosula pests. Attached Figure Description
[0016] Figure 1 This is a system block diagram of a Codonopsis pilosula pest identification system based on image recognition proposed in this invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1As shown, this invention is an image recognition-based pest identification system for Codonopsis pilosula, comprising: Perception Layer: Deploys multispectral imaging modules (visible light / near infrared / hyperspectral cameras) and environmental perception kits (temperature, humidity, light, and wind speed sensors) to collect images of Codonopsis pilosula plants and environmental data in real time; performs preliminary compression and noise filtering on the images through edge computing nodes, synchronously records sensor data and marks timestamps; and transmits the processed multimodal data stream to the preprocessing layer. Preprocessing layer: Dynamically adjust image enhancement strategies based on environmental perception data: improve contrast under low light conditions and use blind deconvolution to deblur under high wind speeds; perform super-resolution reconstruction on images of hidden pests to generate multi-scale feature pyramids; fuse environmental parameters (such as temperature and humidity) with image features to output standardized data adapted to complex field scenarios for analysis by the model layer. Model layer: Image features, environmental features, and biological features (such as the presence of natural enemy insects) are integrated through a cross-modal feature fusion network to generate joint classification probabilities; a dynamic federated learning mechanism is adopted, in which each regional node independently trains a local model, and the aggregation parameters are periodically encrypted and the global model is updated; the model aggregation cycle is dynamically adjusted according to the frequency of environmental changes (such as the rainy season) to improve cross-regional generalization ability; Application layer: Real-time risk heat maps are generated by combining meteorological data with pest and disease spread models to guide precise pesticide application; users submit feedback on misjudgments through the APP, which automatically generates adversarial samples and triggers model iteration; monthly performance reports are output to optimize key parameters (such as the proportion of adversarial samples); finally, a decision report containing identification results, prevention and control suggestions and ecological impact assessment is output, forming a closed loop of "identification-feedback-evolution".
[0019] In one embodiment, the deployment includes a multispectral imaging module with visible light, near-infrared, and hyperspectral cameras, and an environmental sensing kit including temperature, humidity, light intensity, and wind speed sensors, comprising the following steps: Hardware installation: An industrial-grade CCD camera (4096×2160 resolution, 10fps) is installed on a central pillar (3 meters high) in the planting area, covering a diameter of 50 meters with no blind spots. An InGaAs sensor with a wavelength range of 700-1000nm is installed alongside the visible light camera, using a beam splitter for same-view imaging. A pushbroom hyperspectral imager (400-1000nm spectral range, 2.5nm spectral resolution) is installed on a movable track, scanning the entire field once a week (at a speed of 0.1m / s) to generate a hyperspectral data cube. An SHT35-DIS temperature and humidity sensor is installed at the height of the plant canopy (50cm above the ground), with two sensors deployed per acre to avoid direct sunlight. A BH1750FVI light sensor is installed alongside the temperature and humidity sensor to measure canopy light intensity. A three-cup anemometer is installed in an open area of the field (2 meters high), away from obstructions. Configure synchronization device parameters: The visible light camera is set to automatic exposure and automatic white balance, with digital noise reduction turned off to preserve original image details; the near-infrared camera has a fixed exposure time and its gain is adjusted to a signal-to-noise ratio greater than 40dB to ensure clear features on the back of the leaves; the hyperspectral camera is configured with pushbroom speed and exposure time to generate a data cube; the temperature and humidity sensor is set to a sampling frequency of 1 time / minute, a data upload cycle of 5 minutes, and real-time alarms for abnormal values (such as temperature >35℃); the light sensor is configured with dynamic range adjustment to automatically compensate for changes in light intensity (such as switching between cloudy and sunny days); the wind speed sensor is set to an averaging mode (10-minute average value) with threshold triggering (activating the image deblurring algorithm when the wind speed is greater than 5m / s). All devices synchronize timestamps via the NTP protocol to ensure a one-to-one correspondence between images and environmental data; multispectral camera and sensor data are published in topics via ROS (Robot Operating System); Protection calibration: The camera and sensors are enclosed in an IP67-rated enclosure and equipped with an automatic heating and defogging function (activated when the temperature is <10℃). A sunshade is installed on the hyperspectral camera track to prevent direct sunlight from causing the sensor to overheat. The multispectral camera is radiometrically calibrated monthly using a standard reflector to ensure consistent spectral data. The environmental sensor is compared with a standard instrument quarterly, and automatic calibration is triggered when the deviation exceeds 5%.
[0020] Data transmission and storage: Visible / near-infrared images are transmitted in real time to edge computing nodes via a 5G network. The hyperspectral data cube is stored locally and then uploaded to the cloud daily. Environmental data is transmitted to the gateway (with a coverage distance of 1km) via the LoRaWAN protocol and then forwarded to the cloud via a 4G network. Edge nodes are equipped with SSDs to store high-resolution images and raw environmental data from the past 7 days as a cloud backup. The hyperspectral data cube is compressed (H.265 encoding, compression ratio 10:1) and stored on a NAS device with a retention period of 30 days.
[0021] In one embodiment, the dynamic adjustment of the image enhancement strategy based on environmental perception data includes the following steps: Dynamic strategy selection logic: In low-light scenarios (light intensity less than 50 lux), the Retinex algorithm is triggered to enhance contrast and improve the visibility of small pests (such as whiteflies) on the underside of leaves. In high-wind-speed scenarios (wind speed greater than 5 m / s), a blind deconvolution algorithm (kernel size 11×11, number of iterations 50) is applied to remove motion blur and restore the outline features of fast-moving thrips pests. Special treatment for hidden pests: An SRCNN model (3-layer convolution, filter size 9×9 / 1×1 / 5×5) was applied to the images of the underside of the leaves to increase the resolution to 2000 dpi, clearly presenting the fine features such as the antennae of aphids; the reconstructed images were downsampled (scale factor 0.5 / 1.0 / 2.0), and HOG features (cell size 8×8, number of orientations 9) were extracted at each scale and stitched together to form a multi-scale feature vector; Fusion of environmental parameters and image features: Environmental parameters (temperature, humidity, wind speed) are encoded into 128D vectors (LSTM time series encoding), and concatenated with image features (HOG+SIFT, dimension 2048D) to form 2197D features; dimensionality reduction is achieved through a fully connected layer (size 1024), and standardized data (HDF5 format) adapted to complex field scenarios is output for analysis by the model layer; Strategy optimization: Farmers use the app to mark misjudgment cases (such as misjudging ladybugs as aphids), which automatically generates adversarial samples. These adversarial samples are mixed with the original data (in a ratio of 1:5), triggering fine-tuning of the Retinex / blind deconvolution algorithm parameters. The strategy is updated within 72 hours.
[0022] In one embodiment, the super-resolution reconstruction of the image of the concealed pest to generate a multi-scale feature pyramid includes the following steps: Filtering images of hidden pests: The environmental perception suite detects low light or high wind speed, or users manually mark "areas that need enhancement"; it quickly locates hidden areas on the back of leaves and in stem crevices using a shallow CNN model, and marks the locations of potential pests (such as areas where whiteflies congregate). Super-resolution reconstruction (SRCNN model): Extract a hidden area image patch (64×64 pixels). If the original resolution is less than 500dpi, automatically trigger reconstruction. Algorithm flow: First convolutional layer: 9×9 filter extracts low-frequency features, outputs 64 channels, activation function ReLU; The second convolutional layer uses a 1×1 filter to perform non-linear transformation of feature mapping, with 32 output channels. The third convolutional layer: a 5×5 filter reconstructs high-frequency details, outputting a super-resolution image with the same size as the input; Calculate the PSNR (Peak Signal-to-Noise Ratio) of the reconstructed image and the original high-resolution image, with a target value greater than 30 dB; Generate a multi-scale feature pyramid: The reconstructed images were downsampled at three levels (scale factors 0.5 / 1.0 / 2.0) to cover the entire life cycle features of the pest. Scale 0.5: Captures the wing veins or markings of adult insects (such as thrips wing veins); Scale 1.0: Identify the body wall features of larvae (nymphs) (such as the waxy layer of whitefly nymphs); Scale 2.0: Detects egg arrangement patterns (e.g., aphid egg masses). Feature extraction: HOG characteristics: cell size 8×8, number of orientations 9, describing the distribution of local gradient orientations; SIFT features: keypoint detection (Hessian matrix threshold 0.04), generating 128-dimensional descriptors; Feature fusion: The three-scale features (HOG+SIFT) are concatenated into a 2197D vector, and then reduced to 1024D by PCA to form a multi-scale feature pyramid.
[0023] Integration with environmental parameters: The feature pyramid is concatenated with temperature, humidity, and wind speed data (timestamp error <10ms) to generate a joint feature vector; the data is then normalized using the BatchNorm layer to output standardized data suitable for model layer analysis.
[0024] In one embodiment, the integration of image features, environmental features, and biological features through a cross-modal feature fusion network includes the following steps: Multimodal feature extraction preprocessing: The super-resolution reconstructed image was divided into 8×8 cell sizes, and gradient histograms in 9 directions were extracted to generate a 2048-dimensional feature vector. Keypoints were detected using the Hessian matrix to generate a 128-dimensional descriptor, which was then concatenated with the HOG features to form a 2197D vector. Temperature, humidity, and wind speed data (sampling interval of 1 minute) were input into a two-layer LSTM network (hidden layer size 64) to extract the 128D feature vector at the last time step. A pest morphology feature library (such as aphid antenna length and spider mite body color R value) was constructed based on expert knowledge, and a 512D feature vector was generated through KNN matching. Dynamic weighting: The timestamps of image features and environmental features are synchronized using the NTP protocol to ensure that the features correspond to the field conditions at the same time. The locations (pixel coordinates) of pests in the image features are mapped to the coordinate system of the environmental sensor, and local temperature, humidity, and wind speed data are associated. The modal weights (0.6 for image, 0.3 for environment, and 0.1 for organisms) are calculated using Softmax to suppress noisy modalities (such as invalid images under strong winds). The SE module is applied to each channel of the feature vector to improve the response values of key features (such as the wing veins of the whitefly). Cross-modal feature fusion network structure: Image, environmental, and biological features are input into a three-layer fully connected network (size 2048→1024→64), with ReLU activation function. The dimensionality-reduced features (1024+64+256=1344D) are concatenated into a fused feature, which is then input into a two-layer fully connected network (1344→512→256), with ReLU activation function. The loss function includes a classification loss with cross-entropy loss (labeled as pest type) and a contrast loss to increase the distance between features of different categories (margin=2.0) to improve inter-class separability. Generate joint classification probability: The fused features are input into the Softmax layer to generate N-dimensional classification probabilities (N is the number of pest species, such as aphids, spider mites, and whiteflies); the probability distribution is adjusted by temperature scaling to avoid overfitting.
[0025] In one embodiment, the dynamic federated learning mechanism includes the following steps: First, local models are trained. Each edge node independently trains a lightweight CNN model based on local data (image features, environmental parameters, and biological features), using the SGD optimizer and a loss function of cross-entropy + contrastive loss (weight 0.5). Then, the aggregation period is dynamically adjusted: the aggregation period (from 1 hour to 24 hours) is automatically adjusted according to the frequency of environmental changes (such as wind speed fluctuations greater than 2 m / s or sudden changes in temperature and humidity). The more drastic the environmental changes, the higher the aggregation frequency, ensuring that the model can quickly adapt to new scenarios. Next, encrypted parameters are aggregated. Each node uploads its model parameters to the cloud after homomorphic encryption using Paillier. The cloud calculates a weighted average (weight = local data volume / total data volume) to generate global model parameters and distribute them. The entire process does not expose the original data, protecting farmers' privacy. Finally, heterogeneous adaptation is performed. Models from different regions are integrated using a federated averaging algorithm, allowing each node to retain some personalized parameters (such as convolutional kernels for specific pests), improving the generalization ability of the global model in the identification of cross-regional pests such as aphids and spider mites.
[0026] In one embodiment, generating a real-time risk heat map by combining meteorological data with a pest and disease spread model includes the following steps: Multi-source data fusion: Meteorological data (temperature, humidity, rainfall, wind speed) is acquired in real time through the national meteorological station API and spatially interpolated with field environmental data (such as canopy humidity) collected by the sensing layer to generate a meteorological field with a 5m×5m grid; the pest and disease spread model adopts the improved SEIR model; Risk calculation: The model uses the current distribution of pests and diseases (detected via a multispectral imaging module) as initial conditions to simulate the spread path over the next 24 hours; combining meteorological field data with the spread simulation results, a weighted formula is used: Where M is the risk value, P is the diffusion coverage rate, and D is the environmental suitability; environmental suitability is determined by temperature, humidity, and rainfall (e.g., suitability increases by 30% when canopy humidity is greater than 80%). Generate a visual heatmap: The risk values are mapped to a geographic coordinate system, and a pseudo-color heat map (red for high risk and blue for low risk) is generated using OpenCV and overlaid on the field map. The heat map is pushed to the farmer's APP in real time through edge nodes, indicating high-risk areas (such as small-rained plots in Dangzhong participating in bean rotation as grub risk areas) and recommended prevention and control measures (such as timely irrigation to drown the larvae during the 1st-2nd instar of grubs).
[0027] In one embodiment, the user submits feedback on misjudgments via the APP, automatically generating adversarial examples and triggering model iteration, including the following steps: Handling user misjudgment feedback and generating adversarial examples: Farmers upload misclassified images via an app, such as ladybugs being misclassified as aphids, and select the misclassification type (missed / misclassified / location deviation). The app automatically labels the timestamp and geographical location. After the feedback image is reconstructed by a preprocessing layer using super-resolution, it is input into the WGAN-GP model to generate adversarial examples: the perturbation intensity ε=0.1 (L∞ norm constraint) ensures that the visual difference between the adversarial examples and the original images is less than 5% (the threshold that humans can distinguish). The generated samples are mixed with the original labels in a 1:5 ratio to build an incremental training dataset, providing data support for model iteration. Federal update verification: Edge nodes fine-tune their local models using incremental datasets, employing cross-entropy plus contrastive loss as the loss function. The fine-tuned model parameters are uploaded to the cloud via Paillier homomorphic encryption and federated with parameters from other regions, with the weights being the ratio of local data volume to total data volume. The updated global model is then distributed to each node, and the accuracy improvement is verified on a test set containing adversarial examples (e.g., whitefly identification rate increases from 85% to 93%). If the improvement is less than 3%, the perturbation strength is adjusted to ε=0.15 to trigger a second iteration, until the loss value is less than 0.1, forming a closed-loop optimization mechanism of "feedback-generation-iteration-verification".
[0028] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An image recognition-based Codonopsis pilosula pest identification system, characterized in that, Comprise: Sensing layer: Deploy multi-spectral imaging modules including visible light, near-infrared, hyperspectral camera and environmental perception suite including temperature and humidity, light, wind speed sensor, real-time collection of radix ginseng plant image and environmental data; Through edge computing nodes to preliminary compression and noise filtering of image, synchronous recording of sensor data and time stamp; Stream the processed multi-modal data to the preprocessing layer; Preprocessing layer: Dynamically adjust image enhancement strategy based on environmental perception data: increase contrast in low light, use blind deconvolution to remove blur in high wind speed; Super-resolution reconstruction of hidden pest images, generate multi-scale feature pyramid; Fusion of environmental parameters and image features, output standardized data suitable for complex field scene for model layer analysis; Model layer: Through cross-modal feature fusion network to integrate image features, environmental features and biological features, generate joint classification probability; Use dynamic federated learning mechanism, each regional node independently trains local model, regularly encrypts and aggregates parameters and updates global model; Dynamically adjust model aggregation period according to environmental change frequency, improve cross-region generalization ability; Application layer: Combine meteorological data and pest spread model to generate real-time risk heat map to guide precise pesticide application; Users submit false positive feedback through APP, automatically generate adversarial samples and trigger model iteration; Output performance report every month, optimize key parameters; Final output includes identification results, prevention and control suggestions and ecological impact assessment decision report.
2. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The deployment includes multi-spectral imaging modules including visible light, near-infrared, hyperspectral camera and environmental perception suite including temperature and humidity, light, wind speed sensor, comprising the following steps: Hardware installation: Install industrial-grade CCD camera at a height of 3 meters on the central column of the planting area, covering an area of 50 meters in diameter with no dead angle monitoring; Install sensors with a wavelength range of 700-1000nm through a beamsplitter for the same angle of view imaging beside the CCD; Install a push-broom hyperspectral imager on a movable track for scanning the entire field once a week to generate a hyperspectral data cube; Deploy 2 temperature and humidity sensors per mu at the height of the plant canopy to avoid direct sunlight; Install light sensors side by side to measure canopy light intensity; Install three-cup anemometer wind speed sensor in the open area of the field away from obstructions; Configure synchronous device parameters: Set the visible light camera to automatic exposure and automatic white balance, turn off digital noise reduction to preserve the details of the original image; Set the near-infrared camera to a fixed exposure time, adjust the gain to a signal-to-noise ratio greater than 40dB to improve the clarity of the back of the leaf; Configure the push-broom speed and exposure time of the hyperspectral camera to generate a data cube; Set the temperature and humidity sensor to a sampling frequency of 1 per minute, data upload period of 5 minutes, and real-time alarm for abnormal values; Configure the dynamic range adjustment of the light sensor to automatically compensate for the light changes during cloudy / sunny transitions; Set the wind speed sensor to average mode with threshold triggering; Synchronize all devices' time stamps through NTP protocol, and image and environmental data correspond one-to-one; Publish multi-spectral camera and sensor data through ROS topics; Protection calibration: The camera and sensor are encapsulated in an IP67-rated enclosure with automatic heating and defogging. A sunshade is installed on the hyperspectral camera track to reduce direct sunlight from causing the sensor to overheat. The multispectral camera is radiometrically calibrated monthly using a standard reflector to ensure consistent spectral data. The environmental sensor is compared with a standard instrument quarterly, and automatic calibration is triggered when the deviation exceeds 5%. Data transmission and storage: Visible light and near-infrared images are transmitted to edge computing nodes in real time via 5G network. The hyperspectral data cube is stored locally and then uploaded to the cloud daily. Environmental data is transmitted to the gateway via LoRaWAN protocol and then forwarded to the cloud via 4G network. Edge nodes are equipped with SSD hard drives to store high-resolution images and raw environmental data from the past 7 days as a cloud backup. The hyperspectral data cube is compressed and stored on a NAS device and retained for 30 days.
3. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The image enhancement strategy based on environmental perception data includes the following steps: Dynamic strategy selection: In low-light scenarios, the Retinex algorithm is triggered to enhance contrast and improve the visibility of tiny pests on the underside of leaves; in high-wind-speed scenarios, a blind deconvolution algorithm is applied to remove motion blur and restore the outline features of fast-moving pests. Special treatment for hidden pests: The SRCNN model was applied to the images of the underside of the leaves to increase the resolution to 2000 dpi, clearly presenting the subtle features of the insect's antennae; the reconstructed images were downsampled, and HOG features were extracted at each scale and stitched together to form a multi-scale feature vector. Fusion of environmental parameters and image features: Environmental parameters, including temperature, humidity, and wind speed, are encoded into 128D vectors and concatenated with image features to form 2197D features. Dimensionality reduction is achieved through a fully connected layer, outputting standardized data adapted to complex field scenarios for analysis by the model layer. Strategy optimization: Farmers use the app to mark misjudged cases, which automatically generates adversarial examples. These adversarial examples are mixed with the original data at a ratio of 1:5, triggering fine-tuning of the Retinex / blind deconvolution algorithm parameters. The strategy is updated within 72 hours.
4. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The super-resolution reconstruction of images of concealed pests to generate a multi-scale feature pyramid includes the following steps: Filtering images of hidden pests: The environmental perception suite detects low light or high wind speed, or the user manually marks "areas that need enhancement"; it quickly locates hidden areas on the back of leaves and in crevices of stems using a shallow CNN model, and marks the location of potential pests. Super-resolution reconstruction: Extract image patches of hidden areas; if the original resolution is less than 500 dpi, reconstruction will be automatically triggered. Algorithm flow: First convolutional layer: 9×9 filter extracts low-frequency features, outputs 64 channels, activation function ReLU; The second convolutional layer uses a 1×1 filter to perform a non-linear transformation of feature mapping, with 32 output channels. The third convolutional layer: a 5×5 filter reconstructs high-frequency details, outputting a super-resolution image with the same size as the input; Calculate the peak signal-to-noise ratio of the reconstructed image to the original high-resolution image, with a target value greater than 30 dB; Generate a multi-scale feature pyramid: The reconstructed images were downsampled at three levels with scale factors of 0.5, 1.0, and 2.0, covering the characteristics of the entire life cycle of pests: scale 0.5 captures the wing veins and patterns of adult insects, scale 1.0 identifies the body wall features of larvae, and scale 2.0 detects the arrangement pattern of eggs. Feature extraction: HOG characteristics: cell size 8×8, number of orientations 9, describing the distribution of local gradient orientations; SIFT features: key point detection, generating 128-dimensional descriptors; The three-scale features are concatenated into a 2197D vector, which is then reduced to 1024D using PCA to form a multi-scale feature pyramid. The feature pyramid is then concatenated with temperature, humidity, and wind speed data to generate a joint feature vector. The data is then normalized using a BatchNorm layer to output standardized data suitable for model layer analysis.
5. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The method integrates image features, environmental features, and biological features through a cross-modal feature fusion network. Includes the following steps: Multimodal feature extraction preprocessing: The super-resolution reconstructed image is divided into 8×8 cell sizes, and gradient histograms in 9 directions are extracted to generate a 2048-dimensional feature vector. Keypoints are detected using the Hessian matrix to generate a 128-dimensional descriptor, which is then concatenated with the HOG features to form a 2197D vector. Temperature, humidity, and wind speed data are input into a two-layer LSTM network to extract the 128D feature vector at the last time step. A pest morphology feature library is constructed based on expert knowledge, and a 512D feature vector is generated through KNN matching. Dynamic weighting: The timestamps of image features and environmental features are synchronized using the NTP protocol, so that the features correspond to the field conditions at the same time. The pixel coordinates of the pest locations in the image features are mapped to the coordinate system of the environmental sensor, and local temperature, humidity and wind speed data are associated. The modal weights are calculated using Softmax to suppress noisy modes. Apply the SE module to each channel of the feature vector to improve the response values of key features; Cross-modal feature fusion network structure: Image, environmental, and biological features are input into a three-layer fully connected network with ReLU activation function. The dimensionality-reduced features are concatenated into a fused feature, which is then input into a two-layer fully connected network with ReLU activation function. The loss function includes classification loss (cross-entropy loss) and contrast loss (to increase the distance between features of different classes and improve inter-class separability). Generate joint classification probability: The fused features are input into the Softmax layer to generate N-dimensional classification probabilities, where N is the number of mite species, including aphids, spider mites, and whiteflies. The probability distribution is adjusted by temperature scaling.
6. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The dynamic federated learning mechanism includes the following steps: First, local models are trained. Each edge node independently trains a lightweight CNN model based on local data including image features, environmental parameters, and biological features, using the SGD optimizer and a loss function of cross-entropy plus contrastive loss. Then, the aggregation cycle is dynamically adjusted: the aggregation cycle is automatically adjusted according to the frequency of environmental changes; the more drastic the environmental changes, the higher the aggregation frequency, allowing the model to quickly adapt to new scenarios. Next, encrypted parameters are aggregated. Each node uploads its model parameters to the cloud after homomorphic encryption using Paillier. The cloud calculates a weighted average to generate global model parameters, which are then distributed. The entire process does not expose the original data, protecting farmers' privacy. Finally, heterogeneous adaptation is performed. A federated averaging algorithm is used to integrate models from different regions, allowing each node to retain some personalized parameters, improving the generalization ability of the global model in identifying aphids and spider mites across regions.
7. The image recognition-based codonopsis bug identification system according to claim 1, characterized in that, The process of generating a real-time risk heat map by combining meteorological data with a pest and disease spread model includes the following steps: Multi-source data fusion: Meteorological data, including temperature, humidity, rainfall, and wind speed, are acquired in real time through the national meteorological station API. Spatial interpolation is performed with field environmental data collected by the sensing layer to generate a 5m×5m grid meteorological field. The pest and disease spread model adopts an improved SEIR model. Risk Calculation: The model uses the current distribution of pests and diseases as initial conditions to simulate the spread path over the next 24 hours; combining meteorological field data with the spread simulation results, a weighted formula is used: Where M is the risk value, P is the diffusion coverage rate, and D is the environmental suitability; environmental suitability is determined by temperature, humidity, and rainfall. Generate a visual heatmap: Risk values are mapped to a geographic coordinate system, and a pseudo-color heat map is generated using OpenCV and overlaid on a field map. The heat map is then pushed to the farmer's APP in real time through edge nodes, indicating high-risk areas and recommending prevention and control measures. 8.The image recognition-based Codonopsis pilosula pest identification system according to claim 1, characterized in that, The above includes the following steps: Handling user misjudgment feedback and generating adversarial examples: Farmers upload misjudged images via an app, selecting the misjudgment type, including missed judgment, misjudgment, and location deviation. The app automatically labels the timestamps and geographical locations. After the feedback images are reconstructed by a preprocessing layer using super-resolution, they are input into the WGAN-GP model to generate adversarial examples: the perturbation intensity ε=0.1, making the visual difference between the adversarial examples and the original images less than the human-discriminable threshold, 5%. The generated samples are mixed with the original labels in a 1:5 ratio to construct an incremental training dataset, providing data support for model iteration. Federal update verification: Edge nodes fine-tune their local models using incremental datasets, employing cross-entropy plus contrastive loss as the loss function. The fine-tuned model parameters are uploaded to the cloud via Paillier homomorphic encryption and federated with parameters from other regions, with the weight being the ratio of local data volume to total data volume. The updated global model is then distributed to each node, and the accuracy improvement is verified on a test set containing adversarial examples. If the improvement is less than 3%, the perturbation strength is adjusted to ε=0.15 to trigger a second iteration, until the loss value is less than 0.1, forming a closed-loop optimization mechanism of "feedback-generation-iteration-verification".