High-throughput end-to-end aphid honeydew discharge behavior identification method
By constructing an aphid behavior dataset, designing the RAMF algorithm and the RT-DETR-RK50 model, and combining cross-frame processing with a real-time detection platform, the problems of low efficiency and poor real-time performance in the recognition of aphid honeydew behavior were solved, and high-precision, real-time aphid behavior detection was achieved.
Patent Information
- Application Number
- CN202510931031.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies make it difficult to efficiently and in real time identify aphids' honeydew secretion behavior, especially in complex agricultural scenarios where aphids are small, the background is complex, and individuals frequently overlap. Traditional methods are inefficient and have poor real-time performance, and cannot meet the needs of high-throughput detection.
A high-throughput end-to-end aphid honeydew behavior recognition method is adopted, including constructing an aphid behavior dataset, designing a rapid adaptive motion feature fusion algorithm (RAMF), building an RT-DETR-RK50 aphid behavior detection model, a cross-frame processing mechanism and a real-time end-to-end detection platform, and achieving real-time detection through GPU acceleration and optimized memory management.
The system has achieved improved small target detection accuracy, enhanced real-time performance, and enhanced robustness to meet the needs of high-throughput detection. It has also improved the accuracy of aphid honeydew discharge behavior recognition and adapted to complex agricultural scenarios.
Smart Images

Figure CN120766355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine vision and plant protection, and in particular to a high-throughput end-to-end aphid honeydew discharge behavior recognition method. Background Art
[0002] Aphids are globally recognized as significant plant-sucking pests, causing diverse damage to agriculture and ecosystems. Their damage occurs through three primary mechanisms: direct feeding on phloem sap, hindering plant growth and development; secretion of sugar- and amino acid-rich honeydew, which induces sooty mold; and efficient transmission of plant viruses. Due to their high reproductive capacity, polyphagia, and diverse adaptability, aphids have become one of the most destructive pests in global agricultural production.
[0003] The significant threat posed by aphids to agriculture makes in-depth research on their behavior of great scientific and practical significance. Aphid feeding behavior not only reflects the insect's adaptability to its host plant but is also a key indicator for assessing plant resistance mechanisms. Research has found that accurately monitoring aphid feeding behavior can provide important guidance for pest management, resistance breeding, and crop protection.
[0004] Honeydew excretion (HE) behavior, in particular, is directly linked to feeding behavior and is a highly prominent visual feature. Therefore, detecting HE behavior is considered an ideal method for indirectly monitoring aphid feeding status and assessing plant resistance levels. Breeding aphid-resistant crop varieties is considered a more sustainable strategy for controlling these pests. Accurate monitoring and analysis of aphid resistance behavior can provide key technical support for the breeding process of aphid-resistant crop varieties, thereby creating favorable conditions for sustainable agricultural development.
[0005] Traditional methods for monitoring aphid HE behavior, such as manual visual counting or chemical analysis, suffer from low efficiency and poor real-time performance, failing to accurately capture this key characteristic of aphid activity. The rapid development of machine vision and deep learning technologies has provided new solutions to this problem.
[0006] Although some progress has been made in the research of insect behavior recognition, it is difficult to directly apply it to agricultural pests such as aphids, mainly facing the following challenges: (1) Differences between research objects and environment. Existing research focuses on laboratory model organisms such as fruit flies, which have high contrast with the background, while aphids are smaller and semi-transparent, making it difficult to distinguish from the plant background. (2) Complexity of analysis scene. Traditional research usually analyzes single insect behavior under ideal conditions such as culture dishes, but in actual agricultural scenes, aphids often attach to tobacco, cotton or wheat leaves in dense groups, with complex background and frequent overlap between individuals. (3) Speciality of behavior recognition target. Existing researches mostly focus on insect behaviors such as grooming, and rarely involve behaviors with ecological and agricultural significance, especially the HE behavior of aphids, which directly represents the feeding state of aphids. This behavior is not only subtle and fleeting, but also lacks a specialized recognition method.
[0007] In addition, the existing technical method also has limitations in identifying the honeydew discharge behavior of aphids: (1) The motion feature extraction process based on frame difference method is slow and noisy, and the feature map captured by the camera is significantly different from the original RGB video frame, resulting in loss of key spatial information; (2) The two-stage detection process has poor real-time performance and cannot meet the high-throughput detection requirements. Therefore, it is urgent to develop a high-throughput end-to-end aphid honeydew discharge behavior recognition system. SUMMARY
[0008] In view of the problems of low efficiency and poor real-time performance of traditional manual and chemical analysis for aphid honeydew detection, the present application proposes a high-throughput end-to-end intelligent recognition method for aphid honeydew discharge behavior.
[0009] To achieve the above purpose, the present application adopts the following technical scheme: A high-throughput end-to-end aphid honeydew discharge behavior recognition method, comprising the following steps: S1, constructing an aphid behavior dataset: collecting video data containing different developmental stages of aphids, and constructing a fine-grained dataset of aphid crawling, kicking and honeydew discharge behavior under different population density and light intensity conditions; S2, designing a rapid adaptive motion feature fusion algorithm (Rapid Adaptive Motion-Feature Fusion, RAMF): through a 10-frame time window, exponentially increasing weight distribution, adaptive threshold design and GPU acceleration technology, motion feature extraction at 45 frames per second is achieved; S3, constructing an RT-DETR-RK50 aphid behavior detection model: integrating a Kolmogorov-Arnold network module (KAN) at the 4th and 5th stages of the ResNet50 backbone network, replacing the standard bottleneck structure, introducing an activation function based on B-spline basis function, constructing an RK50 module, and strengthening the detection model's ability to capture complex spatial relationships and subtle features; S4. Build a cross-frame processing mechanism: odd-numbered frames detect motion behavior, even-numbered frames identify honeydew targets, combine delayed interpolation post-processing algorithm to eliminate detection flicker, and adopt a hierarchical three-stage analysis strategy to identify honeydew drainage behavior; S5. Build a real-time end-to-end detection platform: Through GPU memory management optimization, persistent CUDA stream parallel execution and other technologies, achieve complete real-time end-to-end processing from video input to behavior detection output.
[0010] Preferably, the RAMF algorithm in step S2 designs an exponentially increasing weight distribution scheme, as shown in formula (1), to ensure that the weight of the latest frame in the time window is higher;
[0011] Where ω(t) represents the time weighting coefficient, t represents the frame index in the time window, and Indicates the size of the time window.
[0012] Preferably, in step S3, the RK50 module realizes kernel adaptation through B-spline basis function, and the activation function The expression of is shown in formula (2):
[0013] Where b(x) is the basis function (usually the basis function is the Sigmoid linear unit), Bi(x) is the B-spline basis function, and ci is the trainable coefficient.
[0014] Preferably, the KAN convolution layer of the RK50 module in step S3 decomposes the convolution into two parallel paths: one is a basic path, which applies a conventional convolution operation to the input transformed by the basic activation function; the other is a spline path, which applies a B-spline basis transformation to the input before convolution; the output of a single group of KAN convolution layers can be expressed as formula (3):
[0015] Where g(⋅) represents the basic activation function, B(x) refers to the spline basis transformation, and Norm is the normalization function.
[0016] Preferably, in step S4, a motion-original video cross-frame detection method is implemented, odd-numbered frames process motion features to identify crawling and kicking behaviors, and even-numbered frames use original optical features to detect translucent honeydew targets. A delayed interpolation algorithm (single-frame results are extended by 1 frame forward and backward) is used to eliminate flickering in detection results, reducing computational overhead by 50%.
[0017] Preferably, in step S4, a hierarchical three-stage analysis strategy is adopted to divide the determination of honeydew excretion behavior into three stages: the first stage is to identify high-frequency micro-movements before honeydew excretion through low-frequency detection; the second stage is to perform honeydew identification and detect physical evidence of the generation of metabolites in the form of single droplets; the third stage is to carry out composite motion evaluation, and associate leg movements with honeydew droplets based on time continuity, and merge them into honeydew excretion behavior.
[0018] Preferably, step S5 also includes constructing a comprehensive acceleration framework, using a multi-threaded parallel pipeline architecture to decouple the serial stages of traditional video processing into concurrent threads of frame acquisition, feature extraction and result rendering, and reducing runtime fragmentation through GPU memory pre-allocation strategy, using persistent CUDA streams to realize asynchronous parallel computing of motion feature extraction and target detection, using batch optimization to reduce task startup overhead, combining the CUDA math library to vectorize and accelerate key operations, and at the same time using just-in-time compilation technology and mixed precision computing to reduce the amount of calculation while maintaining accuracy. The end-to-end real-time behavior recognition system is finally built to achieve real-time end-to-end processing from video input to behavior detection output. At 1080p resolution, the real-time processing speed of RTX4090 reaches 31.82fps.
[0019] Compared with the prior art, the present invention has the following beneficial effects: This paper constructs a high-throughput, end-to-end small-target, fine-grained behavior detection model. Through multi-dimensional technological innovation, it achieves significant improvements in detection accuracy, real-time performance, and robustness in the fields of machine vision and plant protection. Compared with existing technologies, the specific beneficial effects are as follows: 1. Significantly improved small target detection accuracy: Through a fast adaptive motion feature fusion algorithm and RT-DETR-RK50 model optimization, an average detection accuracy of 85.9% was achieved. Compared with models without the RK50 module, mAP50 increased by 2.9%. In particular, the performance of small target detection, such as translucent honeydew droplets (approximately 0.1-0.5mm in diameter), significantly outperformed mainstream algorithms (such as the YOLO series and SSD), addressing the problem of traditional methods' insufficient ability to recognize subtle behaviors.
[0020] 2. End-to-end real-time detection: Through optimizations such as a multi-threaded parallel pipeline architecture (concurrent execution of frame acquisition, feature extraction, and rendering), GPU memory pre-allocation, and persistent CUDA streams, the RTX4090 achieves 31.82 fps per second for the entire process at 1080p resolution. This meets the millisecond-level response requirements from video stream input to detection result output, improving efficiency by 50% compared to the traditional two-stage detection process.
[0021] 3. Enhanced scene robustness: Under brightness variations from -80 to 40 lux, the HE recognition accuracy is >80%, the detection rate is >95% in low-density aphid scenes, and the missed detection rate is <10% in high-density scenes (>50 aphids / field of view). Even under high-density aphids and low-light conditions, the detection robustness remains above 80%, overcoming the challenges of complex backgrounds and overlapping aphids in natural scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a specific flow chart of the present invention; Figure 2 Dataset images for different population densities and light intensity conditions; Figure 3 Provide a RAMF processing workflow and visualization diagram; Figure 4 The distribution of effective behavioral samples for aphids; Figure 5 This is the RT-DETR-RK50 network architecture diagram; Figure 6 This is the network architecture diagram of the conventional module and RK50 module; Figure 7 This is the KAN Conv network model diagram; Figure 8 Visual representation of motion feature extraction under different time windows; Figure 9 Comparison of feature activation heatmaps for different RT-DETR-RK50 variants; Figure 10 Comparison of the confusion matrices of detection models for different aphid behavior categories; Figure 11 Comparison of training loss curves for different model architectures; Figure 12 This is the detection result diagram based on the highest mAP50 indicator model; Figure 13 Visualization results of Grad-CAM++ under different lighting conditions and aphid densities; Figure 14 The diagram shows the honeydew detection process and results in stages. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0024] Example Referring to Figure 1-14 The application provides a high-throughput end-to-end aphid honeydew discharge behavior recognition method, comprising the following steps: S1, constructing an aphid behavior dataset: collecting video data containing aphids at different developmental stages, and constructing an aphid crawling, kicking and honeydew discharge behavior fine-grained dataset covering different population density and light intensity conditions.
[0025] The experimental dataset was collected in the modern greenhouse facilities of the College of Plant Protection of Henan Agricultural University from August to December 2024. The test object was Myzus persicae reared on tobacco leaves, including aphids at different developmental stages. The data collection used a standardized process, and was completed using a high-resolution microscopic imaging system (Sony ILCE-7RM2 camera with Laowa 25mm f / 2.8 ULTRA MACRO 2.5-5X lens, resolution 1920x1080 pixels, sampling rate 30 frames / second) in the same controlled environment. All records were saved in color JPEG format, which optimizes storage efficiency while preserving visual details.
[0026] The dataset covers different population density and light intensity conditions (such as Figure 2 As shown), including 32 independent experimental video groups, each corresponding to a specific light intensity condition, with a single group duration of 30 minutes, and a cumulative effective observation time of 16 hours.
[0027] S2, design a rapid adaptive motion feature fusion algorithm (Rapid Adaptive Motion-Feature Fusion, RAMF): through a 10-frame time window, exponentially increasing weight distribution, adaptive threshold design and GPU acceleration technology, realize 45 frames of motion feature extraction per second.
[0028] To effectively extract the motion feature information of aphids, a rapid adaptive motion feature fusion (RAMF) algorithm suitable for motion target detection is designed, which is realized by the following four steps, and its processing flow and corresponding heat map visualization are shown in Figure 3 .
[0029] 1. Time motion feature processing To accurately capture the motion features of the target, the RAMF algorithm uses a sliding time window based on the frame difference method to process consecutive video frames. The resolution and time window size are determined as key parameters in the frame difference method, which directly affects the extraction of motion features and thus the accuracy of behavior recognition. To determine the optimal time window length, the application explores the influence of different time window lengths T=5, 7, 9, 11, 13, 15, 17 (interval 2 frames) under the premise of keeping the 1080p resolution fixed, and the results are shown in Figure 8 .
[0030] By analyzing the frame differences of a continuous video sequence F and visualizing its motion features, experimental results show that a window length between 9 and 11 frames achieves the optimal balance: it avoids significant motion smearing while ensuring high feature clarity. Therefore, the average value of T = 10 frames is selected as the final window length.
[0031] Within a time window, the differences between adjacent frames are calculated using the following weighted method:
[0032] In this formula, △t[F] represents the weighted difference image of the video sequence F at time t, where F(t) represents the image of the video sequence F at time t, Denotes the difference operator, ω(t) denotes the time weighting coefficient. Here, the difference operator is defined as the absolute difference of pixel intensities:
[0033] In order to highlight the contribution of the latest frame to the current motion state, the RAMF algorithm designs an exponentially increasing weight distribution scheme to reflect the criticality of the latest frame.
[0034]
[0035] In this formula, t represents the frame index within the time window, and Indicates the size of the time window.
[0036] 2. Motion Saliency Calculation In order to characterize the large motion of the target area, the RAMF method first sums the weighted frame differences within the time window and then constructs the initial motion map .
[0037]
[0038] here, Refers to pixels Δt[F] is the weighted difference value at pixel x in the video sequence F at time t. Since background noise can cause interference in real scenes, RAMF proposes an adaptive threshold mechanism based on statistical characteristics to set the threshold.
[0039]
[0040] in and Represents motion graphs The mean and standard deviation of Is an adjustable parameter, which controls the sensitivity of the threshold. is set to 3, which corresponds to 3 Criterion, according to which the motion map The generation of can be expressed by the following formula.
[0041] Based on this threshold, the motion map The generation process can be expressed as:
[0042] Through this threshold mechanism, the background noise can be effectively suppressed while highlighting the area containing the real moving target.
[0043] 3. Adaptive feature fusion based on motion intensity In order to enhance the continuity of feature representation and reduce local noise, the RAMF method performs Apply a Gaussian filter:
[0044] in represents a Gaussian kernel with a covariance matrix of ∑, * represents a convolution operation, and Ω represents an image domain. In the present invention, The size of is 7×7, and ,satisfy conditions, is the identity matrix.
[0045] In order to effectively highlight the motion area while retaining the original image details, the RAMF method designs an adaptive feature fusion scheme based on weighted mixing:
[0046] in represents the original image, Color visualization results representing motion saliency (the yellow channel is used in this invention), is the adaptive mixing coefficient.
[0047] Mixing coefficient Calculated by an adaptive mechanism based on exercise intensity:
[0048] in It is a scaling factor used to adjust the effect of motion significance on the fusion coefficient. is the maximum value of the image pixel intensity (such as 255), is the upper limit of the mixing coefficient to maintain the visibility of the original image information. Set to 1.2, Set to 0.8.
[0049] This adaptive strategy dynamically adjusts the blending ratio to ensure that the original image information is not excessively obscured while also achieving a reasonable enhancement of motion features. Specifically, areas with higher motion intensity receive a higher blending coefficient, with an upper limit of 0.8 to maintain the visibility of the original image information.
[0050] 4. Accelerated processing design Under the conditions of 1080p resolution and 10-frame time analysis window, the RAMF algorithm uses the following acceleration processing to achieve motion feature extraction at 45 frames per second.
[0051] Through parallel processing of intensive image tasks using GPU tensor operations, the pixel-level computing load is transferred from the CPU to the thousands of parallel cores of the GPU; fixed memory technology is used to achieve efficient CPU-GPU data transmission and reduce data handling delays; asynchronous CUDA streams are used to overlap computing and data transmission operations to hide I / O overhead; batch processing and pre-allocation of GPU resources minimize task startup overhead to avoid frequent request / release of video memory; dedicated GPU parallel computing libraries such as the CUDA math library are used to accelerate key operations such as matrix multiplication and convolution. These optimizations work together to break through the real-time and resolution bottlenecks of traditional algorithms, making the RAMF algorithm suitable for real-time monitoring scenarios while maintaining high accuracy. The motion features of aphids extracted by the RAMF algorithm are as follows: Figure 4 shown.
[0052] S3. Construct the RT-DETR-RK50 aphid behavior detection model: Integrate the Kolmogorov-Arnold network module (KAN) in the 4th and 5th stages of the ResNet50 backbone network, replace the standard bottleneck structure, introduce an activation function based on the B-spline basis function, and construct the RK50 module to enhance the detection model's ability to capture complex spatial relationships and subtle features.
[0053] RT-DETR initially used the standard ResNet-50 (R50) as the backbone network for feature extraction. This network uses a layered structure with multiple bottleneck modules stacked in multiple stages. However, the fixed core architecture and rigid activation function limit the model's ability to capture complex spatial relationships and subtle features, which are critical for small object detection.
[0054] To break through the above limitations, the application proposes a new backbone network architecture RT-DETR-RK50, the core of which is to strategically integrate the Kolmogorov-Arnold network (KAN module) into the key deep stage of the ResNet-50 backbone network. Specifically, under the premise of retaining the original architecture in the early stage, the KAN enhanced BottleNeck_KAN module (containing KAGNConv2DLayer) is used to replace the standard BottleNeck module, focusing on optimizing the 4th and 5th stages (as shown in Figure 5 ). The specific architecture of the standard BottleNeck module and the RK50 module is as shown in Figure 6 .
[0055] The core innovation of this structure is to realize kernel adaptability through B-spline basis functions, enabling the network to learn more complex nonlinear feature transformations beyond traditional convolution linear combinations. The expression of the activation function is shown in equation (10), which is parameterized using B-spline.
[0056]
[0057] In the formula, b ( x ) is the basis function (usually this basis function is a Sigmoid linear unit), Bi ( x ) is the B-spline basis function, ci and is a trainable coefficient.
[0058] The KAN convolution layer (KAN Conv) architecture is as shown in Figure 7 , and the working mechanism is to decompose convolution into two parallel paths: one is the basic path, which applies a regular convolution operation to the input after a basic activation function transformation; the other is the spline path, which applies a B-spline basis transformation to the input before convolution. The output of a single KAN convolution layer can be represented as equation (11):
[0059] In the formula, g(⋅) represents the basic activation function, B(x) refers to the spline basis transformation, and Norm is the normalization function.
[0060] KAN Conv improves the function approximation ability through the combination of learnable activation functions, which can more efficiently model complex patterns in data compared to fixed activation functions. The spline-based method can adaptively learn the feature transformation of specific data distribution without relying on predefined activation functions.
[0061] KANConv also supports group convolution, allowing different activation functions to be learned for different feature sets, further enhancing the model's representational capabilities. A dropout mechanism designed specifically for the convolutional dimension provides efficient regularization during training. Essentially, KANConv leverages the mathematical foundations of KAN to enhance the expressive power of convolution operations while maintaining computational efficiency. This makes it particularly suitable for spatiotemporal applications requiring complex function approximation, such as image processing, video analysis, and signal processing.
[0062] Because feature maps at the P4 and P5 stages are reduced in size to 1 / 16 and 1 / 32 of the original image, respectively, they are abstract and contain high-level semantic information. Therefore, the RK50 module is more suitable for processing and enhancing these high-level features. By strategically strengthening the representation capabilities that have the greatest impact on detection performance, RT-DETR-RK50 builds a more efficient feature extraction process without significantly increasing the computational burden.
[0063] S4. Build a cross-frame processing mechanism: odd-numbered frames detect motion behavior, even-numbered frames identify honeydew targets, combine delayed interpolation post-processing algorithm to eliminate detection flicker, and adopt a hierarchical three-stage analysis strategy to identify honeydew drainage behavior.
[0064] Aphid honeydew excretion is a key indicator of population vitality, characterized by small, elusive movements. While frame-subtraction methods successfully extract excellent motion features, the optical properties of aphid honeydew are often obscured by superimposed optical flows, making accurate identification difficult under these complex features. Therefore, this paper proposes a motion-based cross-frame detection method for raw video. This method processes only odd frames during the motion feature extraction phase, enabling detection of crawling and kicking behaviors in odd frames while simultaneously detecting translucent honeydew targets using raw optical features in even frames.
[0065] This cross-frame processing can cause "flickering" in detection results (i.e., discontinuity between adjacent frames). To address this issue, the present invention proposes an improved post-processing algorithm using delayed interpolation, based on the assumption of temporal continuity. If the same behavior is detected in consecutive motion frames (e.g., frames 1 and 3), it is reasonably inferred that the intermediate original video frame (frame 2) also contains the same behavior. By extending the detection results of a single frame forward and backward by one frame, this effectively eliminates flickering in detection results and reduces computational overhead by 50%.
[0066] Aphid honeydew excretion behavior consists of static honeydew secretion (highly reflective droplets) and dynamic pedaling movements (periodic motion trajectories). To address this spatiotemporal heterogeneity, this study employed a hierarchical three-stage analysis strategy to determine honeydew excretion behavior. The first stage involves identifying high-frequency, micro-movements preceding honeydew excretion through low-frequency detection. The second stage involves honeydew identification, detecting physical evidence of metabolite production in the form of single droplets. The third stage involves complex motion assessment, associating leg movements with honeydew droplets based on temporal continuity and combining them into honeydew excretion behavior. In actual detection, if low-frequency motion and honeydew excretion occur simultaneously in adjacent frames within a time window t, this is identified as a complex motion and labeled as "honeydew excretion," thereby establishing a more accurate vitality assessment model.
[0067] S5. Build a real-time end-to-end detection platform: Through GPU memory management optimization, persistent CUDA stream parallel execution and other technologies, achieve complete real-time end-to-end processing from video input to behavior detection output.
[0068] This paper constructs a comprehensive acceleration framework that uses a multi-threaded parallel pipeline architecture to decouple the serial stages of traditional video processing. By implementing concurrent frame acquisition, feature extraction, and result rendering threads, the system effectively masks I / O latency and maximizes computing resource utilization. The specific optimization strategies are as follows: (1) GPU memory management optimization: Reduce runtime memory fragmentation through strategic pre-allocation technology, minimize the overhead of frequent memory operations, and improve memory access efficiency. (2) Persistent CUDA stream parallel execution: Maintain fixed CUDA streams to support parallel asynchronous computing, so that motion feature extraction and target detection inference can be performed simultaneously, fully utilizing the parallel computing capabilities of the GPU. (3) Batch processing optimization: Use batch processing technology to process multiple frames concurrently (the batch size can be up to 24 frames), replacing the traditional frame-by-frame sequential processing mode, significantly reducing computing latency. (4) Vectorized operation: Use vectorized computing for key image processing operations, replacing traditional pixel-by-pixel processing, further exploring GPU parallelism, and accelerating the image processing process. (5) Just-in-time compilation and mixed-precision computing: The detection network is converted into an optimized intermediate representation through just-in-time compilation, which greatly reduces the interpreter overhead; combined with the mixed-precision computing strategy, the amount of computation is reduced while maintaining the detection accuracy.
[0069] By comprehensively applying these technologies, the present invention successfully constructed an efficient end-to-end real-time behavior recognition system, overcoming the limitations of the traditional frame difference method, which is cumbersome to process and has poor real-time performance, and providing a more practical solution for high-definition video behavior analysis of aphids.
[0070] The excellent performance of the RT-DETR-RK50 detection method proposed in the present invention is verified by specific experiments below.
[0071] 1. Ablation Experiment Study of KAN Module Integration in RT-DETR-RK50 To strike a balance between computational efficiency and detection performance, this paper optimizes the deployment strategy of the KAN module in the RT-DETR-RK50 architecture. Although the KAN module's powerful adaptive activation and nonlinear modeling capabilities can significantly improve the performance of the ResNet-50 backbone network, integrating too many KAN modules can lead to a series of problems, including increased computational complexity, increased optimization difficulty, the risk of overfitting, redundant feature representations, chaotic feature hierarchies, unstable gradient flows, and unbalanced resource utilization.
[0072] To address these challenges, we systematically conducted ablation experiments to evaluate the impact of integrating different numbers of KAN modules into the RT-DETR-RK50 architecture on detection performance. By comparing configurations such as RK50-1 (with one KAN module) and RK50-3 (with three KAN modules), we aimed to determine the optimal number and distribution of KAN modules, thereby maximizing detection performance while ensuring computational efficiency.
[0073] 1.1 Optimization of RT-DETR-RK50 Feature Representation Based on Heatmap To verify the effectiveness of feature extraction in ablation studies, we used GradCAM++ for heatmap visualization analysis. The specific parameters were as follows: the method was set to GradCAMPlusPlus, and the 19th layer was set to the 15th, 22nd, and 25th layers, respectively. The backpropagation type (Backward_type) was set to null, the confidence threshold was set to 0.2, and the scale was set to 1.0.
[0074] Figure 9 The heat map visualization results reveal distinct activation characteristics between different RT-DETR backbone network variants: the baseline R50 model exhibits isolated hotspot areas with extremely low connectivity between areas, indicating that although it can effectively recognize features, its contextual association ability is limited; with the gradual integration of KAN modules, the feature representation has changed significantly, among which RK50-1 shows an expanded activation field and preliminary bridging between hotspots, suggesting that its spatial relationship modeling ability has been enhanced, although the activation distribution is still not optimal; RK50-2 presents the most balanced activation topology, characterized by uniform hotspot distribution, comprehensive connectivity and smooth transitions between high and low activation areas. This pattern indicates that it has formed a complex feature hierarchy and achieved optimal integration of contextual information; in contrast, although RK50-3 maintains a strong main activation hotspot, the overall activation pattern is more restricted, suggesting that there may be problems of feature representation redundancy and over-specialization.
[0075]
[0076] Table 1 Ablation experiment results of KAN module integration The quantitative metrics in Table 1 corroborate these visual observations. RK50-2 achieves superior overall detection performance (mAP0.5 of 0.849) with remarkable consistency across behavioral categories, particularly on the challenging honeydew category (mAP0.5 of 0.859 compared to 0.798 for R50). While RK50-3 shows a slight improvement on the crawling (CL) dataset (0.862), its performance on the kicking (LF) dataset degrades significantly (0.693). Notably, RK50-2 achieves a significant improvement in detection capability while maintaining computational efficiency, despite only a 10.6% increase in latency over the baseline model (27.2ms vs. 24.6ms).
[0077] Considering the limited number of training samples for the honeydew class (only 459), relying solely on the mAP metric may not fully reflect the model's performance on minority classes. This paper introduces the F1 score metric to effectively measure the accuracy and completeness of the model's recognition of minority classes.
[0078] As shown in Table 1, the F1 score trends of each model on the honeydew category are generally consistent with mAP@0.5 (Honeydew), indicating a certain correlation between the F1 score and mAP. The RK50-2 model achieved the highest F1 score (0.847) on the honeydew category, demonstrating a good balance between precision and recall, effectively identifying the honeydew category and reducing missed detections (false negatives) and false detections (false positives). In contrast, the R50 model achieved a lower F1 score (0.765) on the honeydew category, possibly due to a greater susceptibility to class imbalance, resulting in some bias in its recognition of the honeydew category.
[0079] 1.2 Analysis of Behavior Classification Accuracy Based on Confusion Matrix Figure 10 The confusion matrix reveals the classification performance of each detection model for three key aphid behaviors (CL, LF, and HE). The RT-DETR-R50-2 model demonstrated excellent classification accuracy, with significantly higher diagonal element values than other variants. In particular, in the honeydew category, the model achieved 86% accuracy, significantly outperforming other RT-DETR backbone network variants (R18: 72%; R50: 74%; R50-1: 69%; R50-3: 81%). Progressive optimization from R18 to each RK50 variant demonstrated a systematic improvement in behavioral differentiation. RK50-2 achieved the best balance in reducing misclassification rates for the three behavioral categories, effectively addressing common issues with other models, such as misclassifying honeydew as background and misidentifying LF as CL.
[0080] 1.3 Analysis of the impact of KAN module integration on model convergence Figure 11 The loss curve in (a) shows that RK50-2 is the optimal RT-DETR-RK50 variant, with a steady-state loss as low as 0.05 and the fastest convergence, particularly in the first 50 epochs. The baseline model, RK50, has a steady-state loss of 0.07. RK50-1, which integrates only one KAN module, shows limited improvement. Notably, despite the addition of KAN modules, RK50-3 performs worse than RK50-2, exhibiting oscillations in the later stages of training, confirming the hypothesis that excessive nonlinearity leads to optimization difficulties. The smooth curve of RK50-2 indicates a more regular loss profile and a more stable optimization process. It continues to improve even when other variants plateau between epochs 75 and 150, confirming that integrating two KAN modules into the RT-DETR backbone achieves an ideal balance between model expressiveness and training stability.
[0081] 2. Comparative Experiments of RT-DETR-RK50 and Current Cutting-Edge Methods To ensure fairness in the evaluation, all experiments were conducted on the same hardware environment equipped with an NVIDIA RTX4090 GPU. (Due to framework compatibility restrictions, SSD and Faster-RCNN were trained on RTX3090.) The study selected representative cutting-edge detection methods, including the YOLO series, the DETR family, Faster R-CNN, and SSD, for comparison.
[0082] As shown in Table 1, while the YOLO family excels in detection speed and model efficiency, its detection accuracy is relatively limited. The proposed method achieves the best overall performance, achieving a significant 2.9% improvement in mAP over the baseline model (RT-DETR-R50), validating its effectiveness. In real-world deployment scenarios, balancing accuracy and efficiency is crucial. RT-DETR-R18 emerged as the best choice in the DETR family, achieving an average mean average precision (mAP) of 82.8% and a latency of 15.9 milliseconds. YOLOv12-M and YOLOv11-L demonstrate an excellent accuracy-speed trade-off within the YOLO family, with YOLOv12-M being particularly suitable for lightweight deployments.
[0083] Figure 12 This visualization shows a comparison of the detection results of five mainstream models for aphid behavior detection. From left to right, they are YOLOv11L, YOLOv12M, RT-DETR-R18, RT-DETR-R50, and the proposed RT-DETR-RK50-2 model. This systematic comparison shows that the RT-DETR-RK50-2 model significantly outperforms the other models in various complex scenarios.
[0084] In scenes with high aphid density, the RT-DETR-RK50-2 demonstrated superior detection and classification capabilities compared to other models. While the YOLO series models exhibited classification errors and missed detections, the RT-DETR series delivered more accurate results and higher confidence scores, consistent with its superior mAP metric. This accuracy advantage stems from the RT-DETR series architecture's prioritization of detection accuracy over speed. While the RT-DETR-R50 performed reasonably well in certain scenarios, its stability declined significantly in complex backgrounds and areas with dense aphid populations, as shown in the second row of the figure, where a significant number of targets were missed. Quantitative results show that the RT-DETR-RK50-2 consistently maintained high confidence scores (mostly above 0.85) under similar conditions, while the comparison models exhibited lower or unstable confidence scores. This directly validates the effectiveness of the KAN module in enhancing feature extraction depth and representation capabilities.
[0085] Figure 10 The confusion matrix reveals significant performance differences among the models for aphid behavior recognition. The YOLO family of models exhibited high honeydew misclassification rates (YOLOv10x: 52.0%, YOLOv11-L: 57.0%, YOLOv12-M: 43.0%), demonstrating that despite their speed advantage, they suffer from fundamental limitations in small object detection. In contrast, RT-DETR-R50-2 achieved 86.0% honeydew detection accuracy, far exceeding the best YOLO family performance. As a key indicator of aphid infestation severity, this key advantage in honeydew detection demonstrates significant practical value.
[0086] Figure 11 The subgraphs further demonstrate these differences through training dynamics. In Figure (b), RK50-2 demonstrates the best convergence, achieving the lowest final loss (approximately 0.05) and converging faster than the R18 and R50 baseline models. In Figure (c), the YOLO family of models, despite a rapid initial drop in loss, eventually stabilizes at a significantly higher loss level (0.9-1.7). These patterns confirm RT-DETR's inherent advantages in fine-grained detection tasks, and the KAN module further enhances this performance.
[0087]
[0088] Table 2 Detection results compared with advanced models comprehensive Figure 12The visualization results and the quantitative analysis data in Table 2 indicate that RT-DETR-RK50-2 achieves an optimal balance of performance in aphid behavior detection tasks—not only achieving the highest mAP50 index (0.88) in honeydew detection, but also maintaining consistent excellent performance in other behavioral categories. This comprehensive and robust detection capability makes it particularly suitable for deployment in aphid behavior monitoring systems in real agricultural environments, providing reliable and accurate phenotyping support for accelerating the breeding of resistant crop varieties.
[0089] 3. Model robustness analysis under aphid density and light conditions To further evaluate the model performance, we used Grad-CAM++ to generate heat maps and conduct supplementary visualization experiments to analyze the model's performance under different lighting conditions and aphid densities ( Figure 13 Specifically, in this paper, experimental group (a) simulated the effects of different brightness levels, using brightness gains of -80, -40, 0, and 40. Experimental group (b) evaluated the model's performance under different aphid densities. Experimental group (c) focused on the HE category, testing its performance under various combinations of brightness and density.
[0090] The experimental results are shown in the figure. In group (a), the activation area of the heat map under normal brightness (brightness gain is 0) is the largest and most concentrated, indicating that the detection performance of the present invention is the best under this condition. As the brightness increases or decreases, the detection performance decreases slightly, but the overall difference is small, indicating that the RT-DETR-RK50 model of the present invention exhibits good robustness to brightness changes. In group (b), the present invention shows that when the aphid density is low, the model can detect most targets normally, but there is one missed detection in the high-density scene, indicating that the model of the present invention has certain limitations in high-density detection. In group (c), the present invention found that although the model can still detect targets under high-density and low-brightness conditions, the activation area of the heat map is smaller. This may be because in the high-density state, the proportion of the HE target category in the total pixels of the image is further reduced, resulting in increased difficulty for the model of the present invention to recognize it.
[0091] 4. Multi-stage analysis of aphid honeydew excretion behavior like Figure 14 As shown in the figure, the cross-frame detection method successfully decomposes aphid honeydew excretion behavior into three different detection stages. Experimental results show that the system is capable of identifying complex behavioral patterns through temporal analysis.
[0092] In stage 1, the system detected periodic leg-flicking (LF) behaviors with a confidence of 0.61 in case 1 and 0.89-0.92 in case 2. The blue bounding boxes effectively located the characteristic leg movements of the aphid during honeydew excretion. Stage 2 focused on identifying tiny honeydew droplets, achieving confidence levels of 0.78 and 0.88 in cases 1 and 2, respectively. The blue bounding boxes in the center column accurately labeled these highly reflective droplets on the plant substrate. In stage 3, the system combined the temporal correlation between LF behavior and honeydew presence, achieving honeydew excretion detection (purple boxes) with a confidence level of 0.83-0.86, while maintaining detection of discrete LF behaviors (confidence of 0.64 in case 2). The images in the right column demonstrate the system's ability to distinguish complex behavioral patterns from individual behavioral components within the same frame, highlighting the advantages of object detection for high-resolution multi-target behavioral recognition.
[0093] Time series detection results confirm that the proposed method successfully addresses the challenges associated with detecting aphid honeydew excretion. By decomposing this complex behavior into its components and leveraging their temporal correlation, the system achieves accurate recognition of both the physical movements (LF) and the physical evidence (honeydew droplets) associated with honeydew excretion events.
[0094] High confidence scores across all stages and cases validate the effectiveness of the cross-frame processing approach in capturing subtle behavioral indicators of aphid population vitality. Figure 9 The visualization results show that even small motions and tiny objects can be reliably detected and classified through the implemented hierarchical detection framework.
[0095] 5. Overall real-time end-to-end detection performance of the system This paper optimizes the detection process by implementing streaming inference that synchronizes feature extraction with detection, enabling the RT-DETR series of models to achieve real-time end-to-end detection. The study evaluated the performance of the three best-performing models based on ResNet18, ResNet50, and an improved RK50 backbone network.
[0096]
[0097] Table 3 Real-time end-to-end detection performance As shown in Table 3, these models were tested for both pure inference speed and real-time end-to-end processing speed on three GPU platforms (RTX3090 and RTX4090) at different resolutions (480p to 1080p), with a time window size set to 10 frames. Experimental results show that although the proposed RK50 model has slightly lower inference speed than the baseline model, it significantly improves detection accuracy (mAP50 reaching 0.849), demonstrating significant performance advantages over the baseline method.
[0098] Architectural optimizations have brought significant performance improvements to the entire model sequence. At 1080p resolution, RT-DETR-RK50 maintained a processing speed of 31.82 frames per second on the RTX4090, exceeding the real-time processing threshold of 30 frames per second. It is worth noting that the data shows that increasing the amount of computation actually improves throughput, especially at lower resolutions - RT-DETR-RK50 on the RTX4090 achieved 50.49 frames per second at 480p resolution and 41.99 frames per second at 720p resolution. This improvement is due to efficient resource utilization achieved through workload distribution through multi-threading and CUDA streams, while overlapping I / O with computation, minimizing GPU idle time, and converting sequential execution bottlenecks into concurrent operations.
[0099] Architectural optimizations have resulted in significant performance improvements across the entire model family, particularly on more powerful hardware. RTX4090 configurations generally outperform RTX3090 by approximately 10-25%. For the same model at 1080p resolution, the RTX4090 achieves a 34% improvement in processing speed (31.82 fps) compared to the RTX3090 (23.73 fps), highlighting the architecture's efficient scalability of computing resources. Experiments demonstrate that a well-designed parallel system enables real-time, end-to-end detection with high accuracy (RT-DETR-RK50 achieves a mAP50 of 0.849) even at resolutions previously considered challenging to achieve both feature extraction and object detection. Comparing the accuracy and inference speed of different models demonstrates the flexibility to adapt models to specific scenarios within the high-throughput, end-to-end RT-DETRs aphid behavior detection framework, expanding its application in areas such as high-fidelity video analysis.
[0100] While the embodiments of the present invention have been described above, the above description is intended to be exemplary, not exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A high-throughput end-to-end aphid honeydew discharge behavior identification method, characterized by: The following steps are involved: S1. Construct an aphid behavior dataset: Collect video data of aphids at different developmental stages and construct a fine-grained dataset covering aphid crawling, kicking, and honeydew discharge behaviors under different population densities and light intensities. S2. Design a Rapid Adaptive Motion-Feature Fusion (RAMF) algorithm: This algorithm uses a 10-frame time window, exponentially increasing weight assignment, adaptive threshold design, and GPU acceleration to achieve motion feature extraction at 45 frames per second. S3. Build the RT-DETR-RK50 aphid behavior detection model: Integrate the Kolmogorov-Arnold network module (KAN) in the 4th and 5th stages of the ResNet50 backbone network, replace the standard bottleneck structure, introduce an activation function based on the B-spline basis function, and build the RK50 module to enhance the detection model's ability to capture complex spatial relationships and subtle features; S4. Build a cross-frame processing mechanism: odd-numbered frames detect motion behavior, even-numbered frames identify honeydew targets, combine delayed interpolation post-processing algorithm to eliminate detection flicker, and adopt a hierarchical three-stage analysis strategy to identify honeydew drainage behavior; S5. Build a real-time end-to-end detection platform: Through GPU memory management optimization, persistent CUDA stream parallel execution and other technologies, achieve complete real-time end-to-end processing from video input to behavior detection output.
2. A high-throughput end-to-end aphid honeydew discharge behavior identification method according to claim 1, characterized in that: In step S2, the RAMF algorithm designs an exponentially increasing weight distribution scheme, as shown in formula (1), to ensure that the weight of the latest frame in the time window is higher; Where ω(t) represents the time weighting coefficient, t represents the frame index in the time window, and Indicates the size of the time window.
3. A high-throughput end-to-end aphid honeydew discharge behavior identification method according to claim 1, characterized in that: In step S3, the RK50 module implements kernel adaptation through the B-spline basis function, and the activation function The expression of is shown in formula (2): Where b(x) is the basis function (usually the basis function is the Sigmoid linear unit), Bi(x) is the B-spline basis function, and ci is the trainable coefficient.
4. A high-throughput end-to-end aphid honeydew discharge behavior identification method according to claim 3, characterized in that: In step S3, the KAN convolution layer of the RK50 module decomposes the convolution into two parallel paths: one is the base path, which applies a conventional convolution operation to the input transformed by the base activation function; the other is the spline path, which applies the B-spline basis transformation to the input before convolution. The output of a single group of KAN convolution layers can be expressed as formula (3): Where g(⋅) represents the basic activation function, B(x) refers to the spline basis transformation, and Norm is the normalization function.
5. The high-throughput end-to-end aphid honeydew excretion behavior identification method according to claim 1, characterized in that: In step S4, a motion-original video cross-frame detection method is implemented, odd-numbered frames are processed with motion features to identify crawling and kicking behaviors, and even-numbered frames use original optical features to detect translucent honeydew targets. A delayed interpolation algorithm is used to extend the single-frame results by 1 frame forward and backward to eliminate flickering in the detection results and reduce computational overhead by 50%.
6. A high-throughput end-to-end aphid honeydew discharge behavior identification method according to claim 1, characterized in that: In step S4, a hierarchical three-stage analysis strategy is adopted to divide the determination of honeydew discharge behavior into three stages: The first stage uses low-frequency detection to identify high-frequency micro-movements before honeydew discharge; The second stage conducts honeydew identification, detecting physical evidence of metabolite production in the form of single droplets; In the third phase, a composite movement assessment was conducted to associate leg movements with honeydew droplets based on temporal continuity and merge them into honeydew excretion behavior.
7. The high-throughput end-to-end aphid honeydew discharge behavior identification method according to claim 1, characterized in that: The step S5 also includes building a comprehensive acceleration framework, using a multi-threaded parallel pipeline architecture to decouple the serial stages of traditional video processing into concurrent threads of frame acquisition, feature extraction and result rendering, and reducing runtime fragmentation through a GPU memory pre-allocation strategy. It uses persistent CUDA streams to achieve asynchronous parallel computing of motion feature extraction and target detection, uses batch optimization to reduce task startup overhead, combines the CUDA math library to vectorize and accelerate key operations, and uses just-in-time compilation technology and mixed precision computing to reduce the amount of computation while maintaining accuracy, ultimately building an end-to-end real-time behavior recognition system.