An artificial intelligence-based analysis method and system for whitefly feeding preference
Patent Information
- Application Number
- CN202610084271.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-01-22
AI Technical Summary
微小目标检测精度不足:烟粉虱体长小于1mm,传统目标检测算法如通用YOLO模型未针对微小目标进行适配,易受植物纹理、环境噪声干扰,导致检测漏检率高、误检率高;
[0013]本发明通过自适应双边滤波降噪、光流引导和植物结构约束补偿植物运动,获取稳定视频;再利用注意力增强YOLO提取烟粉虱三维轨迹,结合轨迹动力学与姿态分析识别取食事件;随后融合多视角植物表型特征、DTW对齐的挥发物特征及轨迹植物交互特征,最后构建强化学习框架,以DQN模型训练实现取食热点精准预测,全流程自动化融合多源数据,有效解决现有烟粉虱取食趋向分析中检测精度低、植物运动干扰难消除、多源数据融合不精准等痛点,通过多源特征深度融合与强化学习实现取食热点精准预测,可自动化完成从数据采集到结果输出的全流程分析,为农业烟粉虱精准防控提供科学靶向依据,助力减少农药滥用、降低防控成本,保障农作物产量与品质,推动农业害虫监测向智能化、精准化升级。
Smart Images

Figure CN121982709B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of insect feeding tendency analysis technology, specifically to an artificial intelligence-based method and system for analyzing the feeding tendencies of whiteflies. Background Technology
[0002] Whiteflies are a global agricultural pest. Both adults and nymphs damage crops by feeding on the sap of host plants and also transmit various plant viruses, severely impacting crop yield and quality. Accurate analysis of whitefly feeding trends, identifying their preferred host plant regions and related influencing factors, is crucial for developing precise control strategies and reducing pesticide overuse. Currently, technologies for analyzing whitefly feeding trends have the following limitations: Insufficient accuracy in detecting tiny targets: Whiteflies are less than 1 mm in length. Traditional target detection algorithms, such as the general YOLO model, are not adapted to tiny targets and are easily affected by plant textures and environmental noise, resulting in high false negative and high false positive rates. Poor handling of background interference: In natural monitoring scenarios, host plants may experience slight swaying, curling, and other non-rigid movements due to environmental factors. Traditional video stabilization technologies are mostly designed for rigid backgrounds and cannot effectively eliminate the interference of plant movement on the extraction of whitefly trajectories, which may easily lead to misjudging plant movement as whitefly behavior. Inaccurate fusion of multi-source data: The feeding tendency of whiteflies is influenced by multiple factors such as plant phenotype and volatiles, but existing technologies mostly analyze data in a single dimension, failing to achieve deep fusion of behavioral data, plant phenotype data, and volatile chemical data, and there is also the problem of asynchronous volatile collection and behavioral monitoring; Poor generalization ability of trend prediction models: Traditional prediction methods mostly use regression, general CNN and other models, without combining the characteristics of whitefly feeding behavior for customized design. They cannot accurately quantify the comprehensive impact of multiple factors on feeding trends, resulting in low accuracy of hotspot prediction.
[0003] Therefore, there is an urgent need for a technology capable of accurately detecting tiny targets, effectively eliminating plant dynamic interference, achieving deep fusion of multi-source data, and accurately predicting the feeding trends of whiteflies, in order to overcome the shortcomings of existing technologies. Summary of the Invention
[0004] To address the shortcomings of existing methods and the needs of practical applications, and in order to solve the aforementioned problems, this invention provides an artificial intelligence-based method for analyzing the feeding tendencies of whiteflies, comprising the following steps: The foreground of whiteflies and the plant background are separated from the monitoring video sequence to obtain the foreground video sequence and background model. Based on the foreground video sequence and background model, a stable video sequence with effective compensation is obtained through optical flow guidance and non-rigid plant motion compensation constrained by plant structure. Using an attention-enhanced YOLO network, the three-dimensional behavioral trajectory of the insect is extracted from the stable video sequence. The trajectory segments marked with feeding events are obtained through trajectory dynamics and posture sequence analysis. The association weights between the feeding event sequence and the dynamic response features of volatiles are mined through multi-head attention to obtain behavioral chemical decision features. Based on the three-dimensional behavioral trajectory and plant phenotypic features, the trajectory-plant interaction features are obtained through a spatiotemporal graph neural network with trajectory points as nodes. Combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory-plant interaction features, the multi-source perception fusion directional decision modeling of deep reinforcement learning is used to predict the distribution of feeding hotspots in whitefly colonies.
[0005] Optionally, separating the whitefly foreground and plant background from the monitoring video sequence to obtain the foreground video sequence and background model includes the following steps: Based on the monitoring video sequence, a clear and detailed denoised image frame sequence of whiteflies is obtained by bilateral filtering; dynamic background modeling and foreground separation are performed based on the denoised image frame sequence to obtain a foreground video sequence and a background model.
[0006] Optionally, the step of obtaining an effectively compensated stable video sequence based on the foreground video sequence and background model through optical flow guidance and non-rigid plant motion compensation constrained by plant structure includes the following steps: Based on the foreground video sequence and background model, optical flow is used to capture the displacement features of plant leaf curling and swaying. Based on the displacement features, the direction and amplitude are compensated using the prior constraints of plant geometry to eliminate the interference of plant movement on the trajectory of whiteflies, thereby obtaining a stable video sequence with effective compensation.
[0007] Optionally, the step of using an attention-enhanced YOLO network to extract the insect's three-dimensional behavioral trajectory from the stabilized video sequence includes the following steps: The attention-enhanced YOLO network is used to obtain the three-dimensional location information of whiteflies from the stabilized video sequence; based on the three-dimensional location information, the three-dimensional behavioral trajectory of the insect is extracted by trajectory tracking with spatiotemporal continuity constraints.
[0008] Optionally, the step of obtaining a feeding event sequence by analyzing the trajectory segments marked with feeding events through trajectory dynamics and attitude sequence analysis includes the following steps: The trajectory dynamics characteristics are analyzed based on the three-dimensional trajectory coordinate sequence to obtain candidate feeding trajectory points; the posture sequence is analyzed based on the candidate feeding trajectory points to obtain posture classification results; the trajectory dynamics characteristics and the posture classification results are weighted and fused to obtain a feeding event sequence.
[0009] Optionally, the step of mining the association weights between the feeding event sequence and the dynamic response features of volatiles through multi-head attention to obtain behavioral chemical decision features includes the following steps: A multi-head attention fusion model is constructed; the feeding event is used as the query end, and the volatile characteristics are used as the key and value ends. The association weights are mined through the multi-head attention fusion model to obtain behavioral chemical decision features.
[0010] Optionally, obtaining trajectory-plant interaction features based on the three-dimensional behavioral trajectory and plant phenotypic features through a spatiotemporal graph neural network with trajectory points as nodes includes the following steps: A spatiotemporal graph is constructed by using trajectory points as nodes in a graph and plant phenotypes as node attributes. The trajectory-plant interaction features of the spatiotemporal graph are then mined using a spatiotemporal graph convolutional network.
[0011] Optionally, the step of combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features, and using deep reinforcement learning to model multi-source perception fusion directional decision-making to predict the feeding hotspot distribution of whitefly populations includes the following steps: A deep reinforcement learning model is constructed. The plant phenotypic features, behavioral chemical decision features, and trajectory plant interaction features are used as the state inputs of the deep reinforcement learning model. The feeding hotspot prediction is used as the action objective, and the distribution of feeding hotspots of the whitefly population is predicted through a reward function.
[0012] Optionally, extracting the plant phenotypic features includes the following steps: Multi-view plant images are fused to obtain a full-view three-dimensional plant phenotypic image; superpixel segmentation is performed on the full-view three-dimensional plant phenotypic image, and phenotypic feature vectors are extracted from the segmentation results.
[0013] This invention acquires stable video by using adaptive bilateral filtering for noise reduction, optical flow guidance, and plant structure constraints to compensate for plant movement. Then, it utilizes attention-enhanced YOLO to extract the three-dimensional trajectory of the whitefly, combining trajectory dynamics and posture analysis to identify feeding events. Subsequently, it integrates multi-view plant phenotypic features, DTW-aligned volatile features, and trajectory-plant interaction features. Finally, it constructs a reinforcement learning framework and trains a DQN model to achieve accurate prediction of feeding hotspots. The entire process is automated, fusing multi-source data, effectively addressing the pain points of existing whitefly feeding trend analysis, such as low detection accuracy, difficulty in eliminating plant movement interference, and inaccurate multi-source data fusion. Through deep fusion of multi-source features and reinforcement learning, it achieves accurate prediction of feeding hotspots and can automatically complete the entire analysis process from data acquisition to result output. This provides a scientific basis for precise control of agricultural whiteflies, helps reduce pesticide overuse, lowers control costs, ensures crop yield and quality, and promotes the intelligent and precise upgrading of agricultural pest monitoring.
[0014] Secondly, to efficiently execute the AI-based whitefly feeding tendency analysis method provided by this invention, this invention also provides an AI-based whitefly feeding tendency analysis system, comprising: an input device, an output device, a processor, and a memory, wherein the input device, output device, processor, and memory are interconnected, and the memory includes program instructions for executing the AI-based whitefly feeding tendency analysis method described in the first aspect of this invention. The AI-based whitefly feeding tendency analysis system of this invention has a compact structure and stable performance, and can stably execute the AI-based whitefly feeding tendency analysis method provided by this invention, further enhancing the overall applicability and practical application capability of this invention. Attached Figure Description
[0015] Figure 1 A flowchart of an artificial intelligence-based method for analyzing the feeding tendencies of whiteflies, provided in an embodiment of the present invention; Figure 2 This is a framework diagram of an artificial intelligence-based whitefly feeding tendency analysis system provided for an embodiment of the present invention. Detailed Implementation
[0016] Specific embodiments of the present invention will now be described in detail. It should be noted that the embodiments described herein are for illustrative purposes only and are not intended to limit the invention. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, well-known circuits, software, or methods have not been specifically described to avoid obscuring the invention.
[0017] Throughout this specification, references to "an embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with that embodiment or example is included in at least one embodiment of the invention. Therefore, the phrases "in an embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily refer to the same embodiment or example. Furthermore, specific features, structures, or characteristics can be combined in one or more embodiments or examples in any suitable combination and / or sub-combination. Moreover, those skilled in the art will understand that the illustrations provided herein are for illustrative purposes and are not necessarily drawn to scale.
[0018] Please see Figure 1 To address the aforementioned problems, this invention provides an artificial intelligence-based method for analyzing the feeding tendencies of whiteflies, such as... Figure 1 As shown, in one embodiment, the method includes the following steps: S1. Separate the foreground of whiteflies and the background of plants from the monitoring video sequence to obtain the foreground video sequence and the background model. Based on the foreground video sequence and the background model, obtain a stable video sequence with effective compensation by optical flow guidance and non-rigid plant motion compensation constrained by plant structure.
[0019] In this embodiment, the step of separating the foreground of whiteflies and the plant background from the monitoring video sequence to obtain the foreground video sequence and the background model includes the following steps: S111. Based on the monitoring video sequence, a clear and detailed whitefly-denoised image frame sequence is obtained by bilateral filtering.
[0020] First, using video analysis tools, the original monitoring video was split into consecutive single-frame images in time-stamp order to ensure that the temporal continuity of the frame sequence was not disrupted. For overexposed frames commonly found in monitoring scenes, adaptive histogram equalization technology was used to adjust the pixel grayscale distribution and repair the problem of lost image details. For blurry frames caused by lens shake or inaccurate focus, Wiener filtering was used for sharpening. At the same time, a grayscale gradient threshold was set to filter out severely abnormal frames that could not be repaired, and the frame sequence was completed by interpolation of adjacent valid frames. Finally, taking into account the movement speed characteristics of whiteflies, the video frame rate was uniformly adjusted to 25fps, and the resolution of all frames was standardized to 1920×1080 pixels to avoid the impact of inconsistent formats on subsequent processing results.
[0021] Furthermore, the image is divided into 3×3 pixel local neighborhoods frame by frame. The size of this neighborhood is chosen based on the tiny scale characteristic of the whitefly, which has a body length of less than 1 mm, to ensure accurate perception of local noise in the image and avoid smoothing of tiny target details due to an excessively large neighborhood. The noise intensity is quantified by calculating the variance of the pixel gray values in each neighborhood (the larger the variance, the higher the degree of dispersion). The gray value variance in areas such as light reflection and dense sensor noise areas is usually greater than 50, and these areas are clearly identified as high noise areas. The bilateral filtering parameters are dynamically adjusted for different noise intensity areas: the spatial kernel radius in high noise areas is adjusted from the default 3 to 5 to enhance the filtering ability of dense noise while ensuring the overall smoothness of the image.
[0022] Furthermore, based on the findings of the whitefly behavior monitoring experiment, which identified high-incidence areas of whitefly activity (such as leaf edges and near leaf veins, where the frequency of whitefly occurrence is more than three times that of other areas), the spatial kernel radius for these areas was reduced to 2, and the gray-scale difference threshold was lowered from the default 25 to 15. By reducing the spatial kernel radius, the coverage of small targets was reduced, and by lowering the gray-scale difference threshold, the subtle gray-scale differences between whiteflies and the background were preserved, thereby maximizing the preservation of key small details such as the edge outline and foot tip of whiteflies.
[0023] Bilateral filtering is performed pixel-by-pixel on each frame of the image according to the adjusted parameters. Specifically, for each target pixel, all pixels within its 3×3 neighborhood are traversed. Spatial weights are calculated based on the spatial distance between neighboring pixels and the target pixel, and simultaneously, grayscale weights are calculated based on the grayscale difference between the two. The two weights are multiplied to obtain the final fusion weight, and the filtered grayscale value of the target pixel is calculated by weighted averaging. This operation, while efficiently filtering Gaussian noise and salt-and-pepper noise, accurately balances the grayscale similarity and spatial proximity of pixels, effectively preventing the subtle grayscale differences between the whitefly and the background from being excessively smoothed and weakened. After the wave is completed, the Canny edge detection algorithm is used to verify each frame of the image. The preset edge contour integrity standard is "the ratio of the number of detected whitefly edge pixels to the number of manually labeled real edge pixels ≥ 90%". If a frame does not meet this standard, it means that the filtering parameters have over-smoothed the target details. At this time, the current filtering parameters of the frame are backtracked and the spatial kernel radius is finely adjusted by ±1 or the gray difference threshold is finely adjusted by ±5. The filtering operation is re-executed on the frame, and the edge detection verification is performed again. The process is iterated until the edge integrity ratio is ≥ 90%, and finally a clear and complete denoised image frame sequence of whitefly details is obtained.
[0024] S112. Perform dynamic background modeling and foreground separation based on the denoised image frame sequence to obtain a foreground video sequence and a background model.
[0025] The first 30 frames of clean plant images without oil fume or whitefly activity were selected to initialize a Gaussian mixture background model. At the same time, the structural prior parameters such as the leaf edge contour coordinates and leaf vein distribution characteristics of the host plant obtained through previous 3D scanning were imported and incorporated into the model initialization process. This limited the grayscale fluctuation range and spatial distribution pattern of the background pixels, and avoided misjudging the texture features of the plant itself as noise or foreground objects.
[0026] Then, the dense optical flow algorithm is used to calculate the displacement vector of each pixel in the plant region of the background model. By cluster analysis, pseudo displacement vectors caused by sudden changes in light are eliminated to obtain the effective displacement field that represents non-rigid movements such as leaf curling and slight swaying. The intensity of plant movement is judged based on the amplitude of the displacement vector. When the intensity of movement is large, the update rate of the background model is reduced to avoid misjudging moving plant leaves as foreground whiteflies. When the intensity of movement is small, the normal update rate is maintained to ensure that the model can adapt to the slow changes in ambient light.
[0027] Each frame of the denoised image is differentially analyzed with the real-time updated background model to obtain a foreground difference image. An adaptive threshold segmentation algorithm is used to automatically determine the segmentation threshold based on the gray-level histogram of the difference image, thereby distinguishing the foreground of the whitefly from the residual background noise. Morphological opening operations are performed on the segmented foreground image, that is, firstly, the tiny background noise spots are removed by erosion, and then the complete morphology of the whitefly target is restored by dilation. Finally, the foreground regions of all frames are stitched together in chronological order to form a foreground video sequence, while the real-time updated background model is saved to provide basic data for subsequent non-rigid motion compensation steps for plants.
[0028] The process of obtaining a stabilized video sequence with effective compensation based on the foreground video sequence and background model, through optical flow guidance and non-rigid plant motion compensation constrained by plant structure, includes the following steps: S121. Based on the foreground video sequence and background model, capture the displacement features of plant leaf curling and swaying using optical flow.
[0029] First, the optical flow calculation range is limited by masking operations, and calculations are performed only on plant organ regions such as leaves and stems in the background model, excluding irrelevant backgrounds such as soil and containers to reduce redundant calculations. Then, a dense optical flow algorithm is used to process the plant region frame by frame. By analyzing the grayscale change trajectory of pixels between adjacent frames, a three-dimensional displacement vector of each pixel is obtained. This vector can accurately represent the movement trajectory of pixels in the horizontal, vertical and depth directions, thereby capturing subtle movements such as leaf curling and swaying. Furthermore, a clustering algorithm was used to classify and aggregate the acquired optical flow vectors. Combined with the optical flow characteristics of normal plant movement in the whitefly monitoring scenario (summarized from previous experiments), screening conditions were set to remove pseudo optical flow vectors caused by sudden changes in light such as light flickering and shadow movement. The remaining effective optical flow vectors were integrated to form a displacement field that can accurately characterize the non-rigid motion state of plants, providing an accurate motion benchmark for subsequent compensation operations.
[0030] S122. Based on the displacement characteristics, the direction and amplitude of the compensation are compensated by prior constraints of the plant geometry, eliminating the interference of plant movement on the trajectory of the whitefly, and obtaining a stable video sequence with effective compensation.
[0031] Based on prior geometric data of the host plant, such as leaf curvature and spatial coordinates of various organs, obtained through 3D scanning technology, plant motion constraints are constructed. The foreground video sequence is matched with the plant motion displacement field at the pixel level to accurately mark the pixel region corresponding to the whitefly foreground target and lock its coordinates, ensuring that the region is not affected during the compensation process. For the pixels in the plant region, the direction and amplitude of their motion are determined according to the displacement field, and a translation compensation operation with equal amplitude is performed in the opposite direction of the motion. The compensation amplitude strictly follows the constraint model to ensure that the morphology of the compensated plant organs conforms to the natural growth law and avoids image distortion caused by overcompensation.
[0032] The compensated image of each frame is compared with the background model by differential operation, and the average gray value of the differential image is calculated as the differential error index. A reasonable error threshold is determined through preliminary experiments. If the differential error of a frame exceeds the threshold, it means that the plant motion compensation in that frame is insufficient. It is necessary to backtrack to the optical flow extraction stage, fine-tune the parameters of the optical flow algorithm, recalculate the optical flow field, and then perform the compensation operation based on the new displacement field until the differential error meets the threshold requirement. Finally, all image frames that passed the compensation verification were collected and stitched together in the order of the original video timestamps. During the stitching process, inter-frame grayscale smoothing was used to avoid obvious brightness jumps between adjacent frames and ensure smooth video playback. The final output was a stable video sequence in which the plant background remained relatively still, and only the autonomous movement trajectory of the whitefly was preserved, providing high-quality data without background motion interference for insect trajectory extraction.
[0033] S2. Using an attention-enhanced YOLO network, extract the three-dimensional behavioral trajectory of the insect from the stabilized video sequence, and obtain the feeding event sequence by analyzing the trajectory segments marked with feeding events through trajectory dynamics and posture sequence analysis.
[0034] In this embodiment, the step of extracting the three-dimensional behavioral trajectory of an insect from the stabilized video sequence using an attention-enhanced YOLO network includes the following steps: S211. Using an attention-enhanced YOLO network, obtain the three-dimensional location information of the whitefly from the stabilized video sequence.
[0035] First, YOLOv8 is selected as the basic detection network, and a small target attention module (SOAM) is embedded after the feature fusion module of its neck. Specifically, the pixel gradient is calculated on the multi-scale feature map output from the neck. The region with drastic gray-level changes is selected by gradient difference (the gray-level change of the whitefly target area is usually higher than that of the plant background) to generate an attention mask. This mask is multiplied pixel by pixel with the original feature map to amplify the feature weight of the whitefly region, while suppressing the weight of redundant feature regions such as plant leaf texture and veins, thereby enhancing the recognizability of the target features.
[0036] Then, model adaptation training was carried out. Monitoring videos of whiteflies under different host plants and different light conditions were collected, and image frames containing whiteflies were extracted and labeled with pixel-level accuracy to ensure coverage of different whitefly postures (crawling, feeding, flying). Based on the actual body size of whiteflies, which is less than 1 mm in length, the default YOLO anchor boxes were adjusted to three tiny sizes: 1×1, 2×2, and 1×2 pixels. A transfer learning strategy was adopted, freezing the parameters of the first 10 layers of the YOLOv8 backbone network (utilizing the pre-trained general feature extraction capabilities), and only training the newly added attention module in the neck and the head detection layer. During the training process, the AdamW optimizer was used, with an initial learning rate of 0.001. The learning rate was reduced every 20 training epochs, and the training was iterated for 100 epochs until the whitefly detection accuracy on the validation set stabilized at over 92%.
[0037] The stabilized video sequence is input frame by frame into the trained attention-enhanced YOLO model to obtain the two-dimensional pixel coordinates and confidence scores of the whitefly in each frame. Low-confidence detection results with a confidence score below 0.85 are discarded, and accurate targets are retained. At the same time, a stereo camera is activated to synchronously acquire depth images. The stereo camera is pre-calibrated using a checkerboard calibration board to ensure the accuracy of depth measurement. Using the principle of triangulation, the two-dimensional pixel coordinates of the whitefly are combined with the depth information of the corresponding position to convert them into three-dimensional world coordinates (x, y, z), realizing the coordinate mapping from the image plane to the real space.
[0038] S212. Based on the three-dimensional position information, extract the three-dimensional behavioral trajectory of the insect through trajectory tracking constrained by spatiotemporal continuity.
[0039] First, the average movement speed of whiteflies in the stabilized video is statistically analyzed through preliminary experiments. The maximum possible movement distance of whiteflies between adjacent frames is calculated in combination with the video frame rate. Based on this, a three-dimensional Euclidean distance threshold is set. For each whitefly target detected in the current frame, the Euclidean distance is calculated with the three-dimensional coordinates of all whitefly targets in the previous frame. If the distance between two targets is less than the set threshold and the similarity of their appearance features (grayscale, outline size) is higher than 0.9, they are determined to be the same insect, and the inter-frame target association is completed. When a whitefly disappears due to being obscured by a leaf, a Kalman filter is activated to predict its location. A filtering model is constructed based on the whitefly's uniform motion characteristics. The three-dimensional coordinates and motion velocity of the last three frames before obscuration are input to predict the possible location of the insect in subsequent frames. A threshold for the number of frames in which the target disappears is set. If the target reappears within the threshold range, the Euclidean distance between the re-detected coordinates and the predicted coordinates is calculated. If the distance is less than the fault tolerance threshold, the target is associated with the original ID to maintain ID continuity. If the target is not detected after exceeding the threshold, the trajectory corresponding to the ID is marked as terminated. All 3D coordinates corresponding to the same ID are sorted according to the timestamp of the video frame to form an initial trajectory sequence. For a small number of missing coordinates in the sequence due to brief occlusion or detection omission, cubic spline interpolation is used to complete them to ensure the continuity of the trajectory. Finally, the complete trajectory coordinate sequence of each whitefly individual is output. Each node in the sequence contains a timestamp and the corresponding 3D spatial coordinates, providing basic data for subsequent feeding event identification.
[0040] Furthermore, the step of obtaining the feeding event sequence by analyzing the trajectory segments marked with feeding events through trajectory dynamics and attitude sequence analysis includes the following steps: S221. Analyze the trajectory dynamics characteristics based on the three-dimensional trajectory coordinate sequence to obtain candidate feeding trajectory points.
[0041] Based on the three-dimensional trajectory coordinate sequence of whiteflies, frame-by-frame dynamic features were extracted: the motion velocity corresponding to each frame was calculated based on the time interval and spatial distance between two adjacent trajectory points; acceleration was further calculated by the change in velocity and time interval between adjacent frames; angular velocity was obtained by combining the spatial direction change of the trajectory points; and the range of dynamic features under feeding conditions was statistically analyzed through preliminary experiments on whitefly feeding behavior to determine feeding judgment thresholds (e.g., velocity below 0.1 mm / s, acceleration below 0.01 mm / s). 2 The dynamic characteristics of the trajectory points are compared frame by frame, and the trajectory points that simultaneously meet two threshold conditions are selected as candidate trajectory points for feeding.
[0042] S222. Based on the feeding candidate trajectory points, perform posture sequence analysis to obtain posture classification results.
[0043] For the selected candidate trajectory points, local images of whiteflies from the corresponding frames are extracted from the stabilized video: a 32×32 pixel local image is cropped with the pixel coordinates of the image mapped back to the 3D coordinates of the candidate trajectory points as the center (ensuring that the whitefly body is completely contained); a lightweight convolutional neural network (CNN) is used to extract pose features. First, the outline of the whitefly body and the surface outline of the host plant are delineated by the edge detection algorithm, and the angle between the two (i.e., the body-plant angle) is calculated; then, the key point detection technology is used to identify the key positions of the whitefly's legs and determine whether the legs are extended or curled up. Finally, the feature information representing the whitefly's pose is integrated. Then, a whitefly pose classification dataset was constructed, collecting a large number of local images of whiteflies in feeding (body attached to plant, legs curled up) and non-feeding (crawling, flying, large angle between body and plant, legs extended) states. A sufficient number of samples were labeled (e.g., 12,000 feeding samples and 8,000 non-feeding samples), and divided into training and test sets in a 7:3 ratio. The aforementioned lightweight CNN was trained using this dataset, with the Adam optimizer used during training. The initial learning rate was set to 0.0005, and the learning rate was reduced every 15 training epochs. The training was iterated for 50 epochs until the pose classification accuracy of the model on the test set stabilized above 92%. The local images corresponding to the candidate trajectory points were input into the trained CNN to obtain the pose classification results and confidence scores for each image.
[0044] S223. The trajectory dynamics features and the posture classification results are weighted and fused to obtain the feeding event sequence.
[0045] A weighted fusion strategy is adopted to integrate dynamic features and posture classification results: the judgment result that the dynamic features meet the feeding threshold is recorded as the basic condition, and the judgment result that the posture classification result is "feeding" and the confidence level is higher than 0.85 is recorded as the core condition; when both conditions are met at the same time, the corresponding trajectory point is judged as a valid feeding point; continuous valid feeding points are combined into feeding trajectory segments, and isolated valid points with a length of less than 3 frames are removed (to avoid misjudgment caused by instantaneous stillness) to ensure the authenticity of feeding trajectory segments.
[0046] For each valid feeding trajectory segment, its corresponding temporal and spatial information is extracted: the start and end times of the feeding event are determined based on the start and end timestamps of the trajectory segment; the average three-dimensional coordinates of all valid points in the trajectory segment are used as the core spatial location of the feeding event; the duration of the trajectory segment (the time difference between the end and start frames) is also calculated; finally, the feeding event sequence for each whitefly is output according to its individual ID, and each event in the sequence includes the start and end times, the three-dimensional core location, and the duration, providing core behavioral event data for subsequent multi-source feature fusion.
[0047] S3. By mining the correlation weights between the feeding event sequence and the dynamic response features of volatiles through multi-head attention, behavioral chemical decision features are obtained.
[0048] Dynamic, asynchronous volatile data is converted into quantitative features to capture the correlation between volatiles and feeding events. Specifically, the dynamic response features of volatiles are extracted, including the following steps: First, the time-series data of volatile concentrations obtained by GC-MS detection were cleaned, and outliers were removed using the 3σ criterion. Then, the Z-score normalization method was used to eliminate the dimensional influence of different volatile component concentrations, making the concentration data of different orders of magnitude comparable. Finally, based on the stabilized video frame rate, the temporal resolution of the volatile time-series data was uniformly adjusted to 25fps using linear interpolation to ensure synchronization with the video timeline and lay the foundation for subsequent time alignment.
[0049] Furthermore, using the feeding event sequence as a reference, the start and end time nodes of each feeding event are extracted to construct a time alignment template. The processed volatile matter time series data and this template are input into a dynamic time warping algorithm. The algorithm automatically stretches or compresses the time axis of the volatile matter time series data through a dynamic programming strategy, calculates the cumulative distance between the data and the template, and finds the optimal alignment path. During the alignment process, it is crucial to ensure that the key nodes of volatile matter concentration changes are accurately matched with the start and end times of the feeding events. Finally, the alignment accuracy is verified through timestamp calibration to ensure that the time deviation between the aligned data and the feeding events is less than 1 frame, thus solving the problem of asynchronous volatile matter collection and feeding behavior monitoring.
[0050] Next, an LSTM autoencoder model was constructed: the model consists of two parts, an encoder and a decoder. The encoder is composed of two layers of LSTM network, with 64 neurons in each layer. It is responsible for progressively encoding the aligned high-dimensional volatile time series data into low-dimensional feature vectors, realizing data dimensionality reduction and key information extraction. The decoder is also composed of two layers of LSTM network, with a structure symmetrical to the encoder. It is responsible for reconstructing the low-dimensional feature vectors into volatile time series data with the same dimension as the original data, thereby verifying the effectiveness of the encoded features. During training, the Adam optimizer was selected, and the initial learning rate was set to 0.001 to ensure stable convergence of the model. Then, feeding event supervision signals were added during training: During model training, the labeled feeding event time markers were incorporated into the training as supervision signals, i.e., the time period during which a feeding event occurred was marked as 1, and the non-feeding time period was marked as 0; a hybrid loss function was constructed, which consisted of two parts: one part was the reconstruction loss between the reconstructed data output by the decoder and the original volatile data, ensuring that the encoded features could accurately represent the dynamics of volatiles; the other part was the cross-entropy loss based on the feeding event markers, guiding the model to learn the correlation between the dynamic changes of volatiles and feeding events; the weight ratio of the reconstruction loss to the cross-entropy loss was set to 7:3, and training was iterated for 80 rounds until the reconstruction error of the model on the validation set was stably lower than 0.05, at which point training was stopped; Finally, the dynamic response features of volatiles are extracted: After training, the network parameters of the encoder are fixed, and all volatile time-series data aligned by DTW are input into the encoder; the output of the last layer of the encoder's LSTM network is extracted as a low-dimensional feature vector that can simultaneously characterize the dynamic change law of volatiles and the response characteristics of feeding events; finally, this feature vector is associated with the corresponding timestamp to ensure that feeding events can be accurately matched in the future, providing standardized chemical feature data for cross-modal fusion.
[0051] Furthermore, the step of mining the association weights between the feeding event sequence and the dynamic response features of volatiles through multi-head attention to obtain behavioral chemical decision features includes the following steps: S31. Construct a multi-head attention fusion model.
[0052] Eight parallel attention heads were constructed and divided into three functional modules according to the dimensions of attention: three attention heads focused on temporal correlations, focusing on the temporal synchronicity between the temporal changes of feeding events and the dynamic changes of volatile concentrations; three attention heads focused on spatial correlations, analyzing the correspondence between the spatial location of feeding events and the intensity of volatile release in different regions of the plant; and two attention heads focused on concentration correlations, exploring the causal relationship between peak / trough values of volatile concentrations and the occurrence and termination of feeding events. Each attention head adopted a scaled dot product attention mechanism to ensure accurate capture of correlation features in different dimensions.
[0053] Within each attention head, the dot product similarity between the query vector and the key vector is first calculated to measure the degree of association between the feeding event and the volatile features. To avoid the gradient vanishing problem caused by excessively large similarity values, the dot product result is scaled by dividing it by the square root of the key vector dimension. Then, the scaled similarity is normalized using the Softmax function to obtain the attention weights for each dimension. The higher the weight value, the greater the influence of the volatile features in the corresponding dimension on the feeding event.
[0054] The attention weights obtained from each attention head are weighted and summed with the corresponding Value vectors to obtain eight intermediate feature vectors that represent the correlation features in different dimensions. Then, these eight intermediate feature vectors are concatenated according to the channel dimension to form a high-dimensional preliminary fusion feature. This feature fully preserves the correlation information between feeding events and volatiles in the three dimensions of time, space, and concentration.
[0055] S32. Using feeding events as the query end and volatile characteristics as the key and value ends, the association weights are mined through the multi-head attention fusion model to obtain behavioral chemical decision features.
[0056] First, the core information of the feeding event is standardized by converting the start and end times and duration of the feeding event into relative timestamps and normalizing them to the 0-1 range. The three-dimensional core location coordinates are also normalized according to the overall spatial range of the plant. Then, these normalized time, location, and duration information are encoded into a fixed-dimensional query vector through two fully connected layers to ensure that the vector can fully represent the key attributes of the feeding event. At the same time, the dimensionality adjustment layer is used to upgrade or reduce the dimensionality of the volatile dynamic response feature vector, converting it into a key vector and value vector that perfectly match the dimension of the query vector, thus ensuring the dimensionality consistency of subsequent attention calculations.
[0057] The pre-fused features after splicing are input into two fully connected layers for dimensionality reduction and feature optimization. The first fully connected layer reduces the high-dimensional features to 128 dimensions and uses the ReLU activation function to enhance the non-linear expressive power of the features while suppressing redundant features. The second fully connected layer further reduces the 128-dimensional features to 64 dimensions and uses the Tanh activation function to map the feature values to the range of -1 to 1, improving the stability of the features. Finally, a standardized 64-dimensional behavioral chemistry decision feature vector is output. This vector can accurately quantify the influence weight of volatiles on the feeding behavior of whiteflies, providing core cross-modal data for multi-source feature fusion in subsequent steps.
[0058] S4. Based on the three-dimensional behavioral trajectory and plant phenotypic features, the trajectory plant interaction features are obtained through a spatiotemporal graph neural network with trajectory points as nodes.
[0059] Extracting the phenotypic features of the plant includes the following steps: S411. Fuse multi-view plant images to obtain a full-view three-dimensional plant phenotypic image.
[0060] First, the simultaneously acquired multi-view plant images (top, side, and oblique views) were preprocessed: image channels were unified by grayscale conversion, and Gaussian filtering was used to remove minor noise generated during acquisition. Then, a checkerboard calibration board was used to correct distortion in each view image to ensure geometric accuracy. Subsequently, the SIFT feature point extraction algorithm was used to extract scale-invariant feature points and corresponding descriptors from each view image. The FLANN matcher was used to perform preliminary matching of feature points between different views, and a distance threshold was set to remove incorrect matching pairs, retaining high-quality matching feature points. Based on the trajectory coordinates of the whitefly, the high-frequency feeding areas with densely distributed trajectory points were statistically analyzed, and the feature points corresponding to these areas were assigned a weight of 1.5 times, while the feature points in non-high-frequency areas retained their original weights. Finally, a weighted average fusion strategy was used to fuse the images from each view according to the feature point weights into a full-view three-dimensional plant phenotypic image. After fusion, the image consistency was verified by reprojection error to ensure that the fused image had no obvious misalignment and completely preserved the spatial structural information of each plant organ.
[0061] S412. Perform superpixel segmentation on the full-view three-dimensional plant phenotypic image and extract phenotypic feature vectors from the segmentation results.
[0062] The SLIC superpixel segmentation algorithm was used to process the fused full-view image. Considering the small-scale characteristics of the whitefly feeding area, the algorithm parameters were adapted as follows: the superpixel block size was set to 10×10 pixels to ensure accurate coverage of the leaf micro-area; the compactness parameter was set to 10 to balance the shape regularity of the superpixel blocks with the image edge fit; after segmentation, background blocks were filtered by color features (HSV color space) and spatial location information: areas with colors close to soil and containers, as well as areas outside the plant outline, were identified as background blocks and removed; only superpixel blocks whose colors conform to the characteristics of plant leaves and stems and are located within the plant outline were retained, and each retained plant organ superpixel block was marked with a unique identifier to facilitate subsequent feature association.
[0063] Then, multi-dimensional phenotypic feature extraction was performed on each plant organ superpixel block: In terms of morphological features, the boundary contour of the superpixel block was obtained through contour extraction algorithm, and the area of the region enclosed by the contour, the perimeter of the contour, and the ratio of perimeter to area (shape factor) were calculated to characterize the size and regularity of the block; in terms of texture features, the superpixel block was converted into a grayscale image, and a grayscale co-occurrence matrix was constructed (selecting a distance of 1 pixel and four directions of 0° / 45° / 90° / 135°). The contrast, entropy, and energy were extracted from the matrix to reflect the roughness of the leaf surface and the texture distribution pattern; in terms of physiological features, the stomatal region was segmented from the superpixel block using a threshold segmentation algorithm and the number was counted to obtain the stomatal density; the chlorophyll content was indirectly characterized by calculating the vegetation index (using the RGB channel brightness value of the superpixel block); Furthermore, the extracted morphological, texture, and physiological features are standardized: the min-max normalization method is used to map all feature values to the 0-1 range, eliminating dimensional differences between different features; then, principal component analysis is used to reduce the dimensionality of the standardized features, retaining principal components that reflect more than 95% of the original information and reducing feature redundancy; the dimensionality-reduced features are then concatenated in the order of morphology, texture, and physiology to generate a fixed-dimensional plant phenotypic feature vector; finally, the feature vector of each superpixel block is associated with its corresponding spatial coordinates (based on the three-dimensional coordinate system of the fused image) to ensure accurate matching of plant phenotypic information corresponding to the whitefly trajectory points.
[0064] The method of obtaining trajectory-based plant interaction features based on the three-dimensional behavioral trajectory and plant phenotypic features through a spatiotemporal graph neural network with trajectory points as nodes includes the following steps: S421. Construct a spatiotemporal graph using trajectory points as nodes of the graph and plant phenotypes as node attributes.
[0065] First, node construction and attribute matching are performed. Each trajectory point in the three-dimensional trajectory coordinate sequence of the whitefly is used as a node in the graph, and a unique identifier is assigned to each node. The three-dimensional spatial coordinates of the trajectory points are accurately matched with the spatial coordinates of the plant superpixel blocks to find the plant superpixel block corresponding to each trajectory point. The plant phenotypic feature vector of the superpixel block is used as the attribute of the corresponding trajectory point node. If a single trajectory point covers multiple superpixel blocks, the phenotypic features of these superpixel blocks are fused using a distance-weighted average method to generate the comprehensive attribute features of the node, ensuring that the node attributes can truly reflect the phenotypic state of the plant area where the trajectory point is located.
[0066] Then, spatiotemporal edges are constructed and weights are set. Two types of edges are constructed to realize spatiotemporal correlation modeling: temporal edges: only connect trajectory point nodes of adjacent timestamps of the same whitefly individual. The weight of the edge is set to the reciprocal of the time interval between two adjacent frames. The smaller the time interval (the smaller the difference between frames), the higher the weight, thereby strengthening the temporal continuity of the trajectory; spatial edges: connect each trajectory point node to the five nearest surrounding plant superpixel block nodes (the superpixel block node attribute is its own phenotypic feature vector). The weight of the edge is set to the reciprocal of the spatial distance between the trajectory point and the center of the superpixel block. The closer the distance (the more direct the phenotypic influence), the higher the weight; the weights of all edges are normalized to the 0-1 range to avoid excessive weight differences affecting model training.
[0067] S422. The trajectory plant interaction features of the spatiotemporal graph are mined through a spatiotemporal graph convolutional network.
[0068] An ST-GCN model with two spatiotemporal convolutional blocks is constructed, and feature extraction and optimization are performed block by block. The first spatiotemporal convolutional block uses a 1×3 convolutional kernel to perform sliding convolution on the temporal node attributes of the same trajectory, capturing the behavioral trajectory change features of whiteflies at different time points. The spatial convolutional layer is based on the constructed graph adjacency matrix (including spatiotemporal edge weights), and uses a graph convolutional kernel to weight and aggregate the attributes of each node and its neighboring nodes, capturing the spatial interaction features between trajectory points and the phenotypes of surrounding plants. After convolution, the feature distribution is standardized by a BatchNorm layer, and then the nonlinear expressive power of the features is enhanced by the ReLU activation function. The second spatiotemporal convolutional block repeats the above operations to further deepen the fusion features of temporal changes and spatial interactions, and filter redundant information. Next, a global average pooling operation is performed on the feature map output by the last layer of ST-GCN. The pooling window covers the entire temporal length of the trajectory, converting the variable-length trajectory temporal features into a fixed-length feature vector. The pooled feature vector is then subjected to L2 normalization to eliminate the influence of dimensions and improve feature stability. Finally, a 128-dimensional trajectory plant interaction feature vector is output. This vector accurately quantifies the spatiotemporal influence of plant phenotypes on the trajectory of whiteflies, providing core interactive feature data for multi-source feature fusion.
[0069] S5. Combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features, use deep reinforcement learning to model multi-source perception fusion directional decision-making and predict the distribution of feeding hotspots in whitefly populations.
[0070] The method of combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features, and using deep reinforcement learning to perform multi-source perception fusion directional decision modeling to predict the feeding hotspot distribution of whitefly populations includes the following steps: S51. Construct a deep reinforcement learning model.
[0071] Specifically, the state space is first constructed by standardizing and preprocessing the three types of input features: plant phenotypic feature vector, behavioral chemistry decision feature vector, and trajectory-plant interaction feature vector. The values are mapped to the 0-1 interval using min-max normalization to eliminate the dimensional differences between different feature sources. Then, the plant phenotypic feature vector is adjusted to a dimension compatible with the other two types of features through a dimension-adjusting fully connected layer. The adjusted three types of features are then concatenated in the order of "plant phenotypic - behavioral chemistry - trajectory-plant interaction" and further fused through a fully connected layer to finally generate a unified state vector, which serves as the input to the reinforcement learning agent.
[0072] Furthermore, an action space is defined, with segmented and preserved plant superpixel blocks as basic action units. Each action corresponds to a binary decision of "determining the superpixel block as a feeding hotspot" or "determining it as a non-feeding hotspot". The dimension of the action space is consistent with the total number of plant superpixel blocks. The reinforcement learning agent can predict feeding hotspots in the entire plant area by outputting the decision probability of each action unit.
[0073] Finally, a reward function is designed, employing a weighted sparse reward mechanism to guide model optimization. When the agent predicts a superpixel block as a feeding hotspot, and this block overlaps with the labeled actual feeding hotspot, a positive reward of +10 is given; when it is predicted as a hotspot but is not actually a hotspot, a negative reward of -5 is given to avoid misjudgment. An additional global matching reward is set, calculated by the cosine similarity between the density distribution of the predicted hotspot and the density distribution of the actual feeding hotspot. If the similarity is higher than 0.8, it indicates a high overall distribution matching degree, and an additional global positive reward of +20 is given. All reward values are discounted to the current training step using a discount factor of 0.95 to balance immediate rewards and long-term decision benefits.
[0074] S52. Using the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features as the state input of the deep reinforcement learning model, and taking the prediction of feeding hotspots as the action target, the distribution of feeding hotspots of the whitefly population is predicted through a reward function.
[0075] In this embodiment, a deep Q-network model is constructed, with a network structure containing three fully connected layers. The state vector is mapped to 128 dimensions and 64 dimensions sequentially, and the final output is a Q-value consistent with the action space dimension (i.e., the probability that each superpixel block is a feeding hotspot). An experience replay buffer with a capacity of 10,000 is set up. In each training iteration, 32 experience samples are randomly sampled from the buffer (the experience sample is a quadruple of "state-action-reward-next state") to avoid the correlation of training data affecting convergence. A dual-network structure of target network and evaluation network is introduced. The target network updates its parameters every 100 training steps to keep it synchronized with the evaluation network. An ε-greedy exploration strategy is adopted to balance exploration and exploitation. The initial value of ε is set to 0.9 (high exploration probability), which decays linearly to 0.1 (high exploitation probability) with the number of training steps, for a total decay of 10,000 steps. An Adam optimizer with a learning rate of 0.0001 is selected, and iterative training is performed for 20,000 steps until the reward value of the model on the validation set is stable and converges without significant fluctuations, at which point training stops.
[0076] The new multi-source feature data is processed into state vectors according to the state space construction process and input into the trained DQN model. The model outputs the Q value of each plant superpixel block. The superpixel blocks are sorted from high to low according to the Q value. The number of feeding hotspots is determined according to the size of the whitefly population (e.g., selecting the top 10%-15% of superpixel blocks). A heatmap visualization tool is used to generate a heatmap of the feeding hotspot distribution, with the color intensity representing the probability of the hotspot (the darker the color, the higher the probability of the area being fed by whiteflies). At the same time, the specific spatial coordinates of the hotspot area, the superpixel block identifier, and the corresponding Q value are output, completing the automated analysis and result presentation of the whitefly feeding tendency.
[0077] This invention first transforms the "prediction of feeding hotspots of whiteflies" into a standardized reinforcement learning problem (defining state, action, and reward) using a deep reinforcement learning model. Then, the DQN model learns the optimal decision-making strategy for this problem through a deep learning mechanism. This facilitates the integration of multi-source features to achieve accurate prediction of feeding hotspots and completes the automated analysis of whitefly feeding trends.
[0078] Please see Figure 2 In an embodiment, to efficiently execute the AI-based whitefly feeding tendency analysis method provided by this invention, this invention also provides an AI-based whitefly feeding tendency analysis system, comprising: an input device, an output device, a processor, and a memory, wherein the input device, output device, processor, and memory are interconnected, and the memory contains program instructions for executing the steps of the AI-based whitefly feeding tendency analysis method. The AI-based whitefly feeding tendency analysis system of this invention has a compact structure and stable performance, and can stably execute the AI-based whitefly feeding tendency analysis method of this invention, further enhancing the overall applicability and practical application capability of this invention.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the present invention.
Claims
1. A method for analyzing the feeding tendencies of whiteflies based on artificial intelligence, characterized in that, Includes the following steps: The foreground of whiteflies and the background of plants are separated from the monitoring video sequence to obtain the foreground video sequence and the background model. Based on the foreground video sequence and the background model, a stable video sequence with effective compensation is obtained through optical flow guidance and non-rigid plant motion compensation constrained by plant structure. Using an attention-enhanced YOLO network, the three-dimensional behavioral trajectory of the insect is extracted from the stabilized video sequence. The trajectory segments marked with feeding events are analyzed by trajectory dynamics and posture sequence to obtain the feeding event sequence. By mining the correlation weights between the feeding event sequence and the dynamic response features of volatiles through multi-head attention, behavioral chemical decision features are obtained. Based on the three-dimensional behavioral trajectory and plant phenotypic features, trajectory plant interaction features are obtained through a spatiotemporal graph neural network with trajectory points as nodes. Combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features, deep reinforcement learning is used to model multi-source perception fusion directional decision-making to predict the distribution of feeding hotspots in whitefly populations. The process of obtaining a stabilized video sequence with effective compensation based on the foreground video sequence and background model, through optical flow guidance and non-rigid plant motion compensation constrained by plant structure, includes the following steps: Based on the foreground video sequence and background model, optical flow is used to capture the displacement features of plant leaf curling and swaying. Based on the displacement characteristics, the direction and amplitude of compensation are compensated by prior constraints of plant geometry, eliminating the interference of plant movement on the trajectory of whiteflies, and obtaining a stable video sequence with effective compensation. The method of obtaining behavioral chemical decision features by mining the association weights between the feeding event sequence and the dynamic response features of volatiles through multi-head attention includes the following steps: Construct a multi-head attention fusion model; Using feeding events as the query endpoint and volatile characteristics as the key and value endpoints, the association weights are mined through the multi-head attention fusion model to obtain behavioral chemical decision features; The method of obtaining trajectory-based plant interaction features based on the three-dimensional behavioral trajectory and plant phenotypic features through a spatiotemporal graph neural network with trajectory points as nodes includes the following steps: A spatiotemporal graph is constructed by using trajectory points as nodes in the graph and plant phenotypes as node attributes. Trajectory plant interaction features of the spatiotemporal graph are mined using a spatiotemporal graph convolutional network.
2. The method for analyzing the feeding tendencies of whiteflies based on artificial intelligence according to claim 1, characterized in that, The process of separating the foreground of whiteflies from the plant background from the monitoring video sequence to obtain the foreground video sequence and the background model includes the following steps: Based on the monitoring video sequence, a clear and denoised image frame sequence with complete details of whiteflies was obtained by bilateral filtering; Dynamic background modeling and foreground separation are performed based on the denoised image frame sequence to obtain a foreground video sequence and a background model.
3. The method for analyzing the feeding tendencies of whiteflies based on artificial intelligence according to claim 1, characterized in that, The method of extracting the three-dimensional behavioral trajectory of insects from the stabilized video sequence using the attention-enhanced YOLO network includes the following steps: The attention-enhanced YOLO network was used to obtain the three-dimensional location information of whiteflies from the stabilized video sequence; Based on the aforementioned three-dimensional position information, the three-dimensional behavioral trajectory of the insect is extracted through trajectory tracking constrained by spatiotemporal continuity.
4. The method for analyzing the feeding tendencies of whiteflies based on artificial intelligence according to claim 1, characterized in that, The process of obtaining a feeding event sequence by analyzing trajectory segments marked with feeding events using trajectory dynamics and attitude sequence analysis includes the following steps: Based on the analysis of the trajectory dynamics characteristics of the three-dimensional trajectory coordinate sequence, candidate feeding trajectory points are obtained; Based on the feeding candidate trajectory points, posture sequence analysis is performed to obtain posture classification results; Combining the trajectory dynamics features and the posture classification results, a feeding event sequence is obtained, including: The judgment result that the dynamic characteristics meet the feeding threshold is recorded as the basic condition, and the judgment result that the posture classification result is feeding and the confidence level is higher than 0.85 is recorded as the core condition. When both conditions are met at the same time, the corresponding trajectory point is determined as a valid feeding point, and consecutive valid feeding points are combined into a feeding trajectory segment. For each valid feeding trajectory segment, the corresponding time and space information is extracted, classified by the individual whitefly ID, and the feeding event sequence of each whitefly is output. Each event in the sequence includes the start and end time, three-dimensional core location, and duration.
5. The method for analyzing the feeding tendencies of whiteflies based on artificial intelligence according to claim 1, characterized in that, The method of combining the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features, and using deep reinforcement learning to perform multi-source perception fusion directional decision modeling to predict the feeding hotspot distribution of whitefly populations includes the following steps: Build deep reinforcement learning models; Using the plant phenotypic features, the behavioral chemical decision features, and the trajectory plant interaction features as the state input of the deep reinforcement learning model, and taking the prediction of feeding hotspots as the action objective, the distribution of feeding hotspots of whitefly populations is predicted through a reward function.
6. The method for analyzing the feeding tendencies of whiteflies based on artificial intelligence according to claim 1, characterized in that, Extracting the phenotypic features of the plant includes the following steps: By fusing plant images from multiple perspectives, a full-view three-dimensional plant phenotypic image can be obtained; Superpixel segmentation is performed on the full-view 3D plant phenotypic image, and phenotypic feature vectors are extracted from the segmentation results.
7. A whitefly feeding tendency analysis system based on artificial intelligence, characterized in that, The artificial intelligence-based whitefly feeding tendency analysis system includes: an input device, an output device, a processor, and a memory, wherein the input device, output device, processor, and memory are interconnected, and the memory includes program instructions for executing the artificial intelligence-based whitefly feeding tendency analysis method according to any one of claims 1-6.
Citation Information
Patent Citations
Cucurbit vegetable virus disease identification method and system based on sensor array construction
CN120354283A
Trajectory prediction method and apparatus, and computer device and storage medium
WO2022222095A1