A deep reinforcement learning-based method for predicting the behavior of avian influenza hosts.

By employing deep reinforcement learning methods, combined with an improved YOLOv5 model and an Actor-Critic strategy, high-precision detection and dynamic modeling of the behavior of avian influenza host birds were achieved. This addresses the problem of insufficient model adaptability in existing technologies and improves the accuracy and stability of avian influenza monitoring.

CN120726693BActive Publication Date: 2026-01-06INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510822606.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2026-01-06
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing technologies lack the ability to jointly model local individual behavior and overall population situation in avian influenza surveillance, and lack an effective online adaptive mechanism for models, resulting in performance degradation and insufficient adaptability of prediction models during long-term deployment.

Method used

We employ a deep reinforcement learning-based approach, using multi-angle variable frame rate cameras to collect video data and combining it with an improved YOLOv5 model for target detection. We construct a dynamic graph neural network and an Actor-Critic policy module to achieve high-precision detection and dynamic modeling of individual and group bird behaviors, and improve the model's adaptability through an online fine-tuning mechanism.

Benefits of technology

It improves the accuracy and stability of bird behavior recognition, enhances the early identification capability of potential avian influenza transmission, breaks through the bottleneck of complex temporal behavior modeling, and realizes continuous performance maintenance of the model in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726693B_ABST
    Figure CN120726693B_ABST
Patent Text Reader

Abstract

The application discloses an avian influenza host bird behavior prediction method based on deep reinforcement learning, comprising the following steps: S1, deploying multi-angle and variable frame rate cameras in the target habitat to collect multi-source video data; S2, performing frame insertion, illumination normalization and background denoising processing on the video data; S3, inputting the pretreated data into an improved YOLOv5 model to obtain target detection results; S4, constructing a dynamic graph neural network based on bird nodes and space-time edges to output a flocking tendency vector; S5, inputting the flight trajectory into a low-level Actor-Critic policy network to generate a local short-term behavior decision; S6, combining the flocking tendency vector and historical error feedback to generate a global multi-step behavior prediction sequence; and S7, comparing the prediction results with actual observation data and triggering online fine-tuning according to the deviation. The application can realize non-invasive and high-precision bird behavior prediction and is suitable for avian influenza monitoring and early warning scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of animal behavior analysis technology, and in particular to a method for predicting the behavior of avian influenza host birds based on deep reinforcement learning. Background Technology

[0002] In the field of avian influenza surveillance and intelligent biological behavior recognition, with the development of artificial intelligence, image processing, and wireless sensing technologies, the recognition and prediction of wildlife behavior based on visual analysis has gradually become a research hotspot. In particular, accurate modeling and prediction of the behavioral patterns of migratory birds and other avian influenza host populations are of great significance for the early detection of disease transmission risks and the optimization of prevention and control strategies. Currently, mainstream bird behavior monitoring technologies can be broadly divided into two categories: individual tracking methods based on sensor tags and group behavior recognition methods based on video images.

[0003] Sensor-tag-based methods achieve long-term tracking of flight trajectories and physiological states by attaching GPS or RFID devices to individual birds. However, these methods suffer from high operational intrusion, high deployment costs, and significant behavioral interference, making large-scale deployment in wild natural environments difficult. In contrast, video image-based behavior recognition methods have greater potential for practical application due to their non-invasiveness and reproducibility. Currently, most research uses target detection algorithms to identify individual birds and combines them with traditional trajectory fitting or cluster analysis methods to classify their behavioral states. However, due to interference factors such as background occlusion, dense distribution of small targets, and fluctuations in ambient lighting in complex natural scenes, existing detection algorithms still have significant shortcomings in terms of accuracy and robustness.

[0004] Furthermore, regarding the temporal prediction of bird behavior, some studies have attempted to introduce graph neural networks and recurrent neural networks to model behavioral sequences. However, these methods generally lack the ability to model group interaction dynamics and struggle to adaptively optimize the prediction model based on environmental changes. As for reinforcement learning methods, some work has preliminarily verified their feasibility in behavior planning, but most methods still focus on optimizing individual behavioral strategies in static environments, and a comprehensive end-to-end collaborative mechanism suitable for dynamic group decision-making and prediction in natural scenarios has not yet been established.

[0005] Existing methods generally lack the ability to jointly model local individual behavior with the overall group situation, and lack an effective online model adaptation mechanism. They cannot make full use of continuously collected observation data for real-time fine-tuning, resulting in performance degradation and insufficient adaptability of prediction models during long-term deployment.

[0006] Therefore, how to provide a method for predicting the behavior of avian influenza hosts based on deep reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] One objective of this invention is to propose a method for predicting the behavior of avian influenza host birds based on deep reinforcement learning. This invention has the capabilities of non-invasive data collection, high-precision detection, dynamic population modeling, and online adaptive prediction, effectively improving the accuracy and stability of bird behavior recognition and multi-walk behavior prediction.

[0008] The method for predicting the behavior of avian influenza host birds based on deep reinforcement learning according to embodiments of the present invention includes the following steps:

[0009] S1. Deploy multi-angle, variable frame rate cameras in the target habitat to collect raw video streams and obtain multi-source video data;

[0010] S2. Perform time-series frame interpolation, ambient lighting normalization, and background modeling and denoising on the multi-source video data to obtain preprocessed video data.

[0011] S3. Input the preprocessed video data into the improved YOLOv5 model to obtain the object detection results;

[0012] S4. Based on the target detection results, construct a dynamic graph neural network with the detected bird individuals as nodes, spatiotemporal location information and relative velocity as edges, and output a clustering situation vector.

[0013] S5. Input the flight trajectory of individual birds in the target detection results into the low-level Actor-Critic strategy module to generate local short-term behavior decision results;

[0014] S6. Feed the clustering situation vector and historical prediction error back into the high-level Actor-Critic strategy module to generate a global multi-step behavior prediction sequence.

[0015] S7. Compare the local short-term behavior decision results and the global multi-step behavior prediction sequence with the real-time collected actual bird behavior observation data, calculate the decision bias and prediction bias, and trigger online fine-tuning based on the decision bias and prediction bias to update the YOLOv5 detection model and Actor-Critic strategy network parameters.

[0016] Optionally, S1 specifically includes:

[0017] S11. Based on DGNSS and ground reflection markers for assisted positioning, several three-dimensional spatial control points are set up in the target habitat, and environmental point clouds and orthophotos are obtained by fusion of lidar and aerial UAV photogrammetry. The target habitat includes foraging grounds, breeding grounds and migratory stopover sites of avian influenza host birds.

[0018] S12. Using environmental point cloud and orthophoto, the particle swarm optimization algorithm is used to calculate the three-dimensional polyhedron of each camera's field of view and solve for the optimal installation position and pitch angle. The pitch angle satisfies the requirement that the field of view coverage is greater than 98% and the occlusion probability is less than 2%.

[0019] S13. Fix the variable frame rate camera at the optimal installation position, and perform internal and external parameter calibration by combining striped structured light with the AprilTag calibration board to obtain the focal length, principal point coordinates, and radial and tangential distortion coefficients of the variable frame rate camera.

[0020] S14. Equip the variable frame rate camera with an electronically controlled liquid lens. The electronically controlled liquid lens achieves focal length switching from 100 mm to 200 mm by adjusting the applied voltage, and dynamically determines the region of interest based on image block histogram equalization analysis and buffers the video frames of the corresponding region at a frequency of 30 fps.

[0021] S15. Clock synchronization is performed using the PTPv2 protocol combined with the GNSS link, and a drift correction algorithm is embedded in the clock synchronization process based on the wireless mesh network to correct the clock deviation in real time.

[0022] S16. Each camera is equipped with an MPPT-controlled solar cell module and a lithium-ion battery storage module. The video resolution and frame rate are automatically adjusted by monitoring the voltage and output power of the battery module and according to the preset power-performance mapping table.

[0023] S17. A circular buffer structure is used to store the original video streams of each camera, and the video streams are fused based on the timestamp and camera viewpoint identifier to generate multi-source video data.

[0024] Optionally, S2 specifically includes:

[0025] S21. Accumulate the pixel-level absolute difference of each pair of adjacent frames in the multi-source video data and compare it with the stillness discrimination threshold. Extract the frames corresponding to the frames whose total pixel difference is less than or equal to the stillness discrimination threshold as key frames, and classify the other frames corresponding to the frames as dynamic frames.

[0026] S22. An integrated visible light and near-infrared spectral band synchronous acquisition device is used to record environmental spectral distribution data. Based on the spectral distribution data and a pre-calibrated pixel-spectral response lookup table, the pixel brightness in key frames and dynamic frames is linearly mapped and normalized to generate a spectral correction frame set.

[0027] S23. Divide the spectral correction frame set into fixed time windows. For each time window, swap the positions of the first frame and the Nth frame in the storage sequence to balance the difference in time interval between frames, and output the reconstructed time series.

[0028] S24. Construct a dual-layer background buffer canvas, wherein the first background canvas is updated every N1 frames for long-term static scene modeling, and the second background canvas is updated every N2 frames for short-term dynamic scene capture. The two background canvases are merged by a pixel-level overlay strategy and then merged with the reconstructed time series frames frame by frame to obtain the foreground frame sequence.

[0029] S25. Add a chain of metadata tags containing UTC timestamp, camera number, ambient light intensity and ambient temperature to each frame in the foreground frame sequence.

[0030] S26. Based on the camera number and frame sequence number, the foreground frame sequence is cross-merged in the storage sequence according to the timestamp and camera view label to generate preprocessed video data.

[0031] Optionally, S3 specifically includes:

[0032] S31. Input the preprocessed video data frame by frame into the YOLOv5s backbone network CSPDarknet, and process it layer by layer through cross-stage local connectivity structure and residual block stacking convolution to obtain the initial feature map:

[0033]

[0034] in, This represents the j-th output feature map of the l-th layer. This represents the i-th input feature map of the (l-1)-th layer. This represents the weights of the convolution kernel connecting the i-th input feature map and the j-th output feature map. M represents the corresponding bias term. j This represents the set of indices of all input feature maps connected to the j-th output feature map;

[0035] S32. Perform global average pooling and global max pooling on the initial feature maps respectively to obtain the average pooled features along the channel dimension. With max pooling characteristics And feature normalization is performed on each channel separately;

[0036] S33. The normalized average pooling features are respectively... and max pooling features The channel attention weight vector M is obtained by independently transforming a multilayer perceptron structure with one-quarter the number of hidden layer neurons compared to the original number of channels, then summing the transformed results channel by channel, and finally activating the structure using the sigmoid function. c (F):

[0037]

[0038] Where σ(·) represents the Sigmoid activation function;

[0039] S34. Perform average pooling and max pooling along the channel dimension on the feature maps enhanced by channel attention to generate spatial feature matrices. and And it forms a spatial attention fusion feature by connecting them along the channel direction;

[0040] S35. Spatial attention fusion features are convolved using a 7×7 two-dimensional convolution kernel with a stride of 1, outputting a single-channel spatial attention weight matrix M. s (F):

[0041]

[0042] Among them, f 7×7 This represents a 7×7 two-dimensional convolution operation, where σ(·) represents the Sigmoid activation function.

[0043] S36, Transfer the channel attention weight vector M c (F) Multiply by the corresponding feature map channel by channel, and then apply the spatial attention weight matrix M s (F) Multiply the feature map at each pixel position to achieve dual feature enhancement at both the channel and spatial levels;

[0044] S37. A path aggregation network is used to fuse feature maps of different scales after dual feature enhancement layer by layer to obtain a multi-scale fused feature map. A convolutional prediction and regression network is used to predict the bounding box coordinates and class confidence of the multi-scale fused feature map and output the target detection results.

[0045] Optionally, S4 specifically includes:

[0046] S41. Using the individual birds detected in the target detection results as nodes, assign each node spatial coordinates, detection bounding box size, and the timestamp of the current frame.

[0047] S42. Determine whether there is a dynamic connection between any two nodes based on the joint determination of the spatial distance threshold and the time interval threshold between nodes. The spatial distance threshold is defined as the sum of the average value and standard deviation of the distance between individual birds, and the time interval threshold is defined as an integer multiple of the acquisition interval between two consecutive video frames.

[0048] S43. For dynamically connected nodes, calculate the magnitude and direction angle of the relative velocity vector between nodes. The direction angle is represented by the angle between the line connecting two adjacent bird individuals and the horizontal direction, and record the rate of change of the angle.

[0049] S44. Based on the spatial distance between nodes and the rate of change of the direction angle of the relative velocity vector, the strength and type of the edge are determined together. The types of edges include approaching, moving away and parallel. The edge strength is the ratio of the magnitude of the relative velocity vector between nodes to the spatial distance.

[0050] S45. Each edge is marked using a hybrid identifier defined by the strength and type of the edge. The mark is stored together with the timestamp of the corresponding node, and a dynamic graph containing the spatiotemporal state history is constructed.

[0051] S46. Calculate the number of nodes in the local neighborhood, the proportion of edge types, and the variance of the velocity angle between the node and its neighboring nodes for each node in the dynamic graph, and aggregate the calculation results into node-level local features for each node.

[0052] S47. The node-level local features are aggregated into a global vector to obtain a clustering trend vector representing the dynamic features of the group.

[0053] Optionally, S5 specifically includes:

[0054] S51. Extract the spatial location coordinates and timestamps of each bird individual in consecutive video frames from the target detection results, and construct the flight trajectory sequence of each bird individual in chronological order. The flight trajectory sequence is represented in the form of a quintuple (x... t ,y t ,v xt ,v yt ,t), where x t ,y t v is the position coordinate. xt ,v yt These are the velocity vector components at the corresponding time points, where t is the timestamp;

[0055] S52. Divide each flight trajectory sequence into segments according to a time sliding window. The length of each segment is a preset number of frames L. Perform first-order difference processing on the velocity vector in each segment to extract velocity change features.

[0056] S53. Combine the spatial coordinates and velocity change characteristics of each trajectory segment with the corresponding time code to form a state vector S. t The input is fed into the lower-level Actor network, and the output is the current action suggestion vector A. t ;

[0057] S54. Construct an instant reward function R, wherein the reward function consists of a prediction accuracy reward term R. acc , Predictive timeliness reward item R time Ecological Value Reward Item R eco With consecutive correct reward item R cont Linear weighted composition:

[0058] R = αR acc +βR time +γR eco +δR cont ;

[0059] Where α, β, γ, δ are weighting coefficients;

[0060] S55, the prediction accuracy reward item R acc Defined as:

[0061]

[0062] S56, the prediction timeliness reward item R time Defined as the value of R when a correct prediction is made t steps ahead. time = +2×t, where t≤5, the ecological value reward item R eco Defined as predicting reproductive or migration behavior, with a value of R. eco =+20;

[0063] S57, the continuous correct reward item R cont Defined as having a value of R when making correct predictions for c consecutive steps. cont =+c 2 ;

[0064] S58. Employ an experience-first replay mechanism to convert the state vector S t Action suggestion vector A t Instant reward R and next state vector S t+1 The data are stored together in the experience replay pool and sorted by sampling priority according to the magnitude of the TD error. The policy is iteratively updated through the double-delay DDPG algorithm. The Critic network uses the minimum mean square error of the target value of the state-action pair as the loss function, and the Actor network updates its parameters by maximizing the expected reward function, and outputs the local short-term behavior decision results.

[0065] Optionally, S6 specifically includes:

[0066] S61. Calculate the historical prediction error feedback vector E t The historical prediction error feedback vector E t The average value within a sliding window of the difference between the predicted location coordinates and the actual observed location coordinates;

[0067] S62. Combine the clustering trend vector with the historical prediction error feedback vector E. t Vector concatenation is performed to generate the input state vector H of the high-level Actor network. t ;

[0068] S63, using state vector H tAs input, the high-level Actor network outputs the current multi-step predicted action distribution probability vector P(a|H) t The action distribution probability vector is a probability vector that includes the predicted probabilities of future multi-step action direction and speed amplitude;

[0069] S64. Define the high-level instantaneous reward function R. H The high-level instant reward function R H It is a weighted sum of the reward for prediction accuracy and the reward for prediction stability;

[0070] S65. The Actor-Critic algorithm is used to train the high-level Critic network, minimizing the expected loss function corresponding to the action distribution probability vector in the Critic network. After the high-level Critic network converges, the expected loss function is calculated based on the action distribution probability vector P(a|H). t The action sequence with the highest probability is selected as the global multi-step action prediction sequence.

[0071] Optionally, S7 specifically includes:

[0072] S71. Calculate the positional deviation Δ between the local short-term behavior decision-making results and the actual observation data. l ocal;

[0073] S72. Calculate the positional deviation Δ between the global multistep behavior prediction sequence and the actual observed data. g lobal;

[0074] S73. Define the positional deviation as the decision deviation E. dec =Δ local and prediction deviation E pred =Δ global Set the decision bias threshold Th dec and prediction deviation threshold Th pred , when E dec or E pred Automatic online fine-tuning is triggered when the threshold is exceeded;

[0075] S74. Based on the triggered bias, update the parameters of the low-level Actor-Critic network using a deep deterministic policy gradient algorithm, where the loss function L of the Critic network is... Q for:

[0076]

[0077] Target value y i The is:

[0078] y i =R i +γQ′(s i+1,μ′(s i+1 |θ μ′ )|θ Q′ );

[0079] Where Q(s) i ,a i |θ Q ) indicates that the Critic network is related to state s i With action a i Value estimation, θ Q For the Critic network parameters, R i For immediate reward, γ is the discount factor, Q′ and μ′ represent the target Critic network and the target Actor network, respectively, and θ Q′ and θ μ′ These represent the corresponding target network parameters;

[0080] S75. Update the gradient through the Actor network policy to maximize the value estimation function output by the Critic network, where the gradient process of the Actor network policy is as follows:

[0081]

[0082] Where, μ(s) i |θ μ ) indicates that the Actor network is in state s i The action output under θ μ These are the parameters of the Actor network;

[0083] S76. Using the local short-term behavioral decision results and the global multi-step as the bias gradient of the prediction sequence, the parameters of the YOLOv5 object detection model are updated by the Adam optimization algorithm.

[0084] The beneficial effects of this invention are:

[0085] (1) This invention introduces the DynamicCBAM++ module based on the channel and spatial joint attention mechanism into the YOLOv5 model and dynamically adjusts the attention window parameters according to environmental statistical features, thereby achieving high-precision detection of birds in complex backgrounds and small targets, effectively improving target recall and detection stability, and enhancing the applicability of the system in the wild natural environment.

[0086] (2) By constructing a dynamic graph neural network based on the spatiotemporal position and relative velocity of individual birds and using the local connection structure to calculate the clustering situation vector, this invention can realize the modeling and expression of the dynamic relationship of the collective behavior of bird flocks, significantly improve the accuracy of group behavior perception, and show stronger identification ability and early warning value in the early stage of potential avian influenza transmission.

[0087] (3) In terms of individual and group behavior prediction, this invention constructs a two-layer Actor-Critic deep reinforcement learning strategy framework and combines TD3 and SAC algorithms to realize hierarchical decision-making and joint optimization of local short-term and global multi-step behaviors. This effectively solves the problems of insufficient behavior prediction accuracy and poor strategy generalization ability in the existing technology and breaks through the bottleneck of complex temporal behavior modeling.

[0088] (4) This invention constructs an online meta-learning adaptive mechanism that includes historical prediction error feedback, and dynamically triggers the fine-tuning process of the YOLOv5 detection model and the Actor-Critic strategy network based on local and global deviations. This enables the model to maintain continuous performance and adapt to dynamic environments in long-term deployment scenarios, thereby effectively improving the stability and practicality of the avian influenza monitoring system. Attached Figure Description

[0089] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0090] Figure 1 This is a system flowchart of the bird behavior prediction method for avian influenza hosts proposed in this invention based on deep reinforcement learning. Detailed Implementation

[0091] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0092] refer to Figure 1 A method for predicting the behavior of avian influenza hosts based on deep reinforcement learning includes the following steps:

[0093] S1. Deploy multi-angle, variable frame rate cameras in the target habitat to collect raw video streams and obtain multi-source video data;

[0094] S2. Perform time-series frame interpolation, ambient lighting normalization, and background modeling and denoising on the multi-source video data to obtain preprocessed video data.

[0095] S3. Input the preprocessed video data into the improved YOLOv5 model to obtain the object detection results;

[0096] S4. Based on the target detection results, construct a dynamic graph neural network with the detected bird individuals as nodes, spatiotemporal location information and relative velocity as edges, and output a clustering situation vector.

[0097] S5. Input the flight trajectory of individual birds in the target detection results into the low-level Actor-Critic strategy module to generate local short-term behavior decision results;

[0098] S6. Feed the clustering situation vector and historical prediction error back into the high-level Actor-Critic strategy module to generate a global multi-step behavior prediction sequence.

[0099] S7. Compare the local short-term behavior decision results and the global multi-step behavior prediction sequence with the real-time collected actual bird behavior observation data, calculate the decision bias and prediction bias, and trigger online fine-tuning based on the decision bias and prediction bias to update the YOLOv5 detection model and Actor-Critic strategy network parameters.

[0100] By introducing an improved YOLOv5 model, target detection is achieved in multi-source video data of avian influenza host birds. A dynamic graph neural network based on spatiotemporal location and relative velocity is used to extract group behavior features. A two-layer Actor-Critic strategy structure is employed to perform hierarchical modeling and prediction of individual and group behaviors. Simultaneously, online fine-tuning is performed based on prediction errors to update the detection and strategy network parameters. These technical solutions synergistically improve the detection accuracy, temporal modeling capability, and model adaptability of bird behavior prediction in complex environments, making it suitable for long-term deployment and avian influenza risk monitoring scenarios with dynamic environmental changes.

[0101] In this embodiment, S1 specifically includes:

[0102] S11. Based on DGNSS and ground reflection markers for assisted positioning, several three-dimensional spatial control points are set up in the target habitat, and environmental point clouds and orthophotos are obtained by fusion of lidar and aerial UAV photogrammetry. The target habitat includes foraging grounds, breeding grounds and migratory stopover sites of avian influenza host birds.

[0103] S12. Using environmental point cloud and orthophoto, the particle swarm optimization algorithm is used to calculate the three-dimensional polyhedron of each camera's field of view and solve for the optimal installation position and pitch angle. The pitch angle satisfies the requirement that the field of view coverage is greater than 98% and the occlusion probability is less than 2%.

[0104] S13. Fix the variable frame rate camera at the optimal installation position, and perform internal and external parameter calibration by combining striped structured light with the AprilTag calibration board to obtain the focal length, principal point coordinates, and radial and tangential distortion coefficients of the variable frame rate camera.

[0105] S14. Equip the variable frame rate camera with an electronically controlled liquid lens. The electronically controlled liquid lens achieves focal length switching from 100 mm to 200 mm by adjusting the applied voltage, and dynamically determines the region of interest based on image block histogram equalization analysis and buffers the video frames of the corresponding region at a frequency of 30 fps.

[0106] S15. Clock synchronization is performed using the PTPv2 protocol combined with the GNSS link, and a drift correction algorithm is embedded in the clock synchronization process based on the wireless mesh network to correct the clock deviation in real time.

[0107] S16. Each camera is equipped with an MPPT-controlled solar cell module and a lithium-ion battery storage module. The video resolution and frame rate are automatically adjusted by monitoring the voltage and output power of the battery module and according to the preset power-performance mapping table.

[0108] S17. A circular buffer structure is used to store the original video streams of each camera, and the video streams are fused based on the timestamp and camera viewpoint identifier to generate multi-source video data.

[0109] By deploying a variable frame rate camera system in the target habitat, combined with DGNSS positioning, point cloud modeling, and particle swarm optimization for site selection, a high-coverage, low-obstruction camera deployment strategy was achieved. Precise geometric calibration was completed using striped structured light and AprilTag calibration. Simultaneously, an electrically controlled liquid lens and dynamic frame rate control mechanism, combined with MPPT power supply components and PTPv2 clock synchronization, improved the light adaptability, power supply stability, and time consistency of video acquisition in the field environment, thereby achieving highly reliable, multi-source, and high-resolution raw video acquisition of bird behavior.

[0110] In this embodiment, S2 specifically includes:

[0111] S21. Accumulate the pixel-level absolute difference of each pair of adjacent frames in the multi-source video data and compare it with the stillness discrimination threshold. Extract the frames corresponding to the frames whose total pixel difference is less than or equal to the stillness discrimination threshold as key frames, and classify the other frames corresponding to the frames as dynamic frames.

[0112] S22. An integrated visible light and near-infrared spectral band synchronous acquisition device is used to record environmental spectral distribution data. Based on the spectral distribution data and a pre-calibrated pixel-spectral response lookup table, the pixel brightness in key frames and dynamic frames is linearly mapped and normalized to generate a spectral correction frame set.

[0113] S23. Divide the spectral correction frame set into fixed time windows. For each time window, swap the positions of the first frame and the Nth frame in the storage sequence to balance the difference in time interval between frames, and output the reconstructed time series.

[0114] S24. Construct a dual-layer background buffer canvas, wherein the first background canvas is updated every N1 frames for long-term static scene modeling, and the second background canvas is updated every N2 frames for short-term dynamic scene capture. The two background canvases are merged by a pixel-level overlay strategy and then merged with the reconstructed time series frames frame by frame to obtain the foreground frame sequence.

[0115] S25. Add a chain of metadata tags containing UTC timestamp, camera number, ambient light intensity and ambient temperature to each frame in the foreground frame sequence.

[0116] S26. Based on the camera number and frame sequence number, the foreground frame sequence is cross-merged in the storage sequence according to the timestamp and camera view label to generate preprocessed video data.

[0117] By distinguishing between dynamic and static frames through pixel-level difference analysis and combining this with a spectrally corrected frame set to improve image brightness consistency, the stability of video data under complex lighting conditions is enhanced. Simultaneously, a dual-layer background caching mechanism is constructed to achieve dynamic modeling and fusion of long-term background and short-term scene data, effectively improving the accuracy and continuity of foreground extraction. Furthermore, by combining time window reconstruction, UTC timestamp marking, and camera number embedding, time alignment and feature unification of multi-source video sequences are achieved.

[0118] In this embodiment, S3 specifically includes:

[0119] S31. Input the preprocessed video data frame by frame into the YOLOv5s backbone network CSPDarknet, and process it layer by layer through cross-stage local connectivity structure and residual block stacking convolution to obtain the initial feature map:

[0120]

[0121] in, This represents the j-th output feature map of the l-th layer. This represents the i-th input feature map of the (l-1)-th layer. This represents the weights of the convolution kernel connecting the i-th input feature map and the j-th output feature map. M represents the corresponding bias term. j This represents the set of indices of all input feature maps connected to the j-th output feature map;

[0122] S32. Perform global average pooling and global max pooling on the initial feature maps respectively to obtain the average pooled features along the channel dimension. With max pooling characteristics And feature normalization is performed on each channel separately;

[0123] S33. The normalized average pooling features are respectively... and max pooling features The channel attention weight vector M is obtained by independently transforming a multilayer perceptron structure with one-quarter the number of hidden layer neurons compared to the original number of channels, then summing the transformed results channel by channel, and finally activating the structure using the sigmoid function. c (F):

[0124]

[0125] Where σ(·) represents the Sigmoid activation function;

[0126] S34. Perform average pooling and max pooling along the channel dimension on the feature maps enhanced by channel attention to generate spatial feature matrices. and And it forms a spatial attention fusion feature by connecting them along the channel direction;

[0127] S35. Spatial attention fusion features are convolved using a 7×7 two-dimensional convolution kernel with a stride of 1, outputting a single-channel spatial attention weight matrix M. s (F):

[0128]

[0129] Among them, f 7×7 This represents a 7×7 two-dimensional convolution operation, where σ(·) represents the Sigmoid activation function.

[0130] S36. Transfer the channel attention weight vector M c (F) Multiply by the corresponding feature map channel by channel, and then apply the spatial attention weight matrix M s (F) Multiply the feature map at each pixel position to achieve dual feature enhancement at both the channel and spatial levels;

[0131] S37. A path aggregation network is used to fuse feature maps of different scales after dual feature enhancement layer by layer to obtain a multi-scale fused feature map. A convolutional prediction and regression network is used to predict the bounding box coordinates and class confidence of the multi-scale fused feature map and output the target detection results.

[0132] By introducing an improved module that integrates channel and spatial attention mechanisms into the YOLOv5 backbone network CSPDarknet, multi-scale pooling and multilayer perceptron are used to calculate channel attention weights. Combined with 7×7 two-dimensional convolution, spatial attention weight maps are generated, and joint enhancement processing of different channels and different spatial regions in the feature map is achieved. This effectively improves the network's ability to perceive features of individual birds in occlusion, small targets and complex backgrounds, and enhances the model's detection accuracy and robustness in real-world environments.

[0133] In this embodiment, S4 specifically includes:

[0134] S41. Using the individual birds detected in the target detection results as nodes, assign each node spatial coordinates, detection bounding box size, and the timestamp of the current frame.

[0135] S42. Determine whether there is a dynamic connection between any two nodes based on the joint determination of the spatial distance threshold and the time interval threshold between nodes. The spatial distance threshold is defined as the sum of the average value and standard deviation of the distance between individual birds, and the time interval threshold is defined as an integer multiple of the acquisition interval between two consecutive video frames.

[0136] S43. For dynamically connected nodes, calculate the magnitude and direction angle of the relative velocity vector between nodes. The direction angle is represented by the angle between the line connecting two adjacent bird individuals and the horizontal direction, and record the rate of change of the angle.

[0137] S44. Based on the spatial distance between nodes and the rate of change of the direction angle of the relative velocity vector, the strength and type of the edge are determined together. The types of edges include approaching, moving away and parallel. The edge strength is the ratio of the magnitude of the relative velocity vector between nodes to the spatial distance.

[0138] S45. Each edge is marked using a hybrid identifier defined by the strength and type of the edge. The mark is stored together with the timestamp of the corresponding node, and a dynamic graph containing the spatiotemporal state history is constructed.

[0139] S46. Calculate the number of nodes in the local neighborhood, the proportion of edge types, and the variance of the velocity angle between the node and its neighboring nodes for each node in the dynamic graph, and aggregate the calculation results into node-level local features for each node.

[0140] S47. The node-level local features are aggregated into a global vector to obtain a clustering trend vector representing the dynamic features of the group.

[0141] By constructing a dynamic graph structure with individual birds in the target detection results as nodes, and combining the spatial distance and time interval between nodes to jointly determine dynamic edge relationships, and introducing the rate of change of orientation angle and the relative velocity magnitude to construct edge weights, the neighborhood movement relationships and structural evolution characteristics within the bird flock can be accurately characterized. Furthermore, by calculating the edge density, edge type ratio, and velocity distribution gradient of local regions, multi-level group interaction features are extracted, ultimately forming a clustering trend vector representing the global collective movement trend, thus improving the accuracy of structured modeling and dynamic perception of group behavior.

[0142] In this embodiment, S5 specifically includes:

[0143] S51. Extract the spatial location coordinates and timestamps of each bird individual in consecutive video frames from the target detection results, and construct the flight trajectory sequence of each bird individual in chronological order. The flight trajectory sequence is represented in the form of a quintuple (x... t ,y t ,v xt ,v yt ,t), where x t ,y t v is the position coordinate. xt ,v yt These are the velocity vector components at the corresponding time points, where t is the timestamp;

[0144] S52. Divide each flight trajectory sequence into segments according to a time sliding window. The length of each segment is a preset number of frames L. Perform first-order difference processing on the velocity vector in each segment to extract velocity change features.

[0145] S53. Combine the spatial coordinates and velocity change characteristics of each trajectory segment with the corresponding time code to form a state vector S. t The input is fed into the lower-level Actor network, and the output is the current action suggestion vector A. t ;

[0146] S54. Construct an instant reward function R, wherein the reward function consists of a prediction accuracy reward term R. acc , Predictive timeliness reward item R time Ecological Value Reward Item R eco With consecutive correct reward item R cont Linear weighted composition:

[0147] R = αR acc +βR time +γR eco +δR cont ;

[0148] Where α, β, γ, δ are weighting coefficients;

[0149] S55, the prediction accuracy reward item R acc Defined as:

[0150]

[0151] S56, the prediction timeliness reward item R time Defined as the value of R when a correct prediction is made t steps ahead. time = +2×t, where t≤5, the ecological value reward item R eco Defined as predicting reproductive or migration behavior, with a value of R. eco =+20;

[0152] S57, the continuous correct reward item R cont Defined as having a value of R when making correct predictions for c consecutive steps. cont =+c 2 ;

[0153] S58. Employ an experience-first replay mechanism to convert the state vector S t Action suggestion vector A t Instant reward R and next state vector S t+1 The data are stored together in the experience replay pool and sorted by sampling priority according to the magnitude of the TD error. The policy is iteratively updated through the double-delay DDPG algorithm. The Critic network uses the minimum mean square error of the target value of the state-action pair as the loss function, and the Actor network updates its parameters by maximizing the expected reward function, and outputs the local short-term behavior decision results.

[0154] By constructing a low-level Actor-Critic policy structure with a five-tuple sequence of flight trajectories as input, and combining local temporal segmentation and velocity change feature extraction mechanisms, a state-action-reward mapping relationship was established. A weighted reward function incorporating four indicators—accuracy, timeliness, ecological value, and continuity—was adopted to refine the evaluation criteria for individual behavior, enhancing the diversity and guidance of the training signal. Combining a priority experience replay mechanism and a TD error ranking sampling method improved the training efficiency and stability of the low-level policy in short-term behavior prediction tasks.

[0155] In this embodiment, S6 specifically includes:

[0156] S61. Calculate the historical prediction error feedback vector E t The historical prediction error feedback vector E t The average value within a sliding window of the difference between the predicted location coordinates and the actual observed location coordinates;

[0157] S62. Combine the clustering trend vector with the historical prediction error feedback vector E. t Vector concatenation is performed to generate the input state vector H of the high-level Actor network. t ;

[0158] S63, using state vector H t As input, the high-level Actor network outputs the current multi-step predicted action distribution probability vector P(a|H) t The action distribution probability vector is a probability vector that includes the predicted probabilities of future multi-step action direction and speed amplitude;

[0159] S64. Define the high-level instantaneous reward function R. H The high-level instant reward function R HThe weighted sum of the prediction accuracy reward and the prediction stability reward:

[0160] R H =ηR accH +θR stab ;

[0161] Among them, R accH R represents the reward for accurate prediction of behavior. stab The prediction stability reward is defined as the reciprocal of the fluctuation range of the prediction error within n consecutive steps, where η and θ are weighting coefficients.

[0162] S65. The Actor-Critic algorithm is used to train the high-level Critic network, minimizing the expected loss function corresponding to the action distribution probability vector in the Critic network:

[0163]

[0164] Among them, Q(H t a) represents the input state vector H of the Critic network. t and the value prediction function for action a;

[0165] S66. After the high-level Critic network converges, according to the action distribution probability vector P(a|H) t The action sequence with the highest probability is selected as the global multi-step action prediction sequence.

[0166] By concatenating the clustering situation vector with historical prediction error feedback to form a state vector, which is then input into a high-level Actor network, the output multi-walk action distribution probability is determined. An immediate reward function is constructed by combining prediction accuracy and stability, and the Actor-Critic method is introduced to train the high-level policy. This method enhances the model's ability to dynamically model complex group behavior, strengthens the policy's sensitivity and constraint to prediction error fluctuations, and ultimately achieves high-confidence predictions of future multi-walk action directions and speeds, improving the stability and accuracy of long-term behavior extrapolation.

[0167] In this embodiment, S7 specifically includes:

[0168] S71. Calculate the positional deviation Δ between the local short-term behavior decision-making results and the actual observation data. local :

[0169]

[0170] Among them, (x li ,y li (x) represents the local short-term predicted location of the i-th bird individual. reali ,y reali() represents the actual location of the corresponding bird individual, and N is the total number of bird individuals observed at the current moment;

[0171] S72. Calculate the positional deviation Δ between the global multistep behavior prediction sequence and the actual observed data. global :

[0172]

[0173] Among them, (x gj,t ,y gj,t (x) represents the global multi-step predicted position of the j-th bird individual in the next t steps, (x) realj,t ,y realj,t () represents the corresponding actual location, M is the number of individual birds, and T is the number of prediction steps;

[0174] S73. Define the positional deviation as the decision deviation E. dec =Δ local and prediction deviation E pred =Δ global Set the decision bias threshold Th dec and prediction deviation threshold Th pred , when E dec or E pred Automatic online fine-tuning is triggered when the threshold is exceeded;

[0175] S74. Based on the triggered bias, update the parameters of the low-level Actor-Critic network using a deep deterministic policy gradient algorithm, where the loss function L of the Critic network is... Q for:

[0176]

[0177] Target value y i The is:

[0178] y i =R i +γQ′(s i+1 ,μ′(s i+1 |θ μ′ )|θ Q′ );

[0179] Where Q(s) i ,a i |θ Q ) indicates that the Critic network is related to state s i With action a i Value estimation, θ Q For the Critic network parameters, R iFor immediate reward, γ is the discount factor, Q′ and μ′ represent the target Critic network and the target Actor network, respectively, and θ Q′ and θ μ′ These represent the corresponding target network parameters;

[0180] S75. Update the gradient through the Actor network policy to maximize the value estimation function output by the Critic network, where the gradient process of the Actor network policy is as follows:

[0181]

[0182] Where, μ(s) i |θ μ ) indicates that the Actor network is in state s i The action output under θ μ These are the parameters of the Actor network;

[0183] S76. Using the local short-term behavioral decision results and the global multi-step as the bias gradient of the prediction sequence, the parameters of the YOLOv5 object detection model are updated by the Adam optimization algorithm.

[0184] By defining the deviation between local short-term behavior and global multi-step behavior prediction of location information, and introducing a deviation threshold triggering mechanism to achieve automatic fine-tuning of the policy network, the parameters of the low-level Actor-Critic network are updated by combining the deep deterministic policy gradient algorithm. Furthermore, the parameters of the YOLOv5 detection model are adjusted by jointly using gradient descent and Adam optimization methods, thereby realizing the adaptive error correction of the model in the continuous observation environment. This effectively improves the system's ability to maintain prediction accuracy and adapt to dynamic environments during long-term operation.

[0185] Example 1: To verify the feasibility of the present invention in practice, the present invention was applied to a bird ecological monitoring system in a certain prefecture-level region to predict in real time the changes in flight behavior and collective dynamic trends of migratory bird groups during their migration, so as to assist in the establishment of a local avian influenza risk early warning mechanism.

[0186] In practice, researchers deployed 12 cameras with automatic zoom and variable frame rate capabilities in a wetland protected area along a migratory bird flyway. They used DGNSS and LiDAR point clouds for spatial calibration to ensure full coverage and low occlusion of the multi-camera deployment. Video data collected by all cameras underwent illumination correction, frame rate reconstruction, and dual-layer background modeling as proposed in this invention, resulting in high-quality pre-processed video input. The target detection module uses a YOLOv5 network with an integrated dynamic CBAM attention mechanism, achieving an average detection accuracy of 87.2% even in scenes with occlusion exceeding 30%. Subsequently, the system extracts group behavior vectors by constructing a dynamic graph neural network and embeds flight trajectories into a low-level Actor-Critic module to predict individual behavior. A high-level policy module is then constructed by combining historical errors and situational vectors to complete multi-step predictions.

[0187] During the model training phase, the TD3 and SAC algorithms were used to stabilize and optimize the upper and lower layer strategies respectively. Online fine-tuning was performed based on a prediction bias triggering mechanism, and the Adam optimizer was used for joint parameter updates. On the 14th day after deployment, the system completed the behavior prediction task for more than 5,000 bird flock segments, with an average inference time of 34 milliseconds per frame and a system stability rate of 99.3%. Among the comparison between actual observed trajectories and predicted trajectories, the average error of local behavior prediction was controlled within ±0.21 meters, and the average error of global behavior prediction was controlled within ±0.28 meters. The model maintained stable prediction capabilities even under dynamic weather conditions.

[0188] The table below shows the error comparison of some typical samples during the experiment and the change in system recall before and after fine-tuning:

[0189] Table 1 Comparison of Sample Error and Recall Rate Before and After Fine-tuning in Migratory Bird Behavior Prediction

[0190]

[0191] As shown in Table 1, the system of this invention exhibits high consistency and reliability in both local trajectory prediction and group multi-walking behavior trend prediction across multiple typical bird behavior samples. Specifically, the average difference between the predicted and measured trajectory errors is no more than 0.14 meters, with a maximum deviation of 0.18 meters. The group behavior prediction error remains within ±0.10 meters across all samples, indicating that the structure combining local Actor-Critic and high-level policy networks with dynamic graph modeling effectively captures the interaction characteristics between individual movement trends and overall group mobility. Particularly in sample BSP-031, the difference between the predicted and measured group behavior error is only 0.05 meters, demonstrating the good stability of the high-level multi-walking behavior distribution prediction network when facing dense group behavior.

[0192] Regarding model fine-tuning, the system can adaptively initiate small-batch updates without interrupting the main workflow by using a joint threshold triggering mechanism for local and global prediction deviations during online monitoring. Taking sample groups BSP-001 to BSP-046 as examples, the average recall rate before fine-tuning was 80.98%, while after 8 to 10 rounds of online local backpropagation updates based on deviation triggering, the average recall rate increased to 88.24%, an improvement of 7.26%. In the prediction results after fine-tuning, the fluctuation values ​​of various errors decreased significantly, indicating that the fine-tuning mechanism not only improved the detection accuracy but also enhanced the robustness of the model under conditions of changing data distribution.

[0193] Of particular note is that on the 19th day of system deployment, two consecutive strong wind events occurred in the area, with maximum wind speeds reaching 17.2 m / s. This caused sudden changes in the direction of bird movement and increased local obstruction, resulting in a short-term drop in detection accuracy to 78.4%. After identifying three consecutive frames of local behavioral deviations exceeding the threshold, the system automatically entered a high-frequency fine-tuning state. Combining real-time data acquisition, it initiated eight rounds of reverse updates and performed gentle parameter replacements, ultimately restoring the accuracy to 86.5% within approximately 2 minutes and maintaining this level for about 6 hours. This demonstrates that the system possesses stable adjustment and self-recovery capabilities under dynamic environmental interference.

[0194] To further quantify the system's error correction effect and improved behavior prediction capability, experiments were conducted to compare the performance of the traditional YOLOv5 detection + LSTM behavior prediction method with the proposed YOLOv5-DCAM + Actor-Critic model on the same dataset. The traditional method showed a local average error of 3.74 meters and a global average error of 6.01 meters in trajectory prediction, while the proposed method showed errors of 2.94 meters and 4.89 meters respectively, representing error reduction rates of 21.4% and 18.6%. Furthermore, the traditional method exhibited a significantly higher false detection rate in scenarios with occlusion rates exceeding 20%, while the proposed solution, through the introduction of spatial attention enhancement structures and group dynamic structure modeling, controlled the average false detection rate to within 6.7%, demonstrating more stable performance.

[0195] In summary, the analysis shows that this invention not only maintains high-precision prediction capabilities for individual and group behavior under normal climatic conditions, but also demonstrates a good adaptive update mechanism in dynamically changing scenarios. Through continuous error feedback control, joint modeling strategies, and high-frequency fine-tuning processes, the system can effectively address typical field monitoring challenges such as long-term deployment, environmental uncertainty, and occlusion interference, demonstrating promising engineering application prospects. This embodiment fully verifies the comprehensive performance advantages of this invention in non-intrusive behavior perception, deep prediction optimization, and real-time model maintenance.

[0196] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for predicting the behavior of avian influenza host birds based on deep reinforcement learning, characterized in that, The method comprises the following steps: S1, deploying multi-angle and variable frame rate cameras in the target habitat, collecting raw video streams, and obtaining multi-source video data; S2, performing time sequence frame insertion, environment light normalization and background modeling denoising on the multi-source video data to obtain preprocessed video data; S3, inputting the preprocessed video data into an improved YOLOv5 model to obtain target detection results; S4, based on the target detection results, constructing a dynamic graph neural network with the detected bird individuals as nodes and the spatio-temporal position information and relative speed as edges, and outputting a flocking trend vector; S5, inputting the flight trajectory of the bird individuals in the target detection results into a low-level Actor-Critic policy module to generate a local short-term behavior decision result; S6, inputting the flocking trend vector and historical prediction error feedback into a high-level Actor-Critic policy module to generate a global multi-step behavior prediction sequence; S7, comparing the local short-term behavior decision result and the global multi-step behavior prediction sequence with the actual bird behavior observation data collected in real time, calculating the decision deviation and the prediction deviation, and triggering online fine-tuning based on the decision deviation and the prediction deviation to update the YOLOv5 detection model and the Actor-Critic policy network parameters.

2. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S1 specifically comprises: S11, based on DGNSS and ground reflection marker assisted positioning, arranging a plurality of three-dimensional space control points in the target habitat, and using a laser radar and aerial unmanned aerial vehicle photogrammetry fusion method to obtain environment point cloud and orthographic image, wherein the target habitat includes foraging ground, breeding ground and migration stopover ground of avian influenza host birds; S12, using the environment point cloud and orthographic image, calculating the three-dimensional polyhedron of each camera view and solving the optimal installation position and pitch angle by a particle swarm optimization algorithm, wherein the pitch angle satisfies the field of view coverage rate greater than 98% and the occlusion probability less than 2%; S13, fixing the variable frame rate camera at the optimal installation position, and performing internal and external parameter calibration through the combination of stripe structured light and AprilTag calibration board to obtain the focal length, principal point coordinates and radial and tangential distortion coefficients of the variable frame rate camera; S14, providing the variable frame rate camera with an electrically controlled liquid lens, which realizes 100mm to 200mm focal length switching by adjusting the applied voltage, and dynamically determines the region of interest based on image block histogram equalization analysis and caches the video frames of the corresponding region at a frequency of 30fps; S15, using PTPv2 protocol combined with GNSS link for clock synchronization, and embedding a drift correction algorithm in the clock synchronization process based on wireless mesh network to correct the clock deviation in real time; S16, providing each camera with an MPPT controlled solar cell module and a lithium ion battery storage module, automatically adjusting the video resolution and frame rate by monitoring the voltage and output power of the battery module according to the preset power-performance mapping table; S17, using a ring buffer structure to store the raw video streams of each camera, and fusing the video streams based on the timestamp and camera view angle identifier to generate multi-source video data.

3. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S2 specifically comprises: S21, accumulate the pixel-level absolute difference value of each pair of adjacent frames in the multi-source video data, and compare it with a static discrimination threshold, and extract the corresponding frames of the frame pair with a pixel difference value sum less than or equal to the static discrimination threshold as key frames, and classify the corresponding frames of other frame pairs as dynamic frames; S22, record the environmental spectrum distribution data by using an integrated visible light and near-infrared spectrum synchronous acquisition device, and linearly map and normalize the pixel brightness in the key frames and the dynamic frames according to the spectrum distribution data and a pre-labeled pixel-spectrum response lookup table, to generate a spectrum correction frame set; S23, divide the spectrum correction frame set according to a fixed time window, exchange the positions of the first frame and the Nth frame in the storage sequence for each time window to balance the frame interval difference, and output a reconstructed time sequence; S24, construct a double-layer background cache canvas, wherein a first background canvas is updated every N1 frames for long-term static scene modeling, and a second background canvas is updated every N2 frames for short-term dynamic scene capture, and the two layers of background canvases are fused by a pixel-level superposition strategy and merged with the reconstructed time sequence frame by frame to obtain a foreground frame sequence; S25, sequentially attach chain metadata tags including UTC time stamps, camera numbers, environmental light intensities and environmental temperatures to each frame in the foreground frame sequence; S26, cross-merge the foreground frame sequence according to the time stamps and camera perspective tags in the storage sequence according to the camera number and frame number, to generate preprocessed video data.

4. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S3 specifically comprises: S31, input the preprocessed video data frame by frame into the backbone network CSPDarknet of YOLOv5s, and obtain an initial feature mapping through cross-stage local connection structure and residual block stacking layer-by-layer convolution processing: wherein, represents the jth output feature map of the lth layer, represents the ith input feature map of the l-1th layer, represents the convolution kernel weight connecting the ith input feature map and the jth output feature map, represents the corresponding bias term, M j represents the index set of all input feature maps connected with the jth output feature map; S32, respectively, global average pooling and global maximum pooling are performed on the initial feature map to obtain average-pooled features in the channel dimension and maximum-pooled features and feature normalization is performed on each channel respectively; S33, respectively normalize the average pooling features and the maximum pooling features The transformed results are added channel by channel, and then activated by a Sigmoid function to obtain a channel attention weight vector M c (F): Wherein, σ(·) represents the Sigmoid activation function; S34, average pooling and maximum value pooling are respectively performed on the channel attention enhanced feature mapping along the channel dimension to generate a spatial feature matrix and And a spatial attention fusion feature is formed in series along the channel direction. S35, the spatial attention fusion feature is subjected to a convolution operation through a two-dimensional convolution kernel with a size of 7x7, a step length of 1, and an output of a single-channel spatial attention weight matrix M s (F): where f 7×7 denotes a 7x7 size two-dimensional convolution operation, and σ(·) denotes a Sigmoid activation function; S36, the channel attention weight vector M c (F) multiply the corresponding feature map by channel by channel, and multiply the spatial attention weight matrix M s (F) multiply the feature map by pixel position by pixel position, while realizing double feature enhancement at the channel level and the spatial level; S37, use a path aggregation network to layer-by-layer fuse different scale feature maps after double feature enhancement, obtain a multi-scale fusion feature mapping, and predict the boundary box coordinates and class confidence of the multi-scale fusion feature mapping through a convolution prediction regression network, and output a target detection result.

5. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S4 specifically comprises: S41, taking the bird individuals detected in the target detection result as nodes, giving each node a spatial coordinate, a detection bounding box size and a time stamp of the current frame; S42, jointly determine whether there is a dynamic connection relationship between any two nodes based on a spatial distance threshold and a time interval threshold, wherein the spatial distance threshold is defined as the sum of the average value and the standard deviation of the distance between the bird individuals, and the time interval threshold is defined as an integer multiple of the interval between two consecutive video frames; S43, for the dynamically connected nodes, calculate the size and direction angle of the relative velocity vector between the nodes, the direction angle is represented by the included angle between the connecting line of the adjacent two bird individuals and the horizontal direction, and record the change rate of the included angle; S44, determine the strength and type of the edge according to the spatial distance between the nodes and the change rate of the direction angle of the relative velocity vector, the type of the edge includes approaching type, moving away type and parallel type, and the edge strength is the ratio of the modulus of the relative velocity vector between the nodes to the spatial distance. S45, label each edge with a hybrid identifier defined by the strength and type of the edge, store the label together with the timestamp of the corresponding node, and construct a dynamic graph containing the history of spatio-temporal states; S46, calculate the number of nodes in the local neighborhood, the proportion of edge types, and the variance of the speed angle between the node and the neighborhood nodes for each node in the dynamic graph, and aggregate the calculation results into node-level local features; S47, aggregate the node-level local features into a global vector in turn to obtain a group gathering trend vector representing the group dynamic characteristics.

6. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S5 specifically comprises: S51. Extract the spatial location coordinates and timestamps of each bird individual in consecutive video frames from the target detection results, and construct the flight trajectory sequence of each bird individual in chronological order. The flight trajectory sequence is represented in the form of a quintuple (x... t ,y t ,v xt ,v yt ,t), where x t ,y t v is the position coordinate. xt ,v yt These are the velocity vector components at the corresponding time points, where t is the timestamp; S52, segment each flight trajectory sequence according to a time sliding window, each segment has a preset frame number L, and perform first-order difference processing on the speed vector in each segment to extract the speed change feature; S53, encode the spatial coordinates, velocity variation characteristics of each trajectory segment and the corresponding time into a state vector S t , input into the low-level Actor network, and output the current action recommendation vector A t ; S54, constructing an immediate reward function R, which is composed of a prediction accuracy reward term R acc , a prediction timeliness reward term R time , an ecological value reward term R eco and a continuous correctness reward term R cont linearly weighted: R = aR + βR + γR + δR acc + βR time + γR eco + δR cont ; Wherein, α, β, γ, δ are weight coefficients; S55, the prediction accuracy reward term R acc is defined as: S56, the prediction and timeliness reward item R time defined as predicting correctly t steps ahead, taking the value R time = + 2 x t, where t < 5, the ecological value reward item R eco defined as predicting breeding or migration behavior, taking the value R eco = + 20; S57、 the continuous correct reward item R cont defined as a continuous c-step correct prediction, taking the value R cont = +c 2 ; S58, the state vector S is updated using an experience priority replay mechanism t , the action recommendation vector A t , the immediate reward R and the next state vector S t+1 are stored in an experience replay pool and are sampled in priority order according to the size of the TD error, policy iteration updates are performed by a double delay DDPG algorithm, the Critic network uses the target value of the state-action pair as the loss function of the minimum mean square error, the Actor network updates the parameters by maximizing the expected reward function, and outputs a local short-term behavior decision result.

7. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S6 specifically comprises: S61, calculate a history prediction error feedback vector E t , the history prediction error feedback vector E t is the average value in the sliding window of the difference between the predicted position coordinates and the actual observed position coordinates; S62, combine the crowd tendency vector with the historical prediction error feedback vector E t Vector splicing is performed to generate the input state vector H of the high-level Actor network t ; S63, with the state vector H t as input, outputs a current multi-step predicted action distribution probability vector P(a|H t ), which is a probability vector containing the predicted probabilities of future multi-step behavior direction and speed magnitude. S64, define a high-level immediate reward function R H , the high-level immediate reward function R H is a weighted sum of a prediction accuracy reward and a prediction stability reward; S65, the high-level Critic network is trained using an Actor-Critic algorithm, and an expected loss function corresponding to the action distribution probability vector in the Critic network is minimized. After the high-level Critic network converges, the action sequence with the highest selection probability is selected as the global multi-step behavior prediction sequence according to the action distribution probability vector P(a|H t ).

8. The deep reinforcement learning-based avian influenza host bird behavior prediction method according to claim 1, wherein, The S7 specifically comprises: S71, calculate the position deviation Δ between the local short-term behavior decision result and the actual observation data l ocal; S72, compute position bias Δ between global multi-step behavior prediction sequence and actual observation data g lobal; S73, define the position deviations as decision deviations E dec = Δ local and prediction deviations E pred = Δ global , set decision deviation threshold Th dec and prediction deviation threshold Th pred , automatically trigger online fine-tuning when E dec or E pred exceeds the respective threshold S74, update the low-level Actor-Critic network parameters based on the triggered deviation through the deep deterministic policy gradient algorithm, wherein the loss function L of the Critic network is Q is: Target value y i The target value y is: y i = R i + γQ ′ (s i+1 , μ ′ (s i+1 | θ μ′ ) | θ Q′ ); Where Q(s) i ,a i |θ Q ) indicates that the Critic network is related to state s i With action a i Value estimation, θ Q For the Critic network parameters, R i For immediate reward, γ is the discount factor, Q′ and μ′ represent the target Critic network and the target Actor network, respectively, and θ Q′ and θ μ′ These represent the corresponding target network parameters; S75, update the gradient of the Actor network policy to maximize the value estimation function output by the Critic network, wherein the Actor network policy gradient process is: where μ(s i |θ μ ) represents the action output of the Actor network at state s i , and θ μ is the parameter of the Actor network. S76, update the parameters of the YOLOv5 target detection model through the Adam optimization algorithm with the bias gradient of the local short-term behavior decision result and the global multi-step behavior prediction sequence.

Citation Information

Patent Citations

  • Power transmission line bird detection method and system based on digital twinning, medium and equipment

    CN117351521A

  • Natural environment bird monitoring method based on multi-modal fusion deep learning and computer device

    CN119027775A