Bird inhabitation multi-mode driving and state monitoring device and method
By combining multimodal sensing units and deterrence execution components, the problem of power outages caused by birds inhabiting power transmission lines and substations is solved. This enables precise bird deterrence and real-time status monitoring of protective structures, providing long-term stable protection and monitoring capabilities.
Patent Information
- Application Number
- CN202511300914.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2026-01-02
AI Technical Summary
In existing technologies, birds' roosting and nesting on power transmission lines and substations leads to frequent power outages, and traditional methods of driving them away are prone to adaptive fatigue and lack real-time status monitoring and remote management capabilities.
The system employs a multimodal sensing unit that integrates infrared pyroelectric sensors, millimeter-wave radar, photosensitive devices, and vibration sensors to detect bird approach. It combines mechanical vibration, directional ultrasonic waves, and high-intensity LED flashes to drive birds away. The central processing unit performs data fusion and strategy optimization to achieve real-time monitoring and remote data transmission.
It enables precise bird control and real-time status monitoring of protective structures, reduces security threats to power systems, provides long-term stable protection and monitoring capabilities, and adapts to complex environmental changes.
Smart Images

Figure CN121242012A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bird deterrence and equipment status monitoring, specifically relating to a multimodal bird habitat deterrence and status monitoring device and method. Background Technology
[0002] Bird activity poses a threat to the safe and stable operation of transmission lines and substations, especially when birds perch or nest on high-voltage equipment, conductors, and insulators. This can not only cause power outages such as flashover and short circuits but also affect the long-term operational stability of the equipment. Furthermore, bird droppings and nesting materials can lead to equipment corrosion, poor heat dissipation, or even fires, seriously threatening the safe operation of the power grid. Traditional bird control measures often rely on simple methods such as installing bird spikes, using reflectors, or sound generators. However, these methods are prone to adaptive fatigue, with their effectiveness diminishing over time. Moreover, they lack real-time status monitoring and remote management capabilities in unattended scenarios.
[0003] To address the aforementioned issues, there is an urgent need for an intelligent device that integrates multiple sensing methods and deterrence techniques. This device should be able to accurately identify bird approach behavior and efficiently drive them away, while also monitoring the operational status of bird-proof structures in real time and transmitting data back to the user. This would achieve long-term, stable, and low-maintenance protection, thereby ensuring the safe and reliable operation of power transmission lines and substations under complex environmental conditions. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention aims to provide a multimodal bird habitat deterrence and status monitoring device. By integrating multiple sensors, it achieves accurate detection of bird approach and combines multi-channel interference methods of touch, hearing, and vision for efficient bird deterrence. Simultaneously, it monitors the operational status of the bird-proof structure in real time and remotely transmits the data back, realizing long-term, stable, protective, and monitoring integration in unattended scenarios, thereby effectively reducing the threat of bird damage to the safe operation of the power system.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A multimodal bird habitat deterrence and status monitoring device includes a multimodal sensing unit, a deterrence execution component, a central processing unit, and a wireless communication module. The multimodal sensing unit is used to collect real-time data on bird approach characteristics, the operational status of the protective structure, and ambient lighting conditions, providing basic data for deterrence strategies and status monitoring. This unit includes an infrared pyroelectric sensor, a millimeter-wave radar, a photosensitive device, and vibration and attitude sensors. The infrared pyroelectric sensor is used to determine whether a target has entered the monitoring area and outputs a trigger signal. The millimeter-wave radar is used to measure the distance, speed, and movement trajectory between the bird and the device. The photosensitive device is used to measure ambient brightness in real-time to assist in determining the bird's activity period and triggering optical deterrence strategies. The vibration and attitude sensors analyze whether the device is affected by external forces by detecting changes in the acceleration and angular velocity of the protective structure.
[0007] The deterrence execution component, based on real-time control commands issued by the central processing unit, performs multi-channel interference to deter birds. This component includes a mechanical vibration device, a directional ultrasonic transmitter, and a high-intensity LED flashing unit. The mechanical vibration device generates vibration signals via a DC eccentric vibration motor. The directional ultrasonic transmitter emits ultrasonic waves that can be perceived by birds, interfering with their hearing through intermittent, frequency-varying, or continuous wave modes. The high-intensity LED flashing unit emits short-duration, high-intensity intermittent flashes. The deterrence execution component dynamically adjusts its output parameters based on the deterrence effect.
[0008] The central processing unit is used to receive data collected by multimodal sensors and perform preprocessing, fusion analysis and drive-away strategy optimization; the wireless communication module automatically selects the optimal transmission method according to the environment, performs real-time data reporting and remote control command issuance; the central processing unit includes a data preprocessing module, a target recognition module, a behavior analysis module, a drive-away strategy optimization module, and a status monitoring and feedback module. The data preprocessing module is used to synchronize, denoise, and normalize multi-source raw data from image, acoustic, radar, and environmental sensors. The target recognition module uses the processed multimodal data as input to accurately determine the species, number, and location of birds. It employs an improved YOLOv7 framework for image target detection, combining backbone feature extraction, neck feature fusion, and head prediction to output candidate target boxes. Overlapping boxes are removed using non-maximum suppression (NMS) to obtain high-confidence detection results, providing a basis for individual tracking in behavior analysis. The behavior analysis module, based on multimodal temporal modeling, infers bird activity patterns to quantify their willingness to stay, foraging, nesting, and flight behaviors. It uses a Transformer Encoder network to model the fused features, outputs the behavior category distribution, and uses maximization decoding to obtain the behavior determination result at the current moment.
[0009] The deportation strategy optimization module is used to abstract the dynamic system formed by multimodal perception, target recognition, and behavior analysis into a Markov decision process (MDP) to solve for the globally optimal deportation strategy. The state space includes information on bird species distribution, population estimation, spatial location, behavioral patterns, environmental conditions, and sensor confidence. The parameterized probabilistic model characterizes the bird state transition patterns and can be extended to a nonlinear continuous action dependency function to adapt to complex environments. Finally, the optimal strategy is solved by combining value iteration, policy iteration, or reinforcement learning algorithms to achieve closed-loop optimization from observation and behavior determination to action execution, and dynamic deportation and early warning management are realized through multi-level response strategies.
[0010] The state monitoring and feedback module continuously tracks the driving effect and changes in bird activity, inputting feedback data into the algorithm model for parameter updates and strategy iterations; it forms a training sample set through simulation and on-site data collection, including image, acoustic, radar, and environmental data, and uses a deep learning framework for incremental training and reinforcement learning optimization.
[0011] The method for multimodal bird habitat deterrence and status monitoring based on the aforementioned device includes the following steps:
[0012] S1: Data Acquisition: The multimodal sensor module is used to monitor bird activity in the target area in real time, including images, acoustics, millimeter-wave radar, and environmental parameters; sensor data is transmitted to the local edge computing device through a multi-channel interface;
[0013] S2: Data Preprocessing: The collected raw multimodal data is denoised, calibrated, and time-aligned; environmental noise and sudden anomalies are removed by Empirical Mode Decomposition (EMD) + Mahalanobis distance method, and Z-score and Min-Max normalization are used to eliminate dimensional differences between modes; and Joint Sparse Representation (JSR) is introduced to enhance cross-modal consistency features, generating a time-consistent, noise-suppressed, and modal-fused multimodal dataset.
[0014] S3: Target Recognition: Bird detection is performed on the processed image data using an improved YOLOv7 network, outputting candidate bounding boxes for bird species, number, and location, and enhancing detection accuracy by combining acoustic and environmental features; overlapping boxes are removed by non-maximum suppression to obtain high-confidence bird target information, providing a basis for individual tracking for behavior analysis;
[0015] S4: Behavioral Analysis: Based on a multimodal temporal modeling framework, bird activity patterns are inferred; image trajectories, acoustic features, and environmental data are fused and modeled using a Transformer temporal network to output bird behavior categories and dwelling intentions; combined with a drive-away demand function, bird population density, nesting tendency, and return probability are quantified to provide a decision-making basis for triggering drive-away strategies;
[0016] S5: Drive-off strategy optimization: The results of multimodal perception, target recognition and behavior analysis are abstracted into Markov decision process (MDP), defining the state space, action space, state transition function and immediate reward function; the optimal drive-off strategy is solved through dynamic programming or reinforcement learning algorithms, and the dynamic selection and intensity adjustment of various drive-off methods such as light, sound and vibration are carried out.
[0017] S6: Status Monitoring and Feedback: Real-time tracking of bird deterrence effects and behavioral changes, using feedback data for algorithm model updates and policy iterations; forming a multimodal training sample set through simulation and field data collection to support incremental training and reinforcement learning optimization;
[0018] S7: Data transmission and synchronization: Transmit multimodal sensor data, target recognition results and drive-away strategy execution information to the central server through the wireless communication module to achieve integrity verification and time synchronization, which facilitates unified management across devices and regions;
[0019] S8: Visualization: Real-time display of bird activity distribution, behavior patterns, and the execution of deterrence strategies on the monitoring platform; and display of bird density maps, behavior category change maps, and deterrence response status through a graphical interface for quick assessment of regional risks and operational decisions;
[0020] S9: Model Optimization: Incremental training and parameter tuning of target recognition, behavior analysis, and deterrence strategy models based on the latest collected data; automatic updates of YOLOv7 network and MDP strategy parameters to adapt to environmental changes and dynamic adjustments in bird behavior;
[0021] S10: System Management and Maintenance: The system monitors the hardware operating status and software module health, and promptly alerts maintenance personnel when anomalies are detected; it supports remote software updates, log viewing, and data management to ensure continuous and efficient operation of the device.
[0022] Furthermore, step S2, data preprocessing, specifically includes:
[0023] Since the sampling frequencies of different modalities differ, the multi-source data are first mapped to a unified time axis; let the true time series be T={t1,t2,…,t…}. N The data for a certain modality is:
[0024]
[0025] Where M is the number of modes, i.e., the total number of different sensors; m represents the mode index; X (m) This represents the time series observation data set of the m-th mode; N represents the observations of the m-th mode within its time range; mThis represents the number of sample points actually collected for the m-th mode; by combining linear interpolation and spline interpolation, an equally spaced sequence with a unified time step Δt is constructed:
[0026]
[0027] In the formula, x represents the estimated value of the m-th mode at a unified time point t obtained through interpolation; (m) (t k ) indicates that the mode was at the original sampling time point t k Observations at; w k The weight corresponding to the k-th interpolation node determines the degree of influence of different sampling points on the target time t; t represents the unified time point to be estimated; t j This represents the j-th sampling time in the original sampling sequence; This means that for all sampling points where j is not equal to k, the difference (tt) will be... j Multiply each term individually; This is a normalization factor used to ensure that at t=t k At that time, weight w k =1 and the remaining weights are zero, thus satisfying the interpolation consistency condition; j and k are both index variables, k specifies the current reference sampling point, and j represents all other sampling points except this point; for environmental noise and sudden anomalies, an empirical mode decomposition (EMD) + adaptive threshold detection method is introduced; EMD decomposition is performed on the input signal s(t):
[0028]
[0029] Among them, IMF i (t) represents the i-th intrinsic mode function, which corresponds to the local features of the signal at different time scales; K represents the number of IMF components extracted after EMD decomposition; r(t) is the residual term; low-frequency components are retained to represent bird behavior trends, while high-frequency noise is suppressed; simultaneously, Mahalanobis distance is used for anomaly detection on the residual sequence:
[0030]
[0031] Among them, D M (x) represents the Mahalanobis distance between sample x and the center of the population distribution; Σ -1 The inverse of the covariance matrix is used for weighted distance calculation; T represents the vector transpose; then the data point is considered an outlier and replaced by interpolation; to eliminate dimensional differences between different modes, Z-score standardization is used:
[0032]
[0033] Where μ and σ are the sample mean and standard deviation, respectively; for sparse events, a further...
[0034] Min-Max normalization enables feature enhancement of rare behavioral patterns;
[0035]
[0036] Where, x min and x max Let x and x' represent the minimum and maximum values of the modal data, respectively; x′ is the normalized value; to preserve the correlation between modalities before fusion, a Joint Sparse Representation (JSR) model is introduced; let the feature vectors of each modality be f. (m) The optimization objective is:
[0037]
[0038] Among them, D (m) is the dictionary matrix for the m-th modality, used to represent the basis space of features; α is the shared sparse coefficient vector, representing the common sparse structure among the multimodalities; Let represent the squared L2 norm, used to measure reconstruction error; ||·||1 represents the L1 norm, introducing sparsity constraints; λ is a regularization parameter used to balance reconstruction accuracy and sparsity; after temporal alignment, noise suppression, normalization, and sparse representation, a temporally consistent, noise-suppressed, and modal fusion multimodal dataset is finally constructed.
[0039]
[0040] Among them, f t Let f1, f2, ..., f be the fused feature vector at time t; d represents the dimension of the feature vector; f1, f2, ..., f T This represents the complete feature sequence at all times 1≤t≤T; this dataset can be directly input into the subsequent target recognition and drive-away decision module.
[0041] Furthermore, the target identification and behavior analysis in steps S3 and S4 specifically include:
[0042] The object detection part adopts the improved YOLOv7 framework, achieving a balance between real-time performance and accuracy; after data preprocessing, the input image has been enhanced and denoised, and its tensor representation is denoted as:
[0043]
[0044] Where H, W, and C represent the height, width, and number of channels of the image, respectively; the YOLOv7 network processes the input image through three stages: feature extraction (Backbone), feature fusion (Neck), and detection / prediction (Head), ultimately outputting a set of bird candidate targets:
[0045]
[0046] Among them, b i =(x i ,y i ,w i ,h i () represents the center coordinates and width and height of the i-th candidate box. For the target category label, s i ∈[0,1] represents the confidence score, indicating the reliability of the detection result; x i y represents the horizontal position of the candidate box center point; i The vertical position of the candidate box center point; w i h is the width of the candidate box. i The height of the candidate box; To represent the set of categories during model training; in the output, the targeting score is first calculated for each candidate box:
[0047]
[0048] Wherein, p(c i |X img ) represents the classification probability, and IoU represents the predicted bounding box b. i With real frame The intersection-over-union ratio (IoU) is calculated; finally, non-maximum suppression (NMS) is used to remove overlapping boxes, resulting in the final set of detection results. After detecting bird targets, it is necessary to further identify their behavioral patterns; within the time window [tk,t], the fused features are represented as follows:
[0049]
[0050] in, Represents image features of consecutive frames. Indicates acoustic characteristics, This represents environmental characteristics; these characteristics are modeled using a lightweight temporal network, the Transformer Encoder.
[0051]
[0052] in, h represents a time series modeling function with parameter θ. t This is the latent vector representation at time t; next, a classification layer is used to output the category distribution of bird behavior:
[0053] P(a|F t =Softmax(W·h) t +b) (15)
[0054] in, Represents the set of candidate behavior categories; by applying P(a|F t Maximize the decoding to obtain the bird's behavior judgment at the current moment:
[0055]
[0056] After completing target detection and behavior recognition, it is necessary to establish a connection between the recognition results and the expulsion strategy; a expulsion demand function D is proposed to quantify whether and how to implement expulsion.
[0057] D=αN+βR nest +γR return (17)
[0058] in, R represents the number of birds currently identified. nest ∈[0,1] represents the strength of nest-building tendency inferred from behavior recognition, R return ∈[0,1] represents the probability of birds returning, reflecting the group's willingness to stay. α, β, and γ are weighting parameters used to balance the contributions of different factors to the need for dispersal. The triggering condition for the dispersal strategy is defined as follows:
[0059] D>θ (18)
[0060] The threshold θ can be dynamically adjusted according to the actual scenario.
[0061] Furthermore, the optimization of the driving strategy in step S5 specifically includes:
[0062] MDP can be formalized as a quintuple:
[0063]
[0064] Among them, the state space It represents a complete description of the system's "perception-environment-habit" at any given moment; the high-dimensional feature vector formed after the fusion of multimodal sensors:
[0065] s t =[species_vec t ,count_vec t pos_grid t behavior t ,env t ,conf t ,h t ], (20)
[0066] Among them, species_vec tThe probability distribution of bird species, count_vec t For quantity estimation, pos_grid t For spatial location grid representation, behavior t For discretized behavior patterns, env t Characterizing environmental factors, conf t h represents the sensor confidence level. t This depicts the habits of bird flocks or their degree of domestication to stimuli; the corresponding movement space. Defined as the control combination of the driving device under different modes, it can be abstracted as an action vector:
[0067]
[0068] in, For optical driving sub-actions, it is a structured sub-vector; To drive the child's movements with hearing; The action is a tactile / mechanical vibratory sub-action; the state transition function P(s′|s,a) characterizes the system dynamics, that is, in state s... t Apply action a t Then transition to the next state s t+1 The probability distribution; the immediate reward function R(s,a) defines the target to be driven away, which can be formalized as:
[0069] R(s,a)=w eff ·f disp (s,a)-w energy ·E(a)-w disturb ·D(a)-w risk ·Ω(s,a)(22)
[0070] Among them, f disp (s,a) represents the expulsion effectiveness, measuring the positive effect of action a on the expulsion target in the current state s; E(a) represents the energy consumption, the energy / power cost consumed by action a during its implementation, reflecting resource consumption and range impact; D(a) represents the environmental disturbance index, measuring the disturbance or adverse impact of action a on non-target objects; Ω(s,a) represents the risk / compliance, measuring the potential risks or non-compliance consequences caused by action a in state s; weight w eff w energy w disturb w risk The positive scalar weights represent the relative importance of different expulsion effects, energy consumption, disturbances, and risks in the overall reward function; finally, the discount factor γ∈(0,1] is used to balance short-term expulsion effects with long-term ecological and safety constraints.
[0071]
[0072] Among them, V π (s t ) represents the expected cumulative reward under policy π; in modeling the optimization of the driving policy, one of the core tasks is to characterize the birds in a specific environmental state s. t With external action a t State transition probability
[0073] P(s′|s,a) is used for reasonable modeling; let the probability of a bird individual of category i leaving at time t be expressed as:
[0074]
[0075] in, For the logical sigmoid function, x t Indicates from state s t Environmental and population characteristics extracted, θ i This represents a vector of species sensitivity parameters, where φ is the external action effect parameter, and b i For bias terms;
[0076] Based on this probability definition, the evolutionary relationship of population size can be further written; if Let represent the number of individuals of category i at time t. Then, the number of individuals at the next time t+1 is approximately:
[0077]
[0078] in, The number of newly arrived or immigrant individuals can be modeled as a Poisson stochastic process; this formula reveals that the dynamic changes in population size are determined by a dual mechanism of "expulsion success rate" and "external immigration".
[0079] Furthermore, if action a t If the signal is not a discrete signal, but a continuous vector with intensity or combination, then the dependence of the departure probability on the action can be extended to a more general nonlinear function, i.e.:
[0080]
[0081] Among them, f ψ (·) can be fitted by a radial basis function (RBF) network, whose parameter ψ is trained in a data-driven manner, thus enabling the model to more flexibly capture complex environment-action-response nonlinear relationships; in the optimization of the expulsion strategy, reward function design, optimality conditions, constraint handling, and behavior modeling together constitute the core of the MDP solution; the instantaneous reward function should quantify the trade-offs between expulsion effect, energy consumption, disturbance, and safety risks, and its general form can be written as:
[0082]
[0083] Where, ΔN (i) t = N (i) tN (i) t+1 represents the number of individuals successfully expelled, E(a t D(a) represents the energy consumption of the action. t NT(s) measures disturbance to the surrounding environment or non-target organisms. t ,a t ) represents the penalty for protecting birds or non-target species, and w is an adjustable weight; in terms of policy solution, the state-value function and action-value function under policy π are defined as follows:
[0084]
[0085] Where π represents the strategy; r t+k The immediate reward obtained at time t+k is given by the reward function R(s,a); γ∈[0,1] is the discount factor used to decay the present value of future returns; condition |s t =s represents the expectation of s t =S is the starting point, and subsequent actions are sampled according to π; s t+1 The system transitions to the next state after action a is applied, and satisfies the Bellman expectation equation:
[0086]
[0087] and the optimal Bellman equation
[0088]
[0089] Where π(a|s) is the probability of choosing action a according to policy π in state s; P(s′|s,a) is the state transition probability, representing the probability distribution of transitioning to s′ after applying a in state s; V * (s) represents the maximum state value that the optimal policy can achieve among all possible policies; max a To maximize all actions; for discrete small state spaces, exact solutions are obtained through value iteration or policy iteration; if the state / action space is large or continuous, a parameterized Q-network approximation method is used.
[0090]
[0091] Where θ is the parameter vector; α > 0 is the learning rate, controlling the step size for parameter updates; max a′ Q(s′,a′;θ) represents the maximum action value of the next state s′ under the estimation. The gradient of the parameters indicates how to adjust the parameters to reduce the estimation error of the current state-action pair; furthermore, to characterize the habitual / domestication effects of birds on stimuli, habitual states are introduced. And correct state transitions and rewards:
[0092]
[0093] Among them, exposure (i) (a t κ is the exposure function, representing the stimulus exposure caused by the action being against the i-th type of target; κ≥0 is the habit growth coefficient, measuring the increase in habit due to each unit of exposure; δ≥0 is the recovery rate coefficient, representing the extent to which habit slowly decays over time under no-exposure conditions; τ is the time step or time interval; δ·τ represents the amount of natural recovery within that time step; the immediate reward multiplies the disengagement efficacy by a decay factor. Encourage diversified strategies to maintain long-term effectiveness;
[0094] Finally, parameter estimation is achieved by fitting the state transition model P(s′|s,a) with cross-entropy:
[0095]
[0096] And update the transition probability using an incremental statistical method:
[0097]
[0098] in, The overall loss function guides the update of parameters θ and φ; θ is the set of parameters of the policy network, which determines the action selection probability distribution π. θ The specific form of (a|s); φ is the parameter of the auxiliary model or discriminant network, used to model the probability distribution of bird responses (such as whether to leave); Let t be the actual departure label of the i-th individual; For the i-th bird predicted by the model in state s t Next, execute action a. t The probability of leaving later; This is a typical cross-entropy loss; reg is a regularization term that constrains the model parameters. The state transition probability is estimated empirically; N(s,a,s′) is the statistical count, representing the number of times "taking action a in state s and reaching s′" has been observed in historical or simulation samples; N(s,a) is the statistical count, representing the total number of times action a has been taken from state s in historical or simulation samples; the expulsion and early warning mechanism is set based on a multi-level response strategy, divided into first-level response, second-level response, and third-level response, which are defined as follows:
[0099] Level 1 Response: If the number of birds or their behavior patterns indicate slight signs of perching, and the demand for driving them away is slightly above the threshold, only low-intensity, low-interference light / sound / vibration stimuli will be initiated to alert operators.
[0100] Level 2 Response: When bird density increases or nesting tendency is obvious, and the demand function for driving away birds is close to the high threshold, moderate-intensity multimodal driving away measures are initiated, and on-site inspections or adjustments to the driving away strategy are recommended.
[0101] Level 3 Response: When birds are densely packed and stay for extended periods, potentially threatening facility safety, and the demand for bird removal exceeds the maximum threshold, a high-intensity combined bird removal operation will be immediately activated, with key areas under close monitoring. An alarm will also be generated to prompt operations personnel to take emergency measures.
[0102] A multimodal sample dataset was constructed by collecting data from simulated environments and real-world scenarios, including image, acoustic, radar, and environmental sensor data, forming labeled data on bird behavior and population distribution; [The dataset was then used...]
[0103] The PyTorch deep learning framework trains the target recognition, behavior analysis, and drive-away strategy optimization modules. By combining hyperparameter tuning, cross-validation, and reinforcement learning strategy iteration, it improves the system's recognition accuracy and strategy effectiveness.
[0104] The beneficial effects of this invention are as follows:
[0105] This invention utilizes multimodal sensors to collect bird images, acoustic, radar, and environmental information, and achieves real-time monitoring through data fusion and wireless transmission. By combining EMD, Mahalanobis distance anomaly detection, and Joint Sparse Representation (JSR) methods, noise and anomalies are effectively suppressed, improving data accuracy. Improved YOLOv7 and a lightweight Transformer are used to identify and predict bird species, numbers, locations, and behavioral patterns. Based on MDP-based deterrence strategy optimization, combined with multi-level response and habit modeling, light, sound, and vibration deterrence methods can be dynamically adjusted to achieve efficient and low-interference deterrence. The system integrates a closed loop of multimodal perception, behavior analysis, and strategy optimization, resulting in high monitoring accuracy, excellent deterrence efficiency, and strong robustness, providing an innovative, efficient, and eco-friendly bird pest control solution for farmland, airports, and power transmission lines. Attached Figure Description
[0106] Figure 1 This is a diagram of the YOLOv7 target recognition framework of the present invention;
[0107] Figure 2 This is a graph showing the training loss curve of the model in this invention;
[0108] Figure 3 This is a graph showing the data curves of the multimodal sensor of the present invention;
[0109] Figure 4 This is a graph showing the target detection and behavior analysis of the present invention.
[0110] Figure 5 This is a visualization curve showing the effect of the driving-off strategy of the present invention. Detailed Implementation
[0111] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are for illustrative purposes only and are not intended to limit the scope of the invention.
[0112] like Figure 1-4 As shown, this invention proposes a multimodal bird roosting and status monitoring device. It achieves accurate detection of bird approach by integrating multiple sensors and efficiently drives birds away using multi-channel interference methods combining tactile, auditory, and visual senses. Simultaneously, it monitors the operational status of the bird-proof structure in real time and remotely transmits the data back, realizing long-term, stable, integrated protection and monitoring in unattended scenarios to effectively reduce the threat of bird damage to the safe operation of power systems. The device includes a multimodal sensing unit, a bird-repelling execution component, a central processing unit, and a wireless communication module.
[0113] Multimodal Sensing Unit: This unit is used to collect real-time data on bird approach characteristics, the operational status of the protective structure, and ambient lighting conditions, providing fundamental data for deterrence strategies and status monitoring. It integrates an infrared pyroelectric sensor, millimeter-wave radar, photosensitive devices, and vibration and attitude sensors. The infrared pyroelectric sensor detects infrared radiation (8μm–14μm) from the bird's body surface to determine if a target has entered the monitoring area and outputs a trigger signal for target confirmation. The millimeter-wave radar can accurately measure the distance, speed, and trajectory between birds and the device, day and night, and in adverse weather conditions, effectively reducing false alarm rates. The photosensitive device measures ambient brightness in real time, assisting in determining the bird's activity period and triggering optical deterrence strategies. The vibration and attitude sensors detect changes in the acceleration and angular velocity of the protective structure to analyze whether the device is affected by external forces, providing a basis for structural health monitoring and deterrence strategy optimization. Multimodal sensor data is synchronously sampled by the acquisition module, Kalman filtered for noise reduction, and time-aligned before being sent to the central processing unit, achieving comprehensive, accurate, and stable monitoring.
[0114] The infrared pyroelectric sensor selected is the Panasonic EKMB1201111 series high-sensitivity dual-element passive infrared pyroelectric sensor, featuring low power consumption, high anti-interference capability, and a wide detection angle. It determines whether a target has entered the monitoring area by detecting the 8μm–14μm infrared signal emitted from the bird's body surface. The dual-element sensing element effectively cancels false alarms caused by background temperature changes. When a bird enters the detection range and causes a temperature change, the PIR sensor converts this into a weak electrical signal, which is then processed by internal amplification and filtering circuits before outputting a trigger signal. This signal is transmitted to the central processing unit as part of the multimodal data, and combined with other sensor data to achieve target confirmation and motion trend analysis.
[0115] The millimeter-wave radar utilizes the Infineon BGT24LTR11 series 24GHz frequency-modulated continuous wave (FMCW) millimeter-wave radar module, possessing strong anti-jamming capabilities and accurate distance / velocity measurement functions. By transmitting and receiving frequency-modulated continuous wave signals, the millimeter-wave radar accurately measures the distance, approach speed, and trajectory of birds relative to the device using the Doppler frequency shift and phase difference of the target's reflected signal. This sensor is unaffected by environmental conditions such as light, rain, and fog, and can operate stably day and night, as well as in adverse weather conditions. Combined with PIR sensor data, false alarms caused by background interference can be effectively reduced.
[0116] The photosensitive device (light sensor) uses the VEML7700 high-resolution digital light sensor, which supports wide dynamic range detection and adapts to different lighting conditions. By measuring changes in ambient brightness in real time, it helps determine the time periods of bird activity (such as dawn and dusk peaks) and trigger optical strategies for deterring birds (such as flashing lights). In low-light environments, the algorithm can automatically increase the flash frequency or combine it with ultrasonic interference to improve the deterrence effect.
[0117] The vibration and attitude sensors utilize the Bosch BMI160 high-precision six-axis IMU sensor, integrating a three-axis accelerometer and a three-axis gyroscope, to detect the vibration state and attitude changes of the protective structure. By detecting the acceleration changes and rotational angular velocities of the protective structure in the X, Y, and Z axes, the system analyzes whether the device has experienced displacement or attitude shifts due to external forces (such as bird strikes or wind effects). This data is not only used to determine direct bird contact behavior but also provides a basis for structural health monitoring and optimization of deterrence patterns.
[0118] All sensor data from the multimodal sensing unit are synchronously sampled by the data acquisition module, and noise reduction and time alignment are performed using algorithms such as Kalman filtering. The central processing unit fuses the multimodal data. First, it confirms the presence and trajectory of the target through a combination of PIR and millimeter-wave radar. Then, it combines illumination data to determine the triggering conditions for a deterrent strategy. Finally, it uses attitude data to determine the structural state and bird contact behavior. This multimodal fusion method significantly improves the accuracy of bird identification and deterrence while reducing the probability of false alarms and missed alarms.
[0119] The bird deterrence execution component, as the core execution unit, is responsible for using multi-channel interference to deter birds based on real-time control commands issued by the central processing unit, thereby reducing their dwelling, nesting, and pecking behaviors and lowering the risk of equipment failure. This component consists of a mechanical vibration device, a directional ultrasonic transmitter, and a high-intensity LED flashing unit. It can be used individually or in combination to implement single-mode, dual-mode, and tri-mode deterrence strategies. The mechanical vibration device generates vibration signals through a DC eccentric vibration motor, making it unsuitable for birds to stay. Its intensity and frequency can be dynamically adjusted to enhance targeting. The directional ultrasonic transmitter emits ultrasonic waves that can be perceived by birds, interfering with their hearing through intermittent, variable-frequency, or continuous wave modes while preventing adaptation. The high-intensity LED flashing unit emits short, high-intensity intermittent flashes, and the visual deterrence effect can be optimized through different wavelengths and flashing modes. The bird deterrence execution component supports adaptive control, dynamically adjusting output parameters according to the deterrence effect to achieve efficient, low-interference, and energy-saving deterrence.
[0120] The mechanical vibration device uses a high-efficiency DC eccentric vibration motor as its core actuator, creating an environment unsuitable for birds at the physical contact level. The motor generates vibration signals of varying intensities by adjusting the mass and rotational speed of the eccentric block. By fixing it to a crossarm, insulator support, or dedicated mounting base, the vibration is transmitted to the birds' usual habitat, forcing them to leave. The device, controlled by a CPU, can achieve periodic or random vibration patterns to disrupt the birds' comfort and stability. Simultaneously, the vibration intensity and frequency can be dynamically adjusted according to bird species, season, and environmental conditions. In conjunction with a status monitoring module, it activates immediately upon detecting bird gathering, improving the targeting and energy efficiency of the deterrent effect.
[0121] The directional ultrasonic transmitter system is equipped with a directional ultrasonic transmitter based on a high-frequency piezoelectric ceramic transducer, which can interfere with birds at the auditory level. The ultrasonic frequency is set within the birds' perceptible range (typically 15kHz-35kHz), and the deterrent effect is enhanced by modulating the waveform, frequency conversion, and directional transmission. The device can be controlled by a CPU to transmit intermittently or at varying frequencies, preventing birds from adapting to a fixed frequency. The algorithm module can determine the species and number of birds based on feedback from cameras and acoustic sensors, and match the most effective deterrent sound wave parameters to achieve precise deterrence.
[0122] The high-intensity LED flashing unit is equipped with a high-brightness LED flashing module, emitting high-intensity, short-duration, intermittent flashing signals that can create strong visual interference. Different wavelengths (such as white, blue, or green light) can be optimized to repel different bird populations. Flashing modes include pulse flashing, random flashing, and gradual brightening and dimming to avoid visual adaptation in birds. The LED flashing unit is CPU-controlled and can work in conjunction with mechanical vibration and ultrasonic transmitters to create complex stimulation in a multi-channel repelling mode, significantly improving the repelling effect. Simultaneously, the system supports automatic adjustment of light intensity at night to reduce the impact on nearby residents.
[0123] The three units of the bird deterrence execution component can be used individually or combined into a multimodal deterrence mode. The central processing unit, based on image, acoustic, and environmental data collected by the status monitoring module, uses YOLOv7 and behavior recognition algorithms to identify the species, number, and activity status of birds, and automatically selects a deterrence strategy. For example:
[0124] Single-modal deterrence: When the number of birds is small and the species are fixed, the lowest energy consumption single deterrence method, such as ultrasonic waves or flashing lights, should be given priority.
[0125] Dual-modal deterrence: When birds show a certain willingness to stay or a tendency to return, a combination of mechanical vibration and ultrasound or ultrasound and flashing light is used.
[0126] Trimodal repulsion: When there are many birds, a clear tendency to nest, or a long stay, three repulsion methods are activated simultaneously to form a strong compound stimulus, forcing the birds to leave.
[0127] In addition, the driving execution component has an adaptive driving mode, which can dynamically adjust parameters based on the driving effect evaluation to avoid excessive interference and energy waste, and achieve the goal of efficient and environmentally friendly driving.
[0128] Central Processing Unit (MCU): The MCU is the core control and data processing module of the device. It is responsible for receiving data collected by multimodal sensors and performing preprocessing, fusion analysis, and deflection strategy optimization. The MCU has built-in target recognition and behavior analysis algorithms, which can determine the species, number, and activity status of birds in real time, and automatically trigger corresponding deflection strategies based on the analysis results. Its edge computing capabilities enable key decisions to be made quickly locally, reducing data transmission latency. Furthermore, the MCU can upload monitoring data and deflection records to a remote operation and maintenance platform via a wireless communication module, enabling remote monitoring, strategy optimization, and closed-loop control, thus ensuring intelligent deflection and status monitoring.
[0129] Wireless communication module: Supports multi-standard transmission, including 4G / 5G cellular networks, LoRa long-range low-power communication, and Wi-Fi local area network, and can automatically select the optimal transmission method according to the environment. The module has built-in data encryption, compression, breakpoint resume, and remote firmware upgrade functions to ensure the security and stability of bird monitoring images, acoustic signals, environmental data, and deterrence records during transmission. Through collaboration with the central processing unit, this module enables real-time data reporting and remote control command issuance, thereby ensuring continuous operation and closed-loop linkage capabilities of the device in unattended scenarios.
[0130] The central processing unit includes a data preprocessing module, a target recognition module, a behavior analysis module, a drive-away strategy optimization module, and a status monitoring and feedback module.
[0131] The data preprocessing module is responsible for synchronizing, denoising, and normalizing raw data from multiple sources, including images, acoustics, radar, and environmental sensors, providing high-quality input for subsequent target recognition and deterrence decisions. The module unifies the sampling frequencies of different modalities through temporal alignment, interpolation modeling, and Lagrange or spline interpolation methods, and employs Empirical Mode Decomposition (EMD) + Mahalanobis distance for anomaly detection and burst event removal. Furthermore, Z-score and Min-Max normalization are introduced to eliminate dimensional differences, and Joint Sparse Representation (JSR) preserves intermodal consistency features, enhancing sparse event features. Finally, a noise-suppressed, modal-fused, and temporally consistent multimodal dataset is constructed, providing a reliable foundation for target recognition and behavior analysis.
[0132] The target recognition module uses processed multimodal data as input to accurately determine the species, number, and location of birds. It employs an improved YOLOv7 framework for image target detection, combining backbone feature extraction, neck feature fusion, and head prediction to output candidate bounding boxes. Non-maximum suppression (NMS) is used to remove overlapping boxes, resulting in high-confidence detection results and providing a foundation for individual tracking in behavioral analysis. Simultaneously, this module integrates acoustic and environmental features to enhance the auxiliary role of multi-source information in target detection, achieving high-precision identification of short-term bird activity and flock distribution.
[0133] The behavior analysis module, based on multimodal temporal modeling, infers bird activity patterns to quantify their willingness to stay, foraging, nesting, and flight. It models the fused features using a Transformer Encoder network, outputting the distribution of behavior categories, and utilizes maximization decoding to obtain the behavior determination result at the current moment. The module further introduces a deterrence demand function, mapping the target recognition results to the deterrence trigger decision, quantifying the demand of bird numbers, nesting tendency, and return probability on the deterrence strategy, providing a basis for strategy optimization.
[0134] The bird deterrence strategy optimization module abstracts the dynamic system formed by multimodal perception, target recognition, and behavior analysis into a Markov Decision Process (MDP) to solve for the globally optimal deterrence strategy. The state space includes information such as bird species distribution, population estimation, spatial location, behavioral patterns, environmental conditions, and sensor confidence levels. The action space defines control combinations of deterrence devices such as light, sound, and vibration, with an immediate reward function that considers deterrence success rate, energy consumption, environmental disturbance, and safety risks. The module characterizes the bird state transition patterns using a parameterized probabilistic model and can be extended to a nonlinear continuous action dependency function to adapt to complex environments. Finally, the optimal strategy is solved by combining value iteration, policy iteration, or reinforcement learning algorithms, achieving closed-loop optimization from observation and behavior determination to action execution. Dynamic deterrence and early warning management are realized through multi-level response strategies (level 1, level 2, and level 3).
[0135] The status monitoring and feedback module continuously tracks the effectiveness of bird deterrence and changes in bird activity, inputting feedback data into the algorithm model for parameter updates and strategy iteration. A training sample set is formed through simulation and on-site data collection, including image, acoustic, radar, and environmental data. Incremental training and reinforcement learning optimization are performed using a deep learning framework to achieve closed-loop intelligent control. This module ensures real-time adjustments to the deterrence strategy, improving the success rate while reducing energy consumption and environmental impact, achieving a complete closed-loop operation of multimodal perception—target recognition—behavior analysis—dynamic deterrence decision-making.
[0136] Accordingly, based on the aforementioned multimodal bird habitat deterrence and status monitoring device, this invention also proposes a multimodal bird habitat deterrence and status monitoring method, which includes the following steps:
[0137] S1: Data Acquisition: The multimodal sensor module is used to monitor bird activity in the target area in real time, including images, acoustics, millimeter-wave radar and environmental parameters; sensor data is transmitted to the local edge computing device through a multi-channel interface to ensure the continuity and real-time nature of data acquisition.
[0138] S2: Data Preprocessing: The collected raw multimodal data undergoes denoising, calibration, and temporal alignment. Environmental noise and sudden anomalies are removed using Empirical Mode Decomposition (EMD) combined with Mahalanobis distance, and Z-score and Min-Max normalization are employed to eliminate dimensional differences between modes. Joint Sparse Representation (JSR) is introduced to enhance cross-modal consistency, generating a temporally consistent, noise-suppressed, and modal-fused multimodal dataset, providing high-quality input for subsequent target recognition and behavior analysis.
[0139] Data preprocessing specifically includes:
[0140] Because different modalities have different sampling frequencies—for example, acoustic sensors have a higher sampling frequency while image sensors have a lower sampling frequency—multi-source data needs to be mapped to a unified time axis first. Let the real time series be...
[0141] T = {t1, t2, ..., t} N The data for a certain modality is:
[0142]
[0143] Where M is the number of modes, i.e., the total number of different sensors; m represents the mode index; X (m) This represents the time series observation data set of the m-th mode; N represents the observations of the m-th mode within its time range; m This represents the number of sample points actually collected for the m-th mode; by combining linear interpolation and spline interpolation, an equally spaced sequence with a unified time step Δt is constructed:
[0144]
[0145] In the formula, x represents the estimated value of the m-th mode at a unified time point t obtained through interpolation; (m) (t k ) indicates that the mode was at the original sampling time point t k Observations at; w kThe weight corresponding to the k-th interpolation node determines the degree of influence of different sampling points on the target time t; t represents the unified time point to be estimated; t j This represents the j-th sampling time in the original sampling sequence; This means that for all sampling points where j is not equal to k, the difference (tt) will be... j Multiply each term individually; This is a normalization factor used to ensure that at t=t k At that time, weight w k =1 and the remaining weights are zero, thus satisfying the interpolation consistency condition; j and k are both index variables, k specifies the current reference sampling point, and j represents all other sampling points except this point; this formula is based on the Lagrange interpolation polynomial, which can reconstruct continuous signals in low-frequency modes and improve the accuracy of cross-modal time alignment.
[0146] To address environmental noise (such as wind noise and mechanical noise) and sudden anomalies (such as false triggering and light flicker), an empirical mode decomposition (EMD) + adaptive threshold detection method is introduced. EMD decomposition is performed on the input signal s(t):
[0147]
[0148] Among them, IMF i (t) represents the i-th intrinsic mode function, corresponding to the local features of the signal at different time scales; K represents the number of IMF components extracted after EMD decomposition; r(t) is the residual term. Low-frequency components (representing bird behavior trends) are retained, while high-frequency noise is suppressed. Simultaneously, Mahalanobis distance is used for anomaly detection on the residual sequence.
[0149]
[0150] Among them, D M (x) represents the Mahalanobis distance between sample x and the center of the population distribution; Σ -1 The inverse of the covariance matrix is used for weighted distance calculation; T represents the vector transpose; then the data point is considered an outlier and replaced by interpolation; to eliminate dimensional differences between different modes, Z-score standardization is used:
[0151]
[0152] Where μ and σ are the sample mean and standard deviation, respectively; for sparse events (such as birds’ short-term roosting behavior), Min-Max normalization is further introduced to enhance the features of rare behavioral patterns.
[0153]
[0154] Where, x min and x max Let f represent the minimum and maximum values of the modality data, respectively; x′ is the normalized value; to preserve the correlation between modalities before fusion, a Joint Sparse Representation (JSR) model is introduced. Let f be the feature vector of each modality. (m) The optimization objective is:
[0155]
[0156] Among them, D (m) is the dictionary matrix for the m-th modality, used to represent the basis space of features; α is the shared sparse coefficient vector, representing the common sparse structure among the multimodalities; Let represent the L2 norm squared, used to measure reconstruction error; ||·||1 represents the L1 norm, introducing sparsity constraints; λ is a regularization parameter used to balance reconstruction accuracy and sparsity. By performing joint sparse representation on multimodal features, not only is feature completion between modalities achieved, but cross-modal consistency features are also highlighted (such as the presence of bird signs in both sound and image). After temporal alignment, noise suppression, normalization, and sparse representation, a temporally consistent, noise-suppressed, and modal-fused multimodal dataset is finally constructed.
[0157]
[0158] Among them, f t Let f1, f2, ..., f be the fused feature vector at time t; d represents the dimension of the feature vector; f1, f2, ..., f T This represents the complete feature sequence at all times 1≤t≤T; this dataset can be directly input into the subsequent target recognition and driving decision module to achieve closed-loop operation.
[0159] S3: Target Recognition: Bird detection is performed on the processed image data using an improved YOLOv7 network, outputting candidate bounding boxes for bird species, number, and location, and enhancing detection accuracy by combining acoustic and environmental features; overlapping boxes are removed by non-maximum suppression (NMS) to obtain high-confidence bird target information, providing a basis for individual tracking for behavior analysis.
[0160] S4: Behavioral Analysis: Based on a multimodal temporal modeling framework, this study infers bird activity patterns. A Transformer temporal network is used to fuse image trajectories, acoustic features, and environmental data to model bird behavior categories (e.g., foraging, resting, nesting, and flying away) and their willingness to stay. A deterrence demand function is combined to quantify bird population density, nesting tendency, and return probability, providing a basis for triggering deterrence strategies.
[0161] Target identification and behavior analysis specifically include:
[0162] After data preprocessing, the system needs to accurately identify bird targets and infer their behavioral patterns to provide a basis for triggering and selecting multimodal deterrence strategies. This section focuses on high-level semantic modeling of multi-source sensory data, i.e., identifying "where the birds are" and "what the birds are doing." Therefore, this module consists of two core subtasks: target detection and behavior recognition. Based on this, a deterrence demand function is introduced to map behavior to decisions. The target detection part uses an improved YOLOv7 framework, aiming to achieve a balance between real-time performance and accuracy. After data preprocessing, the input image has been enhanced and denoised, and its tensor representation is denoted as:
[0163]
[0164] Where H, W, and C represent the height, width, and number of channels of the image, respectively; the YOLOv7 network processes the input image through three stages: feature extraction (Backbone), feature fusion (Neck), and detection / prediction (Head), ultimately outputting a set of bird candidate targets:
[0165]
[0166] Among them, b i =(x i ,y i ,w i ,h i () represents the center coordinates and width and height of the i-th candidate box. For the target category label (such as pigeon, sparrow, magpie, etc.), s i ∈[0,1] represents the confidence score, indicating the reliability of the detection result; x i y represents the horizontal position of the candidate box center point; i The vertical position of the candidate box center point; w i h is the width of the candidate box. i The height of the candidate box; To represent the set of categories during model training; in the output, the targeting score is first calculated for each candidate box:
[0167]
[0168] Wherein, p(c i |X img ) represents the classification probability, and IoU represents the predicted bounding box b. i With real frame The intersection-over-union ratio (IoU) is calculated; finally, non-maximum suppression (NMS) is used to remove overlapping boxes, resulting in the final set of detection results. This step not only provides information on the number and location distribution of birds, but also lays the foundation for individual tracking in subsequent behavioral modeling. After detecting bird targets, it is necessary to further identify their behavioral patterns. Since bird behavior is often closely related to temporal dynamics, acoustic signals, and environmental conditions, this application constructs a multimodal temporal modeling framework. Within the time window [tk,t], the fused features are represented as follows:
[0169]
[0170] in Represents image features (such as trajectory and pose changes) in consecutive frames. Indicates acoustic characteristics (such as call frequency, duration, and intensity). Indicates environmental characteristics (such as temperature, humidity, and air pressure).
[0171] These features are modeled using a lightweight temporal network, the Transformer Encoder:
[0172]
[0173] in, h represents a time series modeling function with parameter θ. t This is the latent vector representation at time t; next, a classification layer is used to output the category distribution of bird behavior:
[0174] P(a|F t =Softmax(W·h) t +b) (15)
[0175] in This represents the set of candidate behavior categories, such as foraging, resting, nesting, and flying away. By analyzing P(a|F... t Maximize the decoding to obtain the bird's behavior judgment at the current moment:
[0176]
[0177] After completing target detection and behavior recognition, it is necessary to establish a connection between the recognition results and the expulsion strategy; a expulsion demand function D is proposed to quantify whether and how to implement expulsion.
[0178] D=αN+βR nest +γR return (17)
[0179] in, R represents the number of birds currently identified. nest∈[0,1] represents the strength of nest-building tendency inferred from behavior recognition, R return ∈[0,1] represents the probability of birds returning, reflecting the group's willingness to stay. α, β, and γ are weighting parameters used to balance the contributions of different factors to the need for dispersal. The triggering condition for the dispersal strategy is defined as follows:
[0180] D>θ (18)
[0181] The threshold θ can be dynamically adjusted according to the actual scenario (such as farmland, airport, power transmission line).
[0182] S5: Drive-off Strategy Optimization: The results of multimodal perception, target recognition, and behavior analysis are abstracted into a Markov Decision Process (MDP), defining the state space, action space, state transition function, and immediate reward function; the optimal drive-off strategy is solved through dynamic programming or reinforcement learning algorithms, and the dynamic selection and intensity adjustment of various drive-off methods such as light, sound, and vibration are carried out.
[0183] The optimization of the expulsion strategy specifically includes:
[0184] In the problem of repelling strategies, the core idea of this invention is to abstract the dynamic process of "multimodal perception—bird identification—behavior analysis—repelling decision" into a Markov Decision Process (MDP), so as to utilize mature sequential decision theory to solve for the globally optimal strategy. Specifically, the MDP can be formalized as a quintuple:
[0185]
[0186] Among them, the state space It represents a complete description of the system's "perception-environment-habit" at any given moment; the high-dimensional feature vector formed after the fusion of multimodal sensors:
[0187] s t =[species_vec t ,count_vec t pos_grid t behavior t ,env t ,conf t ,h t ], (20)
[0188] Among them, species_vec t The probability distribution of bird species, count_vec t For quantity estimation, pos_grid t For spatial location grid representation, behavior tFor discretized behavior patterns, env t Characterizing environmental factors, conf t h represents the sensor confidence level. t This depicts the habits of bird groups or the degree of domestication to stimuli;
[0189] Corresponding motion space Defined as the control combination of the driving device under different modes, such as light intensity, acoustic frequency and sound level, vibration amplitude and duration, etc., it can be abstracted as a motion vector:
[0190]
[0191] in, For optical driving sub-actions, it is a structured sub-vector; To drive the child's movements with hearing; The action is a tactile / mechanical vibratory sub-action; the state transition function P(s′|s,a) characterizes the system dynamics, that is, in state s... t Apply action a t Then transition to the next state s t+1 The probability distribution of the eviction process. The immediate reward function R(s,a) defines the eviction objective, such as improving the eviction success rate, reducing energy consumption and environmental disturbance, and ensuring safety and compliance, and can be formalized as:
[0192] R(s,a)=w eff ·f disp (s,a)-w energy ·E(a)-w disturb ·D(a)-w risk ·Ω(s,a)(22)
[0193] Among them, f disp (s,a) represents the expulsion effectiveness, measuring the positive effect of action a on the expulsion target in the current state s; E(a) represents the energy consumption, the energy / power cost consumed by action a during its implementation, reflecting resource consumption and range impact; D(a) represents the environmental disturbance index, measuring the disturbance or adverse impact of action a on non-target objects; Ω(s,a) represents the risk / compliance, measuring the potential risks or non-compliance consequences caused by action a in state s; weight w eff w energy w disturb w risk The positive scalar weights represent the relative importance of different expulsion effects, energy consumption, disturbances, and risks in the overall reward function; finally, the discount factor γ∈(0,1] is used to balance short-term expulsion effects with long-term ecological and safety constraints.
[0194]
[0195] Where V π (s t Let π be the expected cumulative reward under policy π. Through this formalization, the originally complex multimodal expulsion problem is unified under the MDP framework, which allows subsequent methods such as dynamic programming, value iteration, or reinforcement learning to solve for the optimal expulsion policy, achieving closed-loop optimization from "observation" to "action" and then to "feedback".
[0196] In modeling the optimization of deterrence strategies, one of the core tasks is to characterize birds in specific environmental states. t With external action a t We need to reasonably model the state transition probability P(s′|s,a). In particular, we need to focus on characterizing the dynamics of birds transitioning from a "staying" state to a "leaving" state. For this purpose, a parameterized probability model can be used. Let the probability of a bird of category i leaving at time t be expressed as:
[0197]
[0198] in, For the logical sigmoid function, x t Indicates from state s t Environmental and population characteristics extracted, θ i This represents a vector of species sensitivity parameters, where φ is the external action effect parameter, and b i This is the bias term; the model is essentially equivalent to logistic regression (Logit), which can statistically characterize the strength of the influence of state-action on the driving outcome. Based on this probability definition, the evolutionary relationship of group size can be further written; if Let represent the number of individuals of category i at time t. Then, the number of individuals at the next time t+1 is approximately:
[0199]
[0200] in, The number of newly arrived or migrated individuals can be modeled as a Poisson stochastic process; the formula reveals that the dynamic changes in population size are determined by a dual mechanism of "expulsion success rate" and "external migration".
[0201] Furthermore, if action a t If the signal is not a discrete signal, but a continuous vector with intensity or combination, then the dependence of the departure probability on the action can be extended to a more general nonlinear function, i.e.:
[0202]
[0203] Among them, f ψ(·) can be fitted by a radial basis function (RBF) network, whose parameter ψ is trained in a data-driven manner, thus enabling the model to more flexibly capture complex environment-action-response nonlinear relationships; in the optimization of the expulsion strategy, reward function design, optimality conditions, constraint handling, and behavior modeling together constitute the core of the MDP solution; the instantaneous reward function should quantify the trade-offs between expulsion effect, energy consumption, disturbance, and safety risks, and its general form can be written as:
[0204]
[0205] Where ΔN (i) t = N (i) tN (i) t+1 represents the number of individuals successfully expelled, E(a t D(a) represents the energy consumption of the action. t NT(s) measures disturbance to the surrounding environment or non-target organisms. t ,a t π represents the penalty for protecting birds or non-target species, and w is an adjustable weight. The reward function not only reflects the immediate removal effect but can also introduce a long-term penalty term to address the bird's habitual / domestication effects. Regarding policy solution, the state-value function and action-value function under policy π are defined as follows:
[0206]
[0207]
[0208] Where π represents the strategy; r t+k The immediate reward obtained at time t+k is given by the reward function R(s,a); γ∈[0,1] is the discount factor used to decay the present value of future returns; condition |s t =s represents the expectation of s t =S is the starting point, and subsequent actions are sampled according to π; s t+1 The next state that the system transitions to after action a is applied;
[0209] And it satisfies the Bellman expectation equation:
[0210]
[0211] and the optimal Bellman equation
[0212]
[0213] Where π(a|s) is the probability of choosing action a according to policy π in state s; P(s′|s,a) is the state transition probability, representing the probability distribution of transitioning to s′ after applying a in state s; V *(s) represents the maximum state value that the optimal policy can achieve among all possible policies; max a To maximize all actions; for discrete small state spaces, exact solutions are obtained through value iteration or policy iteration; if the state / action space is large or continuous, a parameterized Q-network approximation method is used.
[0214]
[0215] Where θ is the parameter vector; α > 0 is the learning rate, controlling the step size for parameter updates; max a′ Q(s′,a′;θ) represents the maximum action value of the next state s′ under the estimation. The gradient of the parameters indicates how to adjust the parameters to reduce the estimation error of the current state-action pair; furthermore, to characterize the habitual / domestication effects of birds on stimuli, habitual states are introduced. And correct state transitions and rewards:
[0216]
[0217] Among them, exposure (i) (a t κ is the exposure function, representing the stimulus exposure caused by the action being against the i-th type of target; κ≥0 is the habit growth coefficient, measuring the increase in habit due to each unit of exposure; δ≥0 is the recovery rate coefficient, representing the extent to which habit slowly decays over time under no-exposure conditions; τ is the time step or time interval; δ·τ represents the amount of natural recovery within that time step; the immediate reward multiplies the disengagement efficacy by a decay factor. Encourage diversified strategies to maintain long-term effectiveness;
[0218] Finally, parameter estimation is achieved by fitting the state transition model P(s′|s,a) with cross-entropy:
[0219]
[0220] And update the transition probability using an incremental statistical method:
[0221]
[0222] in, The overall loss function guides the update of parameters θ and φ; θ is the set of parameters of the policy network, which determines the action selection probability distribution π. θ The specific form of (a|s); φ is the parameter of the auxiliary model or discriminant network, used to model the probability distribution of bird responses (such as whether to leave); Let t be the actual departure label of the i-th individual; For the i-th bird predicted by the model in state st Next, execute action a. t The probability of leaving later; This is a typical cross-entropy loss; reg is a regularization term that constrains the model parameters. The state transition probability is estimated empirically; N(s,a,s′) is the statistical count, representing the number of times "taking action a in state s and reaching s′" has been observed in historical or simulation samples; N(s,a) is the statistical count, representing the total number of times action a has been taken from state s in historical or simulation samples; the expulsion and early warning mechanism is set based on a multi-level response strategy, divided into first-level response, second-level response, and third-level response, which are defined as follows:
[0223] Level 1 Response: If the number of birds or their behavior patterns indicate slight signs of roosting, and the demand for driving them away is slightly above the threshold, only low-intensity, low-interference light / sound / vibration stimuli will be initiated to alert operators.
[0224] Level 2 Response: When bird density increases or nesting tendency is obvious, and the demand function for driving away birds is close to the high threshold, moderate-intensity multimodal driving away measures are initiated, and on-site inspections or adjustments to the driving away strategy are recommended.
[0225] Level 3 Response: When birds are densely packed and stay for extended periods, potentially threatening facility safety, and the demand for bird removal exceeds the maximum threshold, a high-intensity combined bird removal operation is immediately activated. Key areas are closely monitored, and an alarm is generated to prompt operations personnel to take emergency measures.
[0226] A multimodal sample dataset was constructed by collecting data from simulated environments and real-world scenarios, including image, acoustic, radar, and environmental sensor data, forming labeled data on bird behavior and flock distribution. The target recognition, behavior analysis, and deterrence strategy optimization modules were trained using the PyTorch deep learning framework. Hyperparameter tuning, cross-validation, and reinforcement learning strategy iteration were combined to improve the system's recognition accuracy and strategy effectiveness.
[0227] The reliability, accuracy, and eco-friendliness of the device under different environmental conditions were verified through simulation tests and actual field tests. Based on feedback, hardware parameters and algorithm strategies were optimized to enable the driving device to achieve closed-loop intelligent operation from multimodal perception and behavior determination to dynamic driving decision-making.
[0228] S6: Status Monitoring and Feedback: Real-time tracking of bird deterrence effects and behavioral changes, using feedback data for algorithm model updates and policy iterations. A multimodal training sample set is formed through simulation and field data collection, supporting incremental training and reinforcement learning optimization to achieve closed-loop intelligent control.
[0229] S7: Data transmission and synchronization: Transmit multimodal sensor data, target recognition results and drive-away strategy execution information to the central server through the wireless communication module to achieve integrity verification and time synchronization, which facilitates unified management across devices and regions;
[0230] S8: Visualization; Real-time display of bird activity distribution, behavior patterns, and the execution of deterrence strategies on the monitoring platform; and display of bird density maps, behavior category change maps, and deterrence response status through a graphical interface for quick assessment of regional risks and operational decisions;
[0231] S9: Model Optimization: Incremental training and parameter tuning of target recognition, behavior analysis, and deterrence strategy models based on the latest collected data; automatic updates of YOLOv7 network and MDP strategy parameters to adapt to environmental changes and dynamic adjustments in bird behavior;
[0232] S10: The system monitors the hardware operating status and software module health, promptly alerting maintenance personnel upon detecting anomalies. It supports remote software updates, log viewing, and data management, ensuring continuous and efficient operation of the device and improving the reliability, accuracy, and eco-friendliness of the repulsion device.
[0233] This invention establishes a closed-loop operation mechanism encompassing multimodal perception, target recognition, behavior analysis, dynamic deterrence, status monitoring, and strategy optimization. This mechanism enables efficient, low-interference intelligent deterrence and continuous status monitoring of bird habitats. The invention utilizes multimodal sensors to collect bird images, acoustic, radar, and environmental information, and achieves real-time monitoring through data fusion and wireless transmission. By combining EMD, Mahalanobis distance anomaly detection, and Joint Sparse Representation (JSR) methods, noise and anomalies are effectively suppressed, improving data accuracy. Improved YOLOv7 and a lightweight Transformer are used to identify and predict bird species, numbers, locations, and behavioral patterns. MDP-based deterrence strategy optimization, combined with multi-level response and habit modeling, dynamically adjusts light, sound, and vibration deterrence methods to achieve efficient, low-interference deterrence. The system integrates multimodal perception, behavior analysis, and strategy optimization in a closed loop, offering high monitoring accuracy, excellent deterrence efficiency, and strong robustness, providing innovative, efficient, and eco-friendly bird control solutions for farmland, airports, and power transmission lines.
Claims
1. A bird habitat multi-modal deterrence and condition monitoring device, characterized by: The device comprises a multi-modal perception unit, a driving execution component, a central processing unit, and a wireless communication module. The multi-modal perception unit is configured to collect bird approaching features, protective structure operation states, and environmental lighting conditions in real time, and provide basic data for driving strategies and state monitoring. The unit comprises an infrared pyroelectric sensor, a millimeter wave radar, a photosensitive device, and a vibration and attitude sensor. The infrared pyroelectric sensor is configured to determine whether a target enters a monitoring area and output a trigger signal. The millimeter wave radar is configured to measure the distance, speed, and motion trajectory between a bird and the device. The photosensitive device is configured to measure environmental brightness in real time, and assist in determining a bird activity time period and triggering an optical driving strategy. The vibration and attitude sensor is configured to detect acceleration and angular velocity changes of the protective structure, and analyze whether the device is affected by external forces. The driving execution component is configured to drive birds away through multi-channel interference according to real-time control instructions from the central processing unit.
2. The bird habitat multi-modal deterrence and status monitoring device of claim 1, wherein: The component comprises a mechanical vibration device, a directional ultrasonic emitter, and a high-intensity LED flashing unit. The mechanical vibration device generates a vibration signal through a direct current eccentric vibration motor. The directional ultrasonic emitter is configured to emit ultrasonic waves that can be perceived by birds, and interfere with the hearing of birds through intermittent, frequency-variable, or continuous wave modes. The high-intensity LED flashing unit is configured to emit short-time, high-intensity, and intermittent flashes. The driving execution component dynamically adjusts output parameters according to driving effects. The central processing unit is configured to receive data collected by multi-modal sensors, and perform preprocessing, fusion analysis, and driving strategy optimization. The wireless communication module automatically selects the optimal transmission mode according to the environment, reports real-time data, and issues remote control instructions. The central processing unit comprises a data preprocessing module, a target recognition module, a behavior analysis module, a driving strategy optimization module, and a state monitoring and feedback module. The data preprocessing module is configured to synchronize, denoise, and normalize multi-source raw data from images, acoustics, radars, and environmental sensors. The target recognition module takes processed multi-modal data as input, and accurately determines the species, quantity, and position of birds. An improved YOLOv7 framework is used for image target detection, combined with Backbone feature extraction, Neck feature fusion, and Head prediction stage output candidate target boxes. Non-maximum suppression (NMS) is used to remove overlapping boxes, obtaining high-confidence detection results and providing a basis for individual tracking for behavior analysis. The behavior analysis module infers bird activity patterns based on multi-modal time series modeling, quantifies their stay willingness, foraging, nesting, and flying away behaviors, and outputs behavior category distributions. A Transformer Encoder network is used to model fused features, output behavior category distributions, and obtain current time behavior determination results using maximum decoding. The driving strategy optimization module is used for abstracting a dynamic system formed by multi-modal perception, target identification and behavior analysis into a Markov decision process (MDP) to solve a globally optimal driving strategy; a state space contains information of bird species distribution, quantity estimation, spatial position, behavior mode, environmental condition and sensor confidence; a parameterized probability model is used for describing bird state transition rules and can be extended into a nonlinear continuous action dependent function to adapt to complex environments; finally, a value iteration, policy iteration or reinforcement learning algorithm is combined to solve an optimal strategy, realize closed-loop optimization from observation, behavior judgment to action execution, and realize dynamic driving and early warning management through a multi-level response strategy; The state monitoring and feedback module continuously tracks driving effects and bird activity changes, and feeds back data to the algorithm model for parameter updating and strategy iteration; simulation and field data collection are used to form a training sample set, including image, acoustic, radar and environmental data, and a deep learning framework is used for incremental training and reinforcement learning optimization.
3. The method of multi-modal bird deterring and condition monitoring of bird habitat implemented by the apparatus of claim 1, wherein: The method comprises the following steps: S1: data acquisition: real-time monitoring of bird activity in a target area by using the multi-modal sensor module, including image, acoustic, millimeter wave radar and environmental parameters; sensor data is transmitted to a local edge computing device through a multi-channel interface; S2: data preprocessing: denoising, calibration and time sequence alignment processing are performed on the collected original multi-modal data; environmental noise and sudden abnormalities are removed by an empirical mode decomposition (EMD) + Mahalanobis distance method, and Z-score and Min-Max normalization are used to eliminate the dimensional differences between modalities; and a joint sparse representation (JSR) is introduced to strengthen the cross-modal consistency feature, to generate a multi-modal data set with time sequence consistency, noise suppression and modal fusion; S3: target identification: using an improved YOLOv7 network to detect birds in the processed image data, outputting bird species, quantity and position candidate boxes, and combining acoustic and environmental features to enhance detection accuracy; overlapping boxes are removed by non-maximum suppression to obtain high-confidence bird target information, providing a basis for individual tracking for behavior analysis; S4: behavior analysis: based on a multi-modal time sequence modeling framework, the bird activity mode is inferred; a Transformer time sequence network is used to fuse and model image trajectories, acoustic features and environmental data, and output bird behavior categories and stay willingness; the bird group density, nesting tendency and return probability are quantified by combining a driving demand function to provide decision basis for driving strategy triggering; S5: driving strategy optimization: the multi-modal perception, target identification and behavior analysis results are abstracted into a Markov decision process (MDP), and a state space, an action space, a state transition function and an immediate reward function are defined; an optimal driving strategy is solved by a dynamic programming or reinforcement learning algorithm, and dynamic selection and intensity adjustment of light, sound and vibration driving methods are performed; S6: state monitoring and feedback: real-time tracking of bird driving effects and behavior changes, and feedback data are used for algorithm model updating and strategy iteration; multi-modal training sample sets are formed by simulation and field data collection, supporting incremental training and reinforcement learning optimization; S7: Data transmission and synchronization: transmit multi-modal sensor data, target recognition results and driving strategy execution information to the central server through the wireless communication module, implement integrity verification and time synchronization, and facilitate unified management across devices and regions; S8: Visual display: real-time presentation of bird activity distribution, behavior patterns and driving strategy execution on the monitoring platform; and display of bird density map, behavior category change graph and driving response state through graphical interface for quick judgment of regional risks and operation decisions; S9: Model optimization: incremental training and parameter tuning of target recognition, behavior analysis and driving strategy models based on the latest collected data; automatic updating of YOLOv7 network and MDP strategy parameters to adapt to environmental changes and dynamic adjustment of bird behavior; S10: System management and maintenance: monitor the running state of hardware and the health of software modules, and remind maintenance personnel in time when abnormalities are found; support remote software update, log viewing and data management to ensure continuous and efficient operation of the device.
4. The bird habitat multi-modal deterrence and status monitoring method of claim 3, wherein: Step S2 data preprocessing specifically includes: Because of the difference of sampling frequency between different modalities, the multi-source data is mapped to a unified time axis first. Let the real time series be T = {t1, t2,..., t N}, and the data of a certain modality be: where M is the number of modalities, i.e., the total number of different sensors; m represents the index of the modality; X (m) represents the time series observation data set of the mth modality; represents the observation value of the mth modality within its time range; N m represents the actual number of sample points collected by the mth modality; a uniform time step Δt is constructed by combining linear interpolation and spline interpolation: wherein, represents the estimated value of the mth modality at the unified time point t obtained by the interpolation method; x (m) (t k ) represents the observed value of the modality at the original sampling time point t k ; w k is the weight corresponding to the kth interpolation node, which determines the influence degree of different sampling points on the target time t; t represents the unified time point to be estimated; t j represents the jth sampling time in the original sampling sequence; represents that for all sampling points not equal to k, the difference (t-t j ) is multiplied item by item; is a normalization factor, which is used to ensure that when t=t k , the weight w k =1 and the rest of the weights are zero, so as to meet the interpolation consistency condition; j and k are both index variables, k specifies the current reference sampling point, and j represents all other sampling points except this point; For environmental noise and sudden abnormalities, the method of empirical mode decomposition EMD + adaptive threshold detection is introduced; EMD decomposition is performed on the input signal s(t): where IMF i (t) represents the i-th intrinsic mode function, which corresponds to the local characteristics of the signal at different time scales; K represents the number of IMF components extracted after EMD decomposition; r(t) is the residual term; the low-frequency component is reserved to represent the bird behavior trend, and the high-frequency noise is suppressed; at the same time, the Mahalanobis distance is used for anomaly detection on the residual sequence: where D M (x) denotes the Mahalanobis distance of sample x from the center of the population distribution;∑ -1 is the inverse of the covariance matrix used for the weighted distance calculation; T denotes vector transposition; if the data point is considered as an outlier and is replaced by interpolation; To eliminate the dimensional differences between different modalities, Z-score standardization is used: Where μ and σ are the sample mean and standard deviation, respectively; for sparse events, further Min-Max normalization is introduced to enhance the characteristics of rare behavior patterns; where x min and x max respectively represent the minimum and maximum values of the modal data; x' is the normalized value; in order to preserve the correlation between the modes before fusion, a joint sparse representation model JSR is introduced; let the feature vector of each mode be f (m) The optimization objective is: where Dm (m) is the dictionary matrix of the m-th modality, which is used to represent the basis space of features; and a is the shared sparse coefficient vector, which represents the common sparse structure among multi-modalities; denotes the square of two-norm, which is used to measure the reconstruction error; ‖·‖1 denotes one-norm, which introduces the sparsity constraint; and λ is the regularization parameter, which is used to balance the reconstruction accuracy and sparsity. After time alignment, noise suppression, normalization and sparse representation, a multi-modal data set X is finally constructed, which is time consistent, noise suppressed and modal fused fusion : where f t is the fusion feature vector at time t; d denotes the dimension of the feature vector; f1,f2,...,f T denotes the complete feature sequence at all time 1≤t≤T; this dataset can be directly input into the subsequent target recognition and evading decision module.
5. The bird habitat multi-modal deterrence and status monitoring method of claim 4, wherein: Target recognition and behavior analysis in steps S3 and S4 specifically include: The improved YOLOv7 framework is used for target detection, balancing real-time performance and accuracy; after data preprocessing, the input image has been enhanced and denoised, and its tensor representation is denoted as: Where H, W and C are the height, width and number of channels of the image, respectively; the YOLOv7 network processes the input image through three stages of feature extraction Backbone, feature fusion Neck and detection prediction Head, and finally outputs a set of bird candidate targets: where b i = (x i , y i , w i , h i ) represents the center coordinates and width-height of the i-th candidate box, is the target class label, s i ∈ [0, 1] is the confidence score, indicating the credibility of the detection result; x i is the horizontal position of the center point of the candidate box; y i is the vertical position of the center point of the candidate box; w i is the width of the candidate box; h i is the height of the candidate box; is the class set representing the model training; In the output result, first calculate the target score of each candidate box: where p(c i |X img ) is the classification probability, and IoU is the intersection over union of the predicted box b i and the ground truth box . Finally, non-maximum suppression (NMS) is used to remove overlapping boxes to obtain the final detection result set After detecting the bird target, further identification of its behavior pattern is needed; set the fused feature representation in the time window [t-k, t] as: wherein, represents an image feature of a consecutive frame, represents an acoustic feature, represents an environmental feature; These features are modeled through a lightweight temporal network Transformer Encoder: where, denotes a time-series modeling function with parameter θ, h t is the hidden vector representation at time t; then, a classification layer is adopted to output the class distribution of the bird behavior: P(a|F t ) = Softmax(W · h t + b) (15) wherein, represents a set of candidate behavior classes; the behavior determination result of the bird at the current time is obtained by decoding the maximization of P(a|F t ). After completing target detection and behavior recognition, the recognition results need to be linked to the driving strategy; a driving demand function D is proposed to quantify whether driving is needed and how to implement it: D = aN + βR nest + γR return (17) wherein, represents the current number of identified birds, R nest represents the nesting tendency strength inferred from the behavior recognition, R return represents the bird return probability, reflecting the group's residence intention, and a, b, g are weight parameters for balancing the contribution of different factors to the driving demand; the driving strategy trigger condition is defined as: D> θ (18) Where the threshold θ can be dynamically adjusted according to the actual scene.
6. The bird habitat multi-modal deterrence and status monitoring method of claim 5, wherein: Step S5 driving strategy optimization specifically includes: MDP can be formalized as a five-tuple: where the state space characterizes the complete description of the "perception-environment-habits" of the system at any instant; the high-dimensional feature vector formed after the fusion of multi-modal sensors: s t = [species_vec t , count_vec t , pos_grid t , behavior t , env t , conf t , h t ], (20) where species_vec t is the bird species distribution probability, count_vec t is the number estimate, pos_grid t is the spatial position grid representation, behavior t is the discretized behavior pattern, env t characterizes the environmental factors, conf t is the sensor confidence, h t then characterizes the habit or degree of domestication of the bird population to the stimulus; Corresponding action space defined as a control combination of the repelling device in different modalities, which can be abstracted as an action vector a t : wherein, is an optical chaser action, a structured sub-vector; is an auditory chaser action; is a tactile / mechanical vibration sub-action; The state transition function P(s' | s, a) then characterizes the system dynamics, i.e. the probability distribution of transitioning to the next state s t after applying the action a t t+1 ; the immediate reward function R(s, a) defines the driving objective, formalized as: R(s, a) = w eff • f disp (s, a) - w energy • E(a) - w disturb • D(a) - w risk • Ω(s, a) (22) where f disp (s, a) is the repelling effectiveness term, measuring the positive effect of action a on the repelling target in the current state s; E(a) is the energy consumption term, the energy / power cost consumed when action a is implemented, reflecting resource consumption and endurance impact; D(a) is the environmental disturbance indicator, measuring the disturbance or adverse impact caused by action a on non-target objects; Ω(s, a) is the risk / compliance term, measuring the potential risk or violation consequences triggered by action a in state s; weights w eff , w energy , w disturb , w risk are positive scalar weights, representing the relative importance of different repelling effects, energy consumption, disturbance, and risk in the overall reward function; finally, the discount factor γ ∈ (0, 1] is used to balance the short-term repelling effect and long-term ecological and safety constraints: where V π (s t ) is the expected cumulative return under policy π; in modeling the policy optimization, one of the core tasks is to model the state transition probability P(s′|s, a) of the bird under the specific environment state s t and external action a t ; Let the probability of a bird individual of category i leaving at time t be denoted as: where, is the logistic sigmoid function, x t denotes the environment and population features extracted from state s t denotes the environment and population features extracted from state s i denotes the species sensitivity parameter vector, φ is the external action effect parameter, b i is the bias term; Based on the probability definition, the evolution relationship of population size can be further written as: Let N; (t) represent the number of individuals of class i at time t, then the number of individuals of class i at the next time t + 1 is approximately: where, represents the number of newly arrived or immigrated individuals, which can be modeled as a Poisson stochastic process; this formula reveals that the dynamic change of population size is determined by the double mechanisms of "driving success rate" and "external immigration"; Further, if action a t and not a discrete signal, but a continuous vector form with intensity or composition, then the dependence of the exit probability on the action can be extended to a more general non-linear function, i.e.: where f ψ (·) can be fitted by a Radial Basis Function (RBF) network, whose parameters ψ are trained in a data-driven fashion, thus making the model more flexible to capture complex environment-action-reaction nonlinearities; In the optimization of the repelling strategy, the reward function design, the optimality condition, the constraint handling and the habit modeling jointly constitute the core of the MDP solution. The immediate reward function r t The trade-off between the repelling effect, energy consumption, disturbance and safety risk should be quantified, and the general form can be written as where ΔN (i) t = N (i) t - N (i) t + 1 represents the number of individuals successfully driven away, E(a t ) is the energy cost of action, D(a t ) measures the disturbance to the surrounding environment or non-target organisms, NT(s t , a t ) is the penalty for protecting birds or non-target species, and w is a tunable weight. In terms of policy solving, define the state value function V π (s) under a policy π as: π (s, a) as: where π is a policy; r t+k is the immediate reward obtained at time t + k, given by the reward function R(s, a); γ ∈ [0, 1] is a discount factor that attenuates the present value of a far-off reward; and the condition |s t = s indicates that the expectation is taken over subsequent actions sampled according to π, starting from s t = s, and s t+1 is the next state to which the system transitions after action a is applied; It satisfies the Bellman expectation equation: And the optimal Bellman equation where π(a|s) is the probability of choosing action a under policy π in state s; P(s'|s,a) is the state transition probability, representing the probability distribution of transitioning to s' after applying a in s; V * (s) is the maximum state value that can be obtained by the optimal policy among all possible policies; max a is the maximization over all actions; For a discrete small state space, accurate solution is obtained through value iteration or policy iteration; if the state / action space is large or continuous, parameterized Q-network approximation method is used: where θ is the parameter vector; α > 0 is the learning rate, controlling the parameter update step size; max a′ Q(s',a';θ) is the maximum action value under the estimate for the next state s'; is the gradient of the parameters, indicating how to adjust the parameters to reduce the estimation error of the current state-action pair; in addition, to characterize the habituation effect of birds to stimuli, a habituation state and correct the state transition and reward: where exposure (i) (a t ) is the exposure function, representing the amount of exposure to the stimulus caused by the action as the ith target; k > 0 is the habit growth coefficient, measuring the magnitude of habit increase caused by each unit of exposure; d > 0 is the recovery rate coefficient, representing the magnitude of slow decay of habit over time in the absence of exposure; t is the time step or interval; d-t represents the amount of natural recovery within that time step; Exponentially decaying incentive to drive away Encourage diversification of actions to maintain long-term effects; Finally, the parameter estimation is achieved by cross-entropy fitting the state transition model P(s'|s, a): And the transition probability is updated using the incremental statistical method: where, is the overall loss function used to guide the update of parameters θ, φ; θ is the parameter set of the policy network, which determines the action selection probability distribution π θ ; φ is the parameter of the auxiliary model or discriminative network, which is used to model the probability distribution of bird responses (e.g., whether to leave or not); is the true leaving label of the ith individual at the tth time step; is the model-predicted probability that the ith bird leaves after performing action a t in state s t ; φ is the parameter of the auxiliary model or discriminative network, which is used to model the probability distribution of bird responses (e.g., whether to leave or not); is the typical cross-entropy loss; reg is the regularization term, which constrains the model parameters; is the empirically estimated state transition probability; N(s, a, s') is the statistical count representing the number of times that "after taking action a in state s, s' is reached" is observed in the historical or simulation samples; N(s, a) is the statistical count representing the total number of times that action a is taken from state s in the historical or simulation samples; Set the driving and warning mechanism based on a multi-level response strategy, which is divided into first-level response, second-level response and third-level response, which are defined as follows: Primary response: The number of birds or behavior patterns show slight habitat signs, the demand function of driving is slightly higher than the threshold, only start low-intensity, low-interference light / sound / vibration stimulation, and prompt the operator to pay attention; Secondary response: The bird aggregation density increases or the nesting tendency is obvious, the demand function of driving is close to the high threshold, the medium-intensity multi-modal driving means is started, and the on-site patrol or adjustment of the driving strategy is suggested; Tertiary response: The bird stays densely and for a long time, or may threaten the safety of the facility, the demand function of driving exceeds the highest threshold, high-intensity combined driving is immediately enabled, key areas are monitored, and an alarm is generated to prompt the operator to take emergency measures; Through the simulation environment and actual scene data collection, a multi-modal sample data set is constructed, including image, acoustic, radar and environmental sensor data, forming labeled bird behavior and group distribution data; using the PyTorch deep learning framework, the target recognition, behavior analysis and driving strategy optimization modules are trained, combined with hyperparameter tuning, cross-validation and reinforcement learning strategy iteration, to improve the recognition accuracy and strategy effect of the system.
Citation Information
Cited By
Protection equipment management system and method based on multi-source data
CN121481189A
Bird voiceprint and vision fusion real-time identification method for oil exploitation operation area
CN121861392A
Bird voiceprint and visual fusion real-time identification method for oil exploitation operation area
CN121861392B
Power transmission line anti-bird control optimization method, device and equipment
CN122123357A
Multi-target recognition result merging method and system of three-dimensional ground penetrating radar
CN122336564A