A method and apparatus for deterring bird adaptation

CN122804764APending Publication Date: 2026-09-25SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611114149.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供一种防止鸟类产生适应性的驱鸟方法及装置,旨在解决现有驱鸟策略因可预测而导致鸟类产生适应性、驱鸟效果随时间衰减的问题

Benefits of technology

本申请实施例通过从预设的多种驱离模态组合中随机选取一种驱离模态组合,为所选中的每种驱离模态分别从其各自预设的工作参数范围内独立随机选取一组工作参数,以及随机确定各所选驱离模态的启动顺序及持续时间,使每次驱离在模态组合、工作参数和启动时序三个层面均不可预测,鸟类无法形成有效的认知关联,从而无法将驱离刺激识别为无害信号并建立适应性习惯,有效防止了鸟类产生适应性,解决了驱鸟效果随时间衰减的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804764A_ABST
    Figure CN122804764A_ABST
Patent Text Reader

Abstract

The application relates to a bird repelling method and device for preventing birds from adapting, which comprises the following steps: monitoring a bird target in real time and acquiring the species, quantity, initial distance and azimuth angle of the bird target, and the environmental light intensity and noise intensity; randomly selecting one from a plurality of repelling mode combinations, independently and randomly selecting working parameters for each repelling mode, randomly determining a starting sequence and a duration, encoding into a bird repelling instruction sequence and executing; evaluating an effectiveness level after repelling ends, and adjusting the weight of the repelling mode combination used in real time; generating a reward value according to the effectiveness level, combining state parameters, action parameters and the reward value into a training sample and storing the training sample into a strategy effect database; and updating the Q value of each action parameter in each state by using a Q-learning algorithm every preset period, and updating the weight of each repelling mode combination according to the Q value. The application makes each repelling unpredictable, effectively prevents birds from adapting, and realizes adaptive optimization of a repelling strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent bird deterrence technology, specifically to a bird deterrence method and device for preventing birds from developing adaptive behaviors. Background Technology

[0002] With the development of power transmission, airport operations, and agricultural planting, the safety hazards and economic losses caused by bird activity are becoming increasingly prominent. Existing bird control technologies mainly fall into two categories: one uses a single deterrent method, such as using an ultrasonic bird deterrent or a fixed-frequency laser bird deterrent alone; the other uses a fixed combination of multiple deterrent methods, such as activating ultrasonic waves and lights simultaneously in a fixed sequence. These existing bird control solutions share a fundamental flaw: the deterrent patterns are fixed or predictable, making it easy for birds to adapt.

[0003] Specifically, whether using a single deterrent method or a fixed combination of multiple deterrent methods, the deterrent strategy is essentially static or repetitive. The modal combination activated each time is fixed, the operating parameters of each modality remain unchanged, and the activation sequence also lacks variation. After prolonged exposure to this fixed deterrent stimulus, birds gradually develop a cognitive habit, recognizing the deterrent stimulus as a harmless signal and thus ceasing to exhibit avoidance behavior. This adaptation leads to a continuous decline in bird deterrent effectiveness over time, eventually resulting in complete failure. Once birds have adapted to the fixed modal combination, even if the equipment is still functioning normally, it cannot effectively deter birds.

[0004] In addition, although some existing devices have simple random functions, such as randomly playing different audio, the random changes in audio types are only changes within a single modality. They lack a mechanism for randomly allocating multiple dispersal methods at a combined level. At the same time, the working parameters and start-up sequence are relatively fixed. Birds can still identify the boundaries and patterns of random changes after repeated exposures, which cannot fundamentally break the learning and adaptation process of birds. Summary of the Invention

[0005] This application provides a bird-repelling method and apparatus to prevent birds from developing adaptations, aiming to solve the problems that existing bird-repelling strategies lead to bird adaptations due to their predictability and that the bird-repelling effect diminishes over time.

[0006] According to a first aspect, embodiments of this application provide a bird-repelling method to prevent birds from developing adaptive behaviors, comprising the following steps:

[0007] The system monitors in real time whether there are bird targets in the protected area. If so, it triggers a bird deterrence event and obtains the type, number, initial distance and orientation angle of the bird target, as well as the current ambient light intensity and noise intensity. In response to the bird deterrence event, a deterrence mode combination is randomly selected from a preset combination of multiple deterrence modes. The multiple deterrence modes combination consists of at least one deterrence mode among laser deterrence mode, ultrasonic deterrence mode, sound deterrence mode and red and blue warning light deterrence mode. The probability of each deterrence mode combination being selected is determined by the weight of each deterrence mode combination. For each of the selected de-icing mode combinations, a set of working parameters is randomly selected independently from its own preset working parameter range, and the start order and duration of each selected de-icing mode are randomly determined; the selected de-icing mode combination, the working parameters of each de-icing mode, the start order and duration are encoded into a bird deterrence command sequence, and each de-icing unit is controlled to perform de-icing actions according to the bird deterrence command sequence. After the bird removal action is completed, the dynamic information of the bird target continues to be monitored in real time. Based on the dynamic information of the bird target, it is determined whether the bird target has left the protected area. The effectiveness level of this removal is evaluated based on the determination result, and the weight of the removal mode combination used in this removal is adjusted in real time according to the effectiveness level of this removal. The effectiveness level includes successful removal, partial removal, and ineffective removal. When it is determined that the bird target has left the protected area, the departure time, departure direction, and departure distance of the bird target are also recorded. According to the effectiveness level, a reward value is generated according to the preset reward rules. The historical success rate of the light intensity, noise intensity, type, quantity, initial distance, and the driving mode combination used in this driving is used as the state parameter. The driving mode combination used in this driving and the working parameters of each driving mode are used as the action parameter. The reward value is used as the evaluation label. The combination is stored in the strategy effect database as a training sample. Every preset period, the Q-learning algorithm is used to update the Q value of each action parameter in each state based on all training samples in the policy effect database, and the weights of each disengagement mode combination are updated based on the updated Q value; wherein, the Q value is used to characterize the cumulative expected reward that can be obtained by selecting the corresponding action parameter in the corresponding state.

[0008] In one specific implementation, the method further includes: Based on the updated weights of each eviction mode combination, perform an overfitting constraint check on each eviction mode combination. The over-adaptation constraint check includes: The number of times any expulsion mode combination can be triggered within a preset time period shall not exceed a preset limit. Once any expulsion mode combination exceeds the preset limit, it shall be prohibited from use for a preset prohibition period. For a combination of expulsion modes that is used more frequently than a preset threshold, when it is selected for expulsion, its operating parameters are randomly reselected within a preset offset range.

[0009] In one specific implementation, the over-adaptive constraint check further includes: When the effectiveness level of three consecutive expulsions is all ineffective or the overall expulsion success rate is less than 80%, the boundary value of the current working parameter range of each expulsion mode is expanded outward by 20%, so that the random selection range of each working parameter is expanded. Within the expanded range, the upper or lower limit value of the preset working parameter range of each expulsion mode is selected first. At the same time, when randomly selecting expulsion mode combinations, the expulsion mode combinations with the highest frequency of use in the first 5 times are excluded.

[0010] In one specific implementation, the real-time monitoring of whether bird targets exist in the protected area specifically includes: The radar monitors the protected area in real time. When the radar detects a target whose distance, speed, and radar cross-section meet preset distance thresholds, speed thresholds, and radar cross-section thresholds, it determines that a bird target has been detected and generates a trigger signal. In response to the trigger signal, the camera is activated to acquire images of the area where the bird target is located, and the image acquired by the camera is used to identify whether the bird target is a bird; the camera has a built-in convolutional neural network model for identifying whether the target in the image is a bird.

[0011] In one specific implementation, each of the de-escalation units includes a laser de-escalation unit, an ultrasonic de-escalation unit, an acoustic de-escalation unit, and a red and blue warning light de-escalation unit. The method further includes: During the execution of the deportation action, the deportation intensity and duration of each deportation unit are dynamically adjusted according to the dynamic information of the bird target. When the bird target approaches, the deportation intensity is increased and the deportation duration is extended. When the bird target moves away, the deportation intensity is reduced and the deportation is terminated early. The pointing direction of the laser deportation unit is also adjusted according to the position change of the bird target.

[0012] In one specific implementation, generating a reward value according to a preset reward rule based on the validity level specifically includes: A base reward value is obtained based on the effectiveness level; where a successful expulsion corresponds to the first base reward value, a partially successful expulsion corresponds to the second base reward value, and an ineffective expulsion corresponds to the third base reward value. If the energy consumption of this expulsion action is lower than the preset energy consumption threshold, then a preset first additional reward value is obtained; wherein the first additional reward value is greater than 0. If the frequency of use of the deportation mode combination used in this expulsion exceeds a preset frequency threshold within a preset time period, a preset second additional reward value is obtained; wherein the second additional reward value is less than 0. The total reward value obtained by adding the base reward value, the first additional reward value, and the second additional reward value is used as the evaluation label.

[0013] In one specific implementation, the step of randomly selecting a drive-off mode combination from a preset combination of multiple drive-off modes includes: A roulette wheel method is used to randomly select one of a variety of pre-set drive-away mode combinations. Radar noise is used as the random seed. Each drive-away mode combination is assigned a corresponding selection interval according to the probability determined by the weight of each drive-away mode combination. The radar noise is used as a random source to generate random hit results that fall into a certain interval. The drive-away mode combination corresponding to the random hit result is used as the selected drive-away mode combination.

[0014] In one specific implementation, the method further includes: Before randomly selecting a drive-off mode combination from a set of preset drive-off mode combinations, the weights of each drive-off mode are biased and adjusted according to the light intensity and the noise intensity. The weights of each drive-off mode combination are then updated based on the adjusted weights. Specifically, if the light intensity is higher than a preset light threshold, the weight of the red-blue warning light drive-off mode is reduced; if the light intensity is lower than a preset low light threshold, the weights of the red-blue warning light drive-off mode and the sound drive-off mode are increased; and if the noise intensity is higher than a preset noise threshold, the weights of the laser drive-off mode and the ultrasonic drive-off mode are increased.

[0015] In one specific implementation, the step of adjusting the weights of the deportation mode combination used in the current deportation in real time according to the effectiveness level of the current deportation includes: After each bird deterrence event, if the effectiveness level of the deterrence is successful, the weight of the deterrence mode combination used in the deterrence event increases by 0.02; if the effectiveness level of the deterrence event is ineffective, the weight of the deterrence mode combination used in the deterrence event decreases by 0.03; the weight of each deterrence mode combination is limited to between 0.1 and 0.4.

[0016] According to a second aspect, embodiments of this application provide a bird deterrent device to prevent birds from developing adaptive behaviors, including a module for performing the bird deterrent method described in the first aspect.

[0017] The bird-repelling method and apparatus according to embodiments of this application for preventing birds from developing adaptive behaviors have the following beneficial effects: This application embodiment randomly selects one of a preset combination of deterrence modes, independently and randomly selects a set of working parameters from its own preset working parameter range for each selected deterrence mode, and randomly determines the start order and duration of each selected deterrence mode. This makes each deterrence unpredictable in terms of mode combination, working parameters, and start order, preventing birds from forming effective cognitive associations and thus preventing them from recognizing the deterrence stimulus as a harmless signal and establishing adaptive habits. This effectively prevents birds from developing adaptations and solves the problem of bird deterrence effect decaying over time.

[0018] This application's embodiments adjust the weights of the bird-repelling mode combinations in real time after each repelling event based on the effectiveness level of that repelling, enabling the system to immediately adjust strategy preferences based on the effect of a single repelling and quickly respond to recent changes in effectiveness. Simultaneously, by using the Q-learning algorithm at preset intervals to update the Q-values ​​of each action parameter in each state based on all training samples in the strategy effectiveness database, and updating the weights of each repelling mode combination based on the updated Q-values, the system can continuously learn from historical repelling effects, with the repelling effectiveness improving rather than decaying over time. The real-time weight adjustment and periodic reinforcement learning optimization work together to balance short-term adaptability and long-term optimality. Furthermore, this application uses the current ambient light intensity and noise intensity as state parameters for reinforcement learning, allowing strategy optimization to incorporate environmental conditions and further improving the environmental adaptability of the bird-repelling strategy. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a bird-repelling method for preventing birds from developing adaptive behaviors, as described in one embodiment of this application.

[0021] Figure 2 This is a schematic diagram of the structure of a bird deterrent device for preventing birds from developing adaptive behavior, according to one embodiment of this application. Detailed Implementation

[0022] The detailed description of the accompanying drawings is intended to illustrate embodiments of this application and is not intended to represent only the forms in which this application can be implemented. It should be understood that the same or equivalent functionality can be achieved by different embodiments intended to be included within the spirit and scope of this application.

[0023] One embodiment of this application provides a bird deterrence method to prevent birds from developing adaptive behaviors. This method is based on a device for preventing adaptive bird behavior, which includes a detection and sensing module, a central control module, and a multimodal deterrence execution module. The detection and sensing module is used to sense bird targets and environmental information within the protected area. The central control module receives data sent by the detection and sensing module and generates a bird deterrence strategy. The multimodal deterrence execution module receives instructions sent by the central control module and executes deterrence actions. These three modules are connected sequentially to form a complete chain of data acquisition, strategy decision-making, and action execution.

[0024] See Figure 1 The method in this application embodiment includes the following steps: Step S10: Monitor in real time whether there are bird targets in the protected area. If so, trigger the generation of a bird deterrent event and obtain the type, number, initial distance and orientation angle of the bird targets, as well as the light intensity and noise intensity of the current environment. Specifically, in this embodiment, the protected area refers to the effective range that the bird-repelling device can cover, which is determined by the performance parameters of each bird-repelling unit in the device. Real-time monitoring refers to the detection and sensing module continuously collecting data at a fixed frequency. The collection frequency can be set according to actual needs, for example, collecting data 10 to 30 times per second.

[0025] The bird species referred to here refers to the biological category to which the target belongs, such as sparrow, magpie, crow, pigeon, falcon, etc. The species is determined by classifying the images captured by the camera using a convolutional neural network model. Before deployment, the convolutional neural network model has been trained using a large number of images labeled with the above bird categories. In actual operation, the model outputs the probability distribution of each category, and the category corresponding to the highest probability value is taken as the recognition result, with a recognition accuracy of no less than 92%.

[0026] The number refers to the number of bird targets that exist simultaneously within the protected area, which is determined by the number of targets in the same direction detected by radar and then verified by the camera.

[0027] The initial distance refers to the distance between the bird target and the device when the bird deterrence event is triggered, which is obtained by radar measurement.

[0028] The azimuth angle refers to the direction of the bird target on the horizontal plane with the device as the origin, which is determined by the radar scanning mechanism.

[0029] The light intensity is collected by a light sensor, which can be installed on the surface of the device housing, and its output is a digital signal or an analog signal.

[0030] The noise intensity is collected by a microphone, which can be placed on the side of the device. The audio signal output by the microphone is converted from analog to digital and then converted to sound pressure level to obtain the noise intensity value.

[0031] Step S20: In response to the bird deterrence event, randomly select one deterrence mode combination from a preset combination of multiple deterrence modes. The multiple deterrence mode combinations consist of at least one deterrence mode among laser deterrence mode, ultrasonic deterrence mode, sound deterrence mode and red and blue warning light deterrence mode. The probability of each deterrence mode combination being selected is determined by the weight of each deterrence mode combination. Specifically, in this embodiment, the repulsion mode refers to an independent repulsion method. The laser repulsion mode stimulates the bird's visual system with a laser beam to produce a repulsion effect. The ultrasonic repulsion mode stimulates the bird's auditory system with high-frequency sound waves to produce a repulsion effect. The sound repulsion mode plays predator calls or warning sounds to produce a repulsion effect. The red and blue warning light repulsion mode produces a repulsion effect with high-frequency flashing colored lights. The above four modes can be used individually or in any combination.

[0032] For example, the preset multiple drive-off mode combinations refer to all non-empty combinations consisting of at least one of the above four modes, with a total of 15 combinations.

[0033] The weight of each disengagement mode combination is calculated from the weights of each disengagement mode it contains. Specifically, the four disengagement modes each correspond to their respective modal weights: the weight of the laser disengagement mode is denoted as w_L, the weight of the ultrasonic disengagement mode as w_U, the weight of the sound disengagement mode as w_S, and the weight of the red-blue warning light disengagement mode as w_Light. The weight of any disengagement mode combination is equal to the product of the weights of each disengagement mode contained in that combination. For example, the combination weight of a combination consisting of the laser disengagement mode and the sound disengagement mode is w_L × w_S; the combination weight of a combination consisting of the laser disengagement mode, the ultrasonic disengagement mode, and the sound disengagement mode is w_L × w_U × w_S; and the combination weight of a combination consisting of all four disengagement modes is w_L × w_U × w_S × w_Light. Each modal weight is normalized so that the sum of the weights of all non-empty combinations is 1.

[0034] Each expulsion mode combination is pre-stored in the storage unit of the central control module. Each expulsion mode combination corresponds to a weight value, which determines the probability that the expulsion mode combination will be selected during random selection. The larger the weight value, the higher the probability that the expulsion mode combination will be selected.

[0035] It should be noted that when a bird deterrence event is triggered, the central control module randomly selects one of the 15 deterrence mode combinations based on the weight values ​​of the current deterrence mode combinations. This random selection process introduces unpredictability, meaning that the deterrence mode combination used for each deterrence is unrelated to previous deterrence events, and birds cannot predict the next deterrence mode combination based on historical experience.

[0036] Step S30: For each of the selected driving mode combinations, a set of working parameters is randomly selected independently from its own preset working parameter range, and the start order and duration of each selected driving mode are randomly determined; the selected driving mode combination, the working parameters of each driving mode, the start order and duration are encoded into a bird deterrence command sequence, and each driving unit is controlled to perform driving actions according to the bird deterrence command sequence; Specifically, in this embodiment, each dispersal mode is set with its own range of operating parameters.

[0037] The operating parameters of the laser expulsion mode include power, scanning path, and scanning speed. Power refers to the output power of the laser beam. The scanning path refers to the trajectory of the laser beam in the target area, including various modes such as line scanning, multi-point scanning, and random path scanning. The scanning speed refers to how fast the laser beam moves along the scanning path. The value of the scanning speed can range from 5° / s (degrees per second) to 20° / s (degrees per second). The value of the laser pulse frequency can range from 1Hz (hertz) to 20Hz (hertz).

[0038] The operating parameters of the ultrasonic expulsion mode include frequency and transmission duration. Frequency refers to the oscillation frequency of the ultrasonic signal, and transmission duration refers to the duration of continuous ultrasonic transmission. The ultrasonic expulsion mode also supports three operating modes: fixed frequency, variable frequency, and sweep frequency. The sweep frequency mode includes two methods: linear sweep frequency and logarithmic sweep frequency. The frequency can be selected from the frequency bands of 15kHz to 20kHz, 20kHz to 25kHz, 25kHz to 30kHz, or 30kHz to 35kHz. The transmission duration can range from 3s to 15s.

[0039] The operating parameters of the sound disengagement mode include audio number, volume, and number of overlay channels. The audio number refers to the index number of each audio file in the preset audio library. The volume refers to the loudness of the sound playback, and the volume value can range from 60dB to 106dB. The number of overlay channels refers to the number of audio files played at the same time, which can be selected from 1, 2, or 3 channels.

[0040] The operating parameters of the red and blue warning light drive-off mode include flashing mode, flashing frequency, and brightness. Flashing modes include alternating red and blue flashing, single-color flashing, random frequency flashing, and gradual flashing. Flashing frequency refers to the speed at which the light alternates between bright and dark, and the value range of flashing frequency can be from 1Hz (Hertz) to 20Hz (Hertz). Brightness refers to the luminous intensity of the light, and the value range of brightness can be from 60LM (lumen) to 120LM (lumen).

[0041] The operating parameters for each mode are independently and randomly selected within their respective preset ranges. For example, the laser power is randomly selected from five preset levels (100mW, 300mW, 500mW, 800mW, and 1000mW), and the scanning speed is randomly selected from a preset speed range (5° / s to 20° / s). The random selection of each parameter is independent and does not affect the others.

[0042] The startup sequence refers to the order in which the selected expulsion modes begin execution. They can start simultaneously or in a preset order. The startup sequence includes four modes: parallel startup, serial startup, interleaved startup, and delayed startup. In interleaved or delayed startup, the startup interval between modes is randomly selected within the range of 0.5 seconds to 2 seconds. The duration refers to the total time for each selected expulsion mode to perform the expulsion action. The duration of each mode can be the same or different, and the total duration of a single expulsion is randomly selected within the range of 3 seconds to 20 seconds.

[0043] The central control module encodes all the above information into a bird deterrence command sequence, which is in a format that the multimodal bird deterrence execution module can parse. After receiving the command sequence, the multimodal bird deterrence execution module controls each deterrence unit to perform the corresponding deterrence action according to the command content.

[0044] Step S40: After the expulsion action is completed, the dynamic information of the bird target continues to be monitored in real time. Based on the dynamic information of the bird target, it is determined whether the bird target has left the protected area. The effectiveness level of this expulsion is evaluated based on the determination result, and the weight of the expulsion mode combination used in this expulsion is adjusted in real time according to the effectiveness level of this expulsion. The effectiveness level includes successful expulsion, partial expulsion, and ineffective expulsion. When it is determined that the bird target has left the protected area, the departure time, departure direction, and departure distance of the bird target are also recorded. Specifically, in this embodiment, the completion of the bird deterrence action means that each deterrence unit has completed all deterrence actions according to the bird deterrence command sequence. At this time, the detection and sensing module does not stop working, but continues to collect dynamic information of the bird targets.

[0045] The central control module determines whether the bird target has left the protected area based on the collected dynamic information. The determination can be based on whether the distance detected by radar exceeds the preset boundary value of the protected area.

[0046] If the bird target has left the protected area and does not return within a certain period of time, it is assessed as successfully driven away. If the bird target briefly leaves and then returns, it is assessed as partially driven away. If the bird target never leaves the protected area, it is assessed as invalid.

[0047] When it is determined that a bird target has left the protected area, the central control module records the departure time, departure direction, and departure distance. The departure time refers to the moment when the target first crosses the boundary of the protected area, the departure direction refers to the direction of movement of the target when crossing the boundary, and the departure distance refers to the farthest distance the target reaches after leaving the boundary.

[0048] The central control module adjusts the weights of the bird deterrence modal combinations used in the current deterrence operation in real time based on the effectiveness level of the deterrence. The weight of the deterrence modal combination is increased for successful deterrence and decreased for ineffective deterrence. For partial deterrence, the weights can remain unchanged or be fine-tuned. This adjustment is executed immediately after the bird deterrence event ends, enabling the system to quickly respond to changes in the effectiveness of a single deterrence operation.

[0049] Step S50: Generate a reward value according to the effectiveness level and a preset reward rule. Use the light intensity, noise intensity, type, quantity, initial distance, and the historical success rate of the expulsion mode combination used in this expulsion as state parameters. Use the expulsion mode combination used in this expulsion and the working parameters of each expulsion mode as action parameters. Use the reward value as an evaluation label. Combine them into a training sample and store it in the strategy effect database. Specifically, in this embodiment, the reward value is a numerical value used to quantify the effectiveness of the expulsion action. Different effectiveness levels correspond to different reward values, with successful expulsions corresponding to higher reward values ​​and ineffective expulsions corresponding to lower reward values ​​or negative values.

[0050] The state parameters describe the environmental and target conditions at the time of the bird deterrence event, including light intensity, noise intensity, bird species, number, initial distance, and the historical deterrence success rate of this modality combination. The historical deterrence success rate refers to the ratio of the number of successful deterrence attempts by this deterrence modality combination in historical events to the total number of attempts.

[0051] The action parameters describe the specific strategy adopted in this expulsion, including the combination of expulsion modes used and the operating parameters of each expulsion mode. There is a one-to-one correspondence between the action parameters and the state parameters.

[0052] The central control module combines state parameters, action parameters, and reward values ​​into a training sample and stores it in the policy effect database. The policy effect database is stored in the central control module's storage unit and is used to accumulate historical data.

[0053] Step S60: Every preset period, the Q-learning algorithm is used to update the Q value of each action parameter in each state based on all training samples in the policy effect database, and the weights of each disengagement mode combination are updated based on the updated Q value; wherein, the Q value is used to characterize the cumulative expected reward that can be obtained by selecting the corresponding action parameter in the corresponding state.

[0054] Specifically, in this embodiment, the preset period refers to the time interval between two batch reinforcement learning updates, which can be set according to actual needs, such as 24 hours.

[0055] Because the state parameters include continuously changing physical quantities (such as light intensity, noise intensity, initial distance, etc.), traditional Q-learning algorithms are not applicable to the method in this embodiment. To construct a finite state-action space to support the storage and retrieval of Q values, the central control module in this embodiment discretizes the continuous state parameters. Light intensity is divided into three states: strong light, normal light, and weak light, according to a preset light threshold. Noise intensity is divided into two states: high noise and normal noise, according to a preset noise threshold. Bird species include six categories: sparrows, magpies, crows, pigeons, falcons, and others. Bird quantity is divided into two states: single bird and multiple birds. Initial distance is divided into two states: near distance and far distance, according to a preset distance threshold. Historical drive-away success rate is divided into three states: high, medium, and low, according to a preset success rate threshold. The drive-away mode combination in the action parameters consists of 15 discrete combinations. The working parameters of each mode are discretized into several levels according to a preset value range, making the action space a finite set.

[0056] Through the discretization process described above, the central control module constructs a Q-table. The Q-table is a two-dimensional table structure, where the row index is the discretized state s and the column index is the discretized action a. Each cell in the table stores the Q-value Q(s,a) corresponding to that state-action pair. All values ​​in the Q-table are set to 0 during system initialization. When it is necessary to query the Q-value of a specific state-action pair, the central control module locates the corresponding row and column in the Q-table based on the state s and action a, and reads the value from that cell; when an update is needed, the updated value is written to that cell.

[0057] When the preset period arrives, the central control module iterates through all training samples in the policy effect database, mapping the continuous state parameters in each sample to the corresponding discrete state, and mapping the continuous action parameters in the sample to the corresponding discrete action, thus obtaining discretized state-action pairs. The central control module updates the Q-value of this state-action pair according to the Q-learning algorithm, with the update formula as follows: Q(s,a) = Q(s,a) + α × [R + γ × max Q(s′,a′) Q(s,a)] Where s represents the current discrete state parameters, a represents the current discrete action parameters, and R represents the total reward value obtained in this expulsion. α is the learning rate (value is 0.1), and γ is the discount factor (value is 0.9). s′ represents the next discrete state after executing action a, constructed and determined by the central control module based on the effect of this expulsion. Specifically, if the expulsion is successful, the initial distance changes from near distance to far distance; if the expulsion is ineffective, the initial distance remains unchanged. The historical expulsion success rate is recalculated based on the result of this expulsion and updated to a new discrete level. Other state parameters (light intensity, noise intensity, bird species, bird numbers) remain consistent with s. max Q(s′,a′) represents the maximum Q value among all possible discrete actions in the next discrete state s′, obtained by querying the Q table.

[0058] In this application, each bird-repelling event is triggered independently, and there is no natural progression relationship between consecutive bird-repelling events. Therefore, s′ does not refer to the next state in the time series, but rather to the state at the end of the bird-repelling event after the current action is performed. The role of s′ is to provide the Q-learning algorithm with transition information that "the state has changed after performing action a", enabling the system to learn the effect of performing a specific action in a specific state.

[0059] After the central control module traverses all training samples, the Q-values ​​of each discrete state-action pair are updated.

[0060] The central control module updates the weights of each expulsion mode combination based on the updated Q-values. Specifically, for each discrete state s, the central control module aggregates the Q-values ​​of all discrete actions in that state according to the mode combination dimension, averages the Q-values ​​of all actions containing the same expulsion mode combination, and obtains the comprehensive Q-value of that mode combination in that state. After traversing all discrete states, the comprehensive Q-values ​​of the same mode combination in different states are averaged to obtain the overall comprehensive Q-value of that mode combination. Then, the overall comprehensive Q-values ​​of all mode combinations are normalized so that the sum of the weights of all mode combinations is 1, resulting in the updated weights of each expulsion mode combination.

[0061] For example, suppose that in a certain state s (e.g., strong light, normal noise, magpie, single, close range, high success rate), the system iterates through all actions, where: The Q values ​​for the two specific actions corresponding to the "laser + sound" combination are 3.225 and 2.100, respectively, and the average value is 2.6625. The Q value of one action corresponding to the "laser" combination is 1.800, and the average value is 1.800. The Q value of a single action corresponding to the "ultrasound" combination is 0.900, and the average value is 0.900. The Q value of an action corresponding to the "sound + light" combination is 0.500, and the average value is 0.500. Other combinations have no action records in this state, and the overall Q value is 0; The sum of the combined Q values ​​of all combinations is 2.6625 + 1.800 + 0.900 + 0.500 = 5.8625.

[0062] After normalization: The weighting for "laser + sound" is approximately 0.45: 2.6625 / 5.8625. "Laser" weight: 1.800 / 5.8625 ≈ 0.31; Weight of "ultrasound": 0.900 / 5.8625 ≈ 0.15; Weighting for "sound + light": 0.500 / 5.8625 ≈ 0.09; Other combinations: 0; Then, based on the weight boundary constraints (0.1 to 0.4), weights exceeding 0.4 are truncated to 0.4, and weights below 0.1 are increased to 0.1. After renormalization, the final weights of each mode combination under state s are obtained. The weights of this state are accumulated into the global weight table, and after traversing all discrete states, the average is taken to obtain the global weights of each mode combination.

[0063] The updated weights are used for the random selection of deterrence modal combinations in subsequent bird deterrence events; higher weights have a greater probability of being selected. By periodically executing the above process, the system's deterrence strategy can continuously learn and optimize from historical experience.

[0064] The central control module also uses a sliding window to calculate the success rate of each expulsion mode combination. The size of the sliding window can be set according to actual needs, such as calculating the success rate of each expulsion mode combination in the most recent 100 expulsion events. When the sliding window success rate of a certain expulsion mode combination is lower than 80%, that expulsion mode combination automatically enters the anti-adaptation reinforcement mode.

[0065] In some embodiments, the method further includes: Step S70: Perform an over-adaptation constraint check on each expulsion mode combination based on the updated weights of each expulsion mode combination. The over-adaptation constraint check includes: The number of times any expulsion mode combination can be triggered within a preset time period shall not exceed a preset limit. Once any expulsion mode combination exceeds the preset limit, it shall be prohibited from use for a preset prohibition period. For a combination of expulsion modes that is used more frequently than a preset threshold, when it is selected for expulsion, its operating parameters are randomly reselected within a preset offset range.

[0066] Specifically, in this embodiment, the over-adaptation constraint check is performed after each weight update. Its purpose is to prevent birds from developing an adaptation to a certain disengagement mode combination due to frequent use.

[0067] The preset time period refers to the time window for counting trigger counts, which can be set to 24 hours (h). The preset maximum number of triggers refers to the maximum number of triggers allowed within this time window, which can be set to 20% of the total number of triggers within this time period. The central control module records the number of triggers for each expulsion mode combination within the preset time period. When the number of triggers for a certain expulsion mode combination reaches the preset maximum number of triggers, that expulsion mode combination is prohibited from participating in random selection for a preset prohibition period. The preset prohibition period can be set to 2 hours (h).

[0068] Applying parameter offsets to drive mode combinations whose usage frequency exceeds a preset threshold means that when such a drive mode combination is selected, its operating parameters are randomly reselected within a preset offset range. For example, the preset offset range for laser power can be ±10%, meaning that if the power of the laser drive mode in the drive mode combination was originally 500mW, it will be randomly reselected within the range of 450mW to 550mW. Parameter offsets cause the same combination to exhibit differentiated drive stimuli when selected at different times, enhancing the unpredictability of the drive stimuli.

[0069] In some embodiments, the over-adaptation constraint check further includes: When the effectiveness level of three consecutive expulsions is all ineffective or the overall expulsion success rate is less than 80%, the boundary value of the current working parameter range of each expulsion mode is expanded outward by 20%, so that the random selection range of each working parameter is expanded. Within the expanded range, the upper or lower limit value of the preset working parameter range of each expulsion mode is selected first. At the same time, when randomly selecting expulsion mode combinations, the expulsion mode combinations with the highest frequency of use in the first 5 times are excluded.

[0070] Specifically, in this embodiment, three consecutive invalid attempts or an overall success rate below 80% indicate that the current strategy set is completely ineffective, requiring a forced adjustment of the strategy space. The aforementioned triggering condition is related to the sliding window statistics in step S60. When the sliding window statistics result meets this condition, the system automatically enters the anti-adaptive reinforcement mode.

[0071] Extending the boundary values ​​of the operating parameter range by 20% means increasing both the upper and lower limits of the current operating parameter range for each disengagement mode by 20%. For example, if the current operating parameter range for a mode is A to B, the extended range is 0.8A to 1.2B. The extended range serves as the valid range for subsequent random selection of operating parameters for that mode.

[0072] Prioritizing the use of the upper or lower limit of the preset operating parameter range within an expanded range means that when randomly selecting operating parameters, the system will choose the upper or lower limit of the expanded range with a higher probability. In other words, it will prioritize trying the highest or lowest intensity parameter settings to explore more extreme eviction stimuli. For example, the laser eviction mode will prioritize the upper limit of its preset operating parameter range, 1000mW, and the ultrasonic eviction mode will prioritize the upper limit of its preset operating parameter range, 35kHz.

[0073] Meanwhile, when randomly selecting expulsion mode combinations, the central control module excludes the top 5 most frequently used expulsion mode combinations. This means these 5 combinations are not included in the random selection under the current reinforcement mode, and the system selects from the remaining combinations. This combination of measures enables the system to quickly escape the current failure strategy space and explore new, effective strategies.

[0074] In some embodiments, step S10, which involves real-time monitoring of whether bird targets exist in the protected area, specifically includes: Step S101: The protected area is monitored in real time by radar. When the radar detects a target whose distance, speed and radar cross-section meet the preset distance threshold, speed threshold and radar cross-section threshold respectively, it determines that a bird target has been detected and generates a trigger signal. Step S102: In response to the trigger signal, the camera is woken up to acquire images of the area where the bird target is located, and the image acquired by the camera is used to identify whether the bird target is a bird; the camera has a built-in convolutional neural network model for identifying whether the target in the image is a bird.

[0075] Specifically, in this embodiment, the radar employs a frequency-modulated continuous wave (FMCV) system. Its working principle involves transmitting a continuous wave signal whose frequency varies linearly with time, receiving the echo signal reflected from the target, and calculating the target's distance by comparing the frequency difference between the transmitted and echo signals. The relationship between the frequency difference and the target distance is as follows: R = c×Δf / (2B×T_sweep); Where c is the speed of light, Δf is the frequency difference between the transmitted signal and the echo signal, B is the sweep bandwidth, and T_sweep is the sweep period.

[0076] The radar operates at a frequency of 24 GHz (gigahertz), with a 360° scanning range, a range resolution of no more than 0.5 m (meters), and a velocity resolution of no more than 0.1 m / s (meters per second). During detection, the radar can reliably distinguish between different targets such as birds, insects, fallen leaves, and wind-blown branches, with a false alarm rate of less than 1 time per day. The radar continuously detects targets during scanning. When a target is detected at a distance no greater than a preset range threshold (e.g., 20 m), with a velocity within a preset velocity threshold range (e.g., 1 m / s to 15 m / s), and with a radar cross-section (RCS) no less than a preset RCS threshold (e.g., -30 dBsm per square meter), the radar identifies the target as a bird and generates a trigger signal. The radar cross-section reflects the target's ability to reflect radar waves; different objects have different RCS values, with bird targets typically having an RCS value between -30 dBsm and -20 dBsm per square meter. Radar velocity measurement is achieved by measuring the Doppler frequency shift of the echo signal. The relationship between target velocity and Doppler frequency shift is as follows: v = f_d×c / (2f_0); Where f_d is the Doppler frequency shift and f_0 is the radar transmission frequency.

[0077] In response to a trigger signal, the central control module wakes up the camera. The camera, in standby mode (low power), is switched to operating mode by the trigger signal. The camera acquires image data of the target area. The camera's built-in convolutional neural network model processes the image data and outputs a classification result. This model is pre-trained using a large number of images labeled with both bird and non-bird targets, enabling it to identify whether a target is a bird. The model can employ a lightweight architecture (such as MobileNetV2) to reduce inference latency, with inference time controlled to within 150ms.

[0078] If the model output indicates that the target is a bird, a bird deterrent event is triggered, and the subsequent deportation process begins. If the model output indicates that the target is not a bird, no bird deterrent event is triggered, and the system remains in monitoring mode.

[0079] In some embodiments, each of the decoy units includes a laser decoy unit, an ultrasonic decoy unit, an acoustic decoy unit, and a red and blue warning light decoy unit; The method further includes: Step S310: During the execution of the deportation action, the deportation intensity and duration of each deportation unit are dynamically adjusted according to the dynamic information of the bird target. When the bird target approaches, the deportation intensity is increased and the deportation duration is extended. When the bird target moves away, the deportation intensity is reduced and the deportation is terminated early. The pointing direction of the laser deportation unit is adjusted according to the position change of the bird target.

[0080] Specifically, in this embodiment, dynamic adjustment occurs during the execution of the deterrent action. While executing the deterrent action, the central control module continuously receives dynamic information about the bird target from the detection and sensing module. This dynamic information includes the target's position, speed, and direction of movement.

[0081] The central control module determines whether the target is approaching or moving away from the protected area based on its movement trend relative to the protected area. If the target moves towards the protected area, it is determined to be approaching; if the target moves away from the protected area, it is determined to be moving away.

[0082] When a target approaches, the central control module increases the deterrence intensity of each deterrence unit. For laser deterrence units, increasing the deterrence intensity means increasing the laser power or accelerating the scanning speed. For ultrasonic deterrence units, increasing the deterrence intensity means increasing the sound pressure level or adjusting the frequency. For acoustic deterrence units, increasing the deterrence intensity means increasing the volume or switching to a more intimidating audio mode. For red and blue warning light deterrence units, increasing the deterrence intensity means increasing the flashing frequency or brightness. Simultaneously, the central control module extends the deterrence duration of each unit, i.e., extends the duration of each mode.

[0083] When the target is far away, the central control module reduces the deterrence intensity of each deterrence unit and terminates the deterrence action prematurely. Reducing the deterrence intensity is the opposite of increasing it as described above. Terminating the deterrence action prematurely means stopping the deterrence action before the preset duration is reached, in order to save energy.

[0084] Laser deterrence units are typically equipped with a pointing adjustment device that allows the laser beam to rotate horizontally and vertically. The central control module, based on real-time target azimuth and distance information detected by radar, controls the pointing adjustment device to adjust the laser beam's direction, ensuring it is always pointed at the target area.

[0085] In some embodiments, generating a reward value according to a preset reward rule based on the effectiveness level specifically includes: A base reward value is obtained based on the effectiveness level; where a successful expulsion corresponds to the first base reward value, a partially successful expulsion corresponds to the second base reward value, and an ineffective expulsion corresponds to the third base reward value. If the energy consumption of this expulsion action is lower than the preset energy consumption threshold, then a preset first additional reward value is obtained; wherein the first additional reward value is greater than 0. If the frequency of use of the deportation mode combination used in this expulsion exceeds a preset frequency threshold within a preset time period, a preset second additional reward value is obtained; wherein the second additional reward value is less than 0. The total reward value obtained by adding the base reward value, the first additional reward value, and the second additional reward value is used as the evaluation label.

[0086] Specifically, in this embodiment, the reward value is calculated by consisting of two parts: a basic reward value and an additional reward value.

[0087] The base reward value is determined by the effectiveness level. A successful expulsion corresponds to the first base reward value, a partial successful expulsion corresponds to the second base reward value, and an ineffective expulsion corresponds to the third base reward value. The first, second, and third base reward values ​​decrease sequentially, meaning a successful expulsion yields the highest base reward, and an ineffective expulsion yields the lowest. In a specific example, the first base reward value is +10, the second base reward value is +3, and the third base reward value is -5.

[0088] The first additional reward value is used to incentivize low-energy-consumption de-electrode operations. The central control module monitors energy consumption during each de-electrode action, which can be calculated from the operating current and voltage of each de-electrode unit. If the total energy consumption of this de-electrode action is lower than a preset energy consumption threshold, the first additional reward value is obtained. The first additional reward value is a positive number, in a specific example, +2. If the total energy consumption of this de-electrode action is not lower than the preset energy consumption threshold, the first additional reward value is not obtained.

[0089] The second additional reward value is used to penalize the frequent repetition of the same dispersal mode combination. The central control module counts the frequency of use of the dispersal mode combination employed within a preset time period. If the frequency exceeds a preset frequency threshold, the second additional reward value is obtained. The second additional reward value is a negative number, which is -4 in a specific example. If the frequency does not exceed the preset frequency threshold, the second additional reward value is not obtained.

[0090] The central control module adds the base reward value to the first additional reward value and the second additional reward value to obtain the total reward value, which is used as the evaluation label. For example, if a eviction is evaluated as a successful eviction (base reward value +10), the energy consumption is below the threshold (first additional reward value +2), and the usage frequency does not exceed the threshold (no second additional reward value is obtained), then the total reward value is +12.

[0091] In some embodiments, step S20 randomly selects one expulsion mode combination from a preset plurality of expulsion mode combinations, including: A roulette wheel method is used to randomly select one of a variety of pre-set drive-away mode combinations. Radar noise is used as the random seed. Each drive-away mode combination is assigned a corresponding selection interval according to the probability determined by the weight of each drive-away mode combination. The radar noise is used as a random source to generate random hit results that fall into a certain interval. The drive-away mode combination corresponding to the random hit result is used as the selected drive-away mode combination.

[0092] Specifically, in this embodiment, the roulette wheel selection method is a probabilistic random selection method. The central control module assigns a corresponding selection interval within the range of 0 to 1 for each dispersal mode combination based on its weight. The larger the weight value, the longer the corresponding interval, meaning a higher probability that the dispersal mode combination will be selected.

[0093] For example, suppose there are three combinations with weights of 0.5, 0.3, and 0.2, then the selected intervals for the three combinations are [0, 0.5), [0.5, 0.8), and [0.8, 1.0], respectively. The interval length is proportional to the weight.

[0094] Radar noise refers to the inherent electrical signal fluctuations generated by the thermal motion of electronic components within a radar receiver. This signal is inherently random and unpredictable. The central control module collects the instantaneous voltage value of this noise signal from the output of the radar receiver, converts the voltage value into a random number within the range of 0 to 1, and uses this number as the random hit result.

[0095] The random hit result is compared with the selected interval of each drive-off mode combination to determine which combination's interval the random number falls into. This drive-off mode combination is the one selected for this test. Because radar noise is unpredictable, the random hit result cannot be pre-calculated or inferred, thus ensuring the unpredictability of the selection result.

[0096] In some embodiments, the method further includes: Step S110: Before randomly selecting a drive-off mode combination from a preset variety of drive-off mode combinations, the weights of each drive-off mode are biased and adjusted according to the light intensity and the noise intensity, and the weights of each drive-off mode combination are updated according to the biased weights of each drive-off mode; wherein, if the light intensity is higher than a preset light threshold, the weight of the red-blue warning light drive-off mode is reduced; if the light intensity is lower than a preset low light threshold, the weights of the red-blue warning light drive-off mode and the sound drive-off mode are increased; if the noise intensity is higher than a preset noise threshold, the weights of the laser drive-off mode and the ultrasonic drive-off mode are increased.

[0097] Specifically, in this embodiment, bias adjustment refers to modifying the basic weights of each disengagement mode according to environmental conditions, so that the random selection results are more adapted to the current environment.

[0098] When the light intensity exceeds the preset light threshold, it indicates a strong light environment. The visual stimulation effect of the red and blue warning lights in strong light environments is weakened by ambient light; therefore, the weight of the red and blue warning light distancing mode is reduced to decrease the probability of this mode being selected in strong light environments. The preset light threshold can be set to 500 lux.

[0099] When the light intensity is below the preset low-light threshold, it indicates that the current environment is either low-light or nighttime. The visual impact of red and blue warning lights is significantly enhanced in low-light environments, and the deterrent effect of sound warnings is also more pronounced at night. Therefore, increasing the weight of these two modes increases their probability of being selected in low-light environments. The preset low-light threshold can be set to 10 lux.

[0100] When the noise intensity exceeds the preset noise threshold, it indicates a high-noise environment. The effect of sound-based repelling is masked by ambient noise; therefore, increasing the weights of the laser and ultrasonic repelling modes increases their probability of selection in high-noise environments. The preset noise threshold can be set to 70 dB.

[0101] The central control module recalculates the weights of each disengagement mode combination based on the bias-adjusted modal weights, and then performs the random selection in step S20 based on the updated combination weights.

[0102] In some embodiments, step S40 adjusts the weights of the deportation mode combination used in the current deportation in real time according to the effectiveness level of the current deportation, specifically including: After each bird deterrence event, if the effectiveness level of the deterrence is successful, the weight of the deterrence mode combination used in the deterrence event increases by 0.02; if the effectiveness level of the deterrence event is ineffective, the weight of the deterrence mode combination used in the deterrence event decreases by 0.03; the weight of each deterrence mode combination is limited to between 0.1 and 0.4.

[0103] Specifically, in this embodiment, the real-time adjustment is performed after each bird-repelling event, which is a short-cycle, rapid adjustment mechanism.

[0104] The magnitude of the weight increase or decrease is a fixed value. A successful removal increases the weight by 0.02, while an invalid removal decreases it by 0.03. The penalty is slightly larger than the reward, aiming to make the system eliminate invalid strategies more quickly.

[0105] The weight of each rejection mode combination is always limited to between 0.1 and 0.4. The lower limit of 0.1 ensures that each combination retains a certain probability of being selected, preventing any combination from being completely excluded, thus maintaining strategy diversity. The upper limit of 0.4 prevents any combination from having an excessively high weight, avoiding over-reliance on a single combination, and preventing birds from becoming adapted to that rejection mode combination due to frequent use.

[0106] If the weight increases to more than 0.4, it is truncated to 0.4; if the weight decreases to less than 0.1, it is increased to 0.1.

[0107] In some embodiments, the method includes: During the bird deterrence action, the central control module repeatedly re-executes the random selection of modal combinations, the random selection of working parameters, and the random determination of the start sequence and duration at random intervals, and re-encodes them into a new bird deterrence command sequence for execution; the random interval for each re-selection is independently and randomly selected within the range of 3 to 8 seconds.

[0108] Specifically, in this embodiment, the central control module also supports dynamic reselection during the execution of the bird deterrence action. Specifically, during the duration of a single deterrence action, the central control module performs multiple strategy reselections at random intervals. These random intervals are randomly selected within the range of 3 to 8 seconds, and each random interval is determined independently. That is, during the execution of this bird deterrence command sequence, the central control module does not remain fixed to the initially selected single modal combination. Instead, it repeatedly re-executes the random selection of modal combinations, random selection of operating parameters, and random determination of the start order and duration at random intervals, and re-encodes these into a new bird deterrence command sequence, which is then sent to the multimodal deterrence execution module for execution. This means that the modal combinations and operating parameters during a single deterrence action may also change.

[0109] For example, suppose in a bird deterrence event, the initially selected deterrence mode combination is "laser + sound," with operating parameters of 800mW laser power, audio number #07, and a total duration of 15 seconds. At the 4-second mark, the system triggers dynamic reselection (this time with a random interval of 4 seconds), reselecting the deterrence mode combination as "laser + ultrasound," with operating parameters of 500mW laser power and 25kHz ultrasound frequency. The remaining duration is re-determined based on the preset values ​​after the reselection. At the 10-second mark, the system triggers dynamic reselection again (this time with a random interval of 6 seconds), reselecting the deterrence mode combination as "laser + sound + light," with operating parameters of 1000mW laser power, audio number #12, and light flashing frequency of 8Hz. Execution ends at the 15-second mark.

[0110] The dynamic reselection mechanism operates on a completely randomized rhythm, with each reselection interval determined independently and randomly, independent of previous historical intervals. This mechanism ensures unpredictable variations in the modal combinations, working parameters, and temporal sequence of the deterrence stimuli experienced by birds, even within the duration of the same bird deterrence event. This further enhances the unpredictability of the deterrence stimuli and prevents birds from adapting to a specific combination or parameter during a single deterrence event. The central control module records the combination and parameters after each dynamic reselection and incorporates them into the execution record of the current bird deterrence event for subsequent effect evaluation and reinforcement learning training.

[0111] Another embodiment of this application provides a bird deterrent device to prevent birds from developing adaptive behavior, including a module for performing a bird deterrent method to prevent birds from developing adaptive behavior as described in the above embodiments of this application.

[0112] For example, such as Figure 2 As shown, the bird deterrent device includes a detection and sensing module 1, a central control module 2, and a multimodal bird deterrent execution module 3.

[0113] The detection and sensing module 1 is used to monitor the presence of bird targets within the protected area in real time, and to obtain the species, number, initial distance, and azimuth angle of the bird targets. The detection and sensing module 1 includes a millimeter-wave radar, a camera, a light sensor, and a microphone.

[0114] The millimeter-wave radar emits electromagnetic waves and receives echoes reflected from targets. It processes the echo signals to extract information about the target's distance, velocity, and radar cross-section. The radar generates a trigger signal when it detects a bird target that meets preset conditions. A camera, equipped with a built-in convolutional neural network model, captures images of the target area in response to the trigger signal. The camera processes the captured images and identifies whether the target is a bird. A light sensor, mounted on the device's outer casing, collects ambient light intensity and outputs an electrical signal. A microphone, located on the side of the device, collects ambient noise and outputs an audio signal.

[0115] The central control module 2 is electrically connected to the detection and sensing module 1 and the multimodal expulsion execution module 3, respectively. The central control module 2 includes a processor, a storage unit, a communication unit and a power management unit.

[0116] The processor runs a random policy generation engine and an adaptive optimization module to execute the steps described in the above embodiments. Specifically, the random policy generation engine is used to randomly select the expulsion mode combination, randomly select the working parameters, and randomly determine the start order and duration. The adaptive optimization module is used to execute the reinforcement learning algorithm and anti-overfitting constraint check.

[0117] The storage unit stores an audio library, a strategy parameter library, and a strategy effect database. The audio library contains at least 28 types of audio files, including bird predator calls, warning sounds, and sudden noises.

[0118] The communication unit supports wired and wireless communication methods, including 4G / 5G mobile communication, WiFi, and LoRa, and is used to interact with the remote monitoring platform to realize functions such as remote monitoring, policy configuration, data viewing, and firmware upgrades.

[0119] The power management unit provides a stable power supply to each module and manages energy consumption. It supports both solar panel and battery power supply modes, as well as mains power, allowing for flexible selection based on the deployment scenario. In solar power mode, the power management unit monitors battery level and solar input power in real time, automatically adjusting the operating status of each module to achieve intelligent power consumption management.

[0120] The multimodal bird deterrence execution module 3 is used to receive the bird deterrence command sequence sent by the central control module 2 and execute the bird deterrence action. The multimodal bird deterrence execution module 3 includes a laser deterrence unit, an ultrasonic deterrence unit, an acoustic deterrence unit, and a red and blue warning light deterrence unit.

[0121] The laser disengagement unit is used to generate and emit laser beams with a wavelength of 520nm. The power can be adjusted between five levels: 100mW, 300mW, 500mW, 800mW, and 1000mW. It supports modes such as line scanning, multi-point scanning, and random path scanning. The horizontal scanning angle is 180°, and the vertical scanning angle is ±30°.

[0122] The ultrasonic drive unit is used to generate and emit ultrasonic signals with a frequency range of 15kHz to 35kHz. It supports three working modes: fixed frequency, variable frequency, and sweep frequency. The sound pressure level is not less than 90dB and the coverage range is 360°.

[0123] The sound de-escalation unit is used to play de-escalation audio. The preset audio library contains at least 28 audio types and supports single audio playback, multi-audio overlay, and random sequence playback modes. The sound pressure level is not less than 106dB@1m.

[0124] The red and blue alarm light deterrent unit is used to generate high-frequency flashing colored lights. It adopts a high-brightness LED light source, including one set of red LEDs and one set of blue LEDs. It supports modes such as alternating red and blue flashing, single-color flashing, random frequency flashing and gradual flashing, with a brightness greater than 100 LM (lumens).

[0125] Each bird deterrence unit executes its own deterrence action according to the working parameters, start order, and duration specified in the bird deterrence instruction sequence.

[0126] It should be noted that the bird-repelling device in this embodiment corresponds to the bird-repelling method in the above embodiments. Therefore, any details not described in the bird-repelling device in this embodiment can be obtained by referring to the bird-repelling method in the above embodiments, and will not be elaborated in this embodiment.

[0127] In summary, the bird-repelling method and apparatus provided in this application for preventing birds from developing adaptive behaviors have at least the following advantages: The bird-repelling method and apparatus of this application randomly selects a repelling modality combination from a preset combination of repelling modalities. For each selected repelling modality, a set of working parameters is independently and randomly selected from its respective preset working parameter range. The activation sequence and duration of each selected repelling modality are also randomly determined, making each repelling operation unpredictable at the levels of modality combination, working parameters, and activation sequence. Simultaneously, during a single repelling operation, the system dynamically reselects the modality combination and working parameters multiple times at random intervals, causing the stimulus pattern to continuously change during each repelling operation. Birds cannot form effective cognitive associations, thus failing to recognize the repelling stimulus as a harmless signal and establish adaptive habits, effectively solving the problem of bird-repelling effectiveness decaying over time.

[0128] The bird-repelling method and apparatus of this application adjust the weights of the used de-icing modal combinations in real time according to the effectiveness level after each bird-repelling event, enabling the system to quickly respond to recent changes in effectiveness based on the effect of a single de-icing event. Simultaneously, by using the Q-learning algorithm at preset intervals to batch learn from all training samples in the policy effect database, the Q-values ​​of each state-action pair are updated, and the weights of each de-icing modal combination are normalized and updated, allowing the system to continuously learn from historical de-icing effects. The real-time weight adjustment and periodic reinforcement learning optimization work together to balance short-term adaptability and long-term optimality.

[0129] The bird-repelling method and apparatus of this application use a dual-sensor fusion mechanism of radar priority detection and image secondary confirmation to effectively distinguish birds from insects, fallen leaves and other disturbances, keeping the false alarm rate at an extremely low level and avoiding energy waste and ineffective stimulation caused by frequent false triggers.

[0130] The bird-repelling method and apparatus of this application adjust the modal weights by biasing the ambient light intensity and noise intensity, making the random selection results more adaptable to the current environmental conditions, thereby further improving the bird-repelling effect and energy utilization efficiency.

[0131] The bird deterrence method and apparatus of this application prevent birds from adapting to the strategy itself by means of an anti-over-adaptation constraint check, including limiting the number of times a modal combination is triggered within a preset time period, applying parameter offsets to frequently used modal combinations, and activating an enhancement mode when multiple consecutive failures or a low overall success rate.

[0132] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A bird-repelling method to prevent birds from developing adaptive behaviors, characterized in that, Includes the following steps: The system monitors in real time whether there are bird targets in the protected area. If so, it triggers a bird deterrence event and obtains the type, number, initial distance and orientation angle of the bird target, as well as the current ambient light intensity and noise intensity. In response to the bird deterrence event, a deterrence mode combination is randomly selected from a preset combination of multiple deterrence modes. The multiple deterrence modes combination consists of at least one deterrence mode among laser deterrence mode, ultrasonic deterrence mode, sound deterrence mode and red and blue warning light deterrence mode. The probability of each deterrence mode combination being selected is determined by the weight of each deterrence mode combination. For each of the selected de-icing mode combinations, a set of working parameters is randomly selected independently from its own preset working parameter range, and the start order and duration of each selected de-icing mode are randomly determined; the selected de-icing mode combination, the working parameters of each de-icing mode, the start order and duration are encoded into a bird deterrence command sequence, and each de-icing unit is controlled to perform de-icing actions according to the bird deterrence command sequence. After the bird removal action is completed, the dynamic information of the bird target continues to be monitored in real time. Based on the dynamic information of the bird target, it is determined whether the bird target has left the protected area. The effectiveness level of this removal is evaluated based on the determination result, and the weight of the removal mode combination used in this removal is adjusted in real time according to the effectiveness level of this removal. The effectiveness level includes successful removal, partial removal, and ineffective removal. When it is determined that the bird target has left the protected area, the departure time, departure direction, and departure distance of the bird target are also recorded. According to the effectiveness level, a reward value is generated according to the preset reward rules. The historical success rate of the light intensity, noise intensity, type, quantity, initial distance, and the driving mode combination used in this driving is used as the state parameter. The driving mode combination used in this driving and the working parameters of each driving mode are used as the action parameter. The reward value is used as the evaluation label. The combination is stored in the strategy effect database as a training sample. Every preset period, the Q-learning algorithm is used to update the Q value of each action parameter in each state based on all training samples in the policy effect database, and the weights of each disengagement mode combination are updated based on the updated Q value; wherein, the Q value is used to characterize the cumulative expected reward that can be obtained by selecting the corresponding action parameter in the corresponding state.

2. The bird-repelling method according to claim 1, characterized in that, The method further includes: Based on the updated weights of each eviction mode combination, perform an overfitting constraint check on each eviction mode combination. The over-adaptation constraint check includes: The number of times any expulsion mode combination can be triggered within a preset time period shall not exceed a preset limit. Once any expulsion mode combination exceeds the preset limit, it shall be prohibited from use for a preset prohibition period. For a combination of expulsion modes that is used more frequently than a preset threshold, when it is selected for expulsion, its operating parameters are randomly reselected within a preset offset range.

3. The method according to claim 2, characterized in that, The over-adaptation constraint check also includes: When the effectiveness level of three consecutive expulsions is all ineffective or the overall expulsion success rate is less than 80%, the boundary value of the current working parameter range of each expulsion mode is expanded outward by 20%, so that the random selection range of each working parameter is expanded. Within the expanded range, the upper or lower limit value of the preset working parameter range of each expulsion mode is selected first. At the same time, when randomly selecting expulsion mode combinations, the expulsion mode combinations with the highest frequency of use in the first 5 times are excluded.

4. The method according to claim 1, characterized in that, The real-time monitoring of whether bird targets exist in the protected area specifically includes: The radar monitors the protected area in real time. When the radar detects a target whose distance, speed, and radar cross-section meet preset distance thresholds, speed thresholds, and radar cross-section thresholds, it determines that a bird target has been detected and generates a trigger signal. In response to the trigger signal, the camera is activated to acquire images of the area where the bird target is located, and the image acquired by the camera is used to identify whether the bird target is a bird; the camera has a built-in convolutional neural network model for identifying whether the target in the image is a bird.

5. The method according to claim 1, characterized in that, Each of the de-escalation units includes a laser de-escalation unit, an ultrasonic de-escalation unit, an acoustic de-escalation unit, and a red and blue warning light de-escalation unit; The method further includes: During the execution of the deportation action, the deportation intensity and duration of each deportation unit are dynamically adjusted according to the dynamic information of the bird target. When the bird target approaches, the deportation intensity is increased and the deportation duration is extended. When the bird target moves away, the deportation intensity is reduced and the deportation is terminated early. The pointing direction of the laser deportation unit is also adjusted according to the position change of the bird target.

6. The method according to claim 1, characterized in that, The step of generating a reward value according to the validity level and a preset reward rule specifically includes: A base reward value is obtained based on the effectiveness level; where a successful expulsion corresponds to the first base reward value, a partially successful expulsion corresponds to the second base reward value, and an ineffective expulsion corresponds to the third base reward value. If the energy consumption of this expulsion action is lower than the preset energy consumption threshold, then a preset first additional reward value is obtained; wherein the first additional reward value is greater than 0. If the frequency of use of the deportation mode combination used in this expulsion exceeds a preset frequency threshold within a preset time period, a preset second additional reward value is obtained; wherein the second additional reward value is less than 0. The total reward value obtained by adding the base reward value, the first additional reward value, and the second additional reward value is used as the evaluation label.

7. The method according to claim 1, characterized in that, The step of randomly selecting a drive-off mode combination from a preset combination of multiple drive-off modes includes: A roulette wheel method is used to randomly select one of a variety of pre-set drive-away mode combinations. Radar noise is used as the random seed. Each drive-away mode combination is assigned a corresponding selection interval according to the probability determined by the weight of each drive-away mode combination. The radar noise is used as a random source to generate random hit results that fall into a certain interval. The drive-away mode combination corresponding to the random hit result is used as the selected drive-away mode combination.

8. The bird-repelling method according to claim 1, characterized in that, The method further includes: Before randomly selecting a drive-off mode combination from a set of preset drive-off mode combinations, the weights of each drive-off mode are biased and adjusted according to the light intensity and the noise intensity. The weights of each drive-off mode combination are then updated based on the adjusted weights. Specifically, if the light intensity is higher than a preset light threshold, the weight of the red-blue warning light drive-off mode is reduced; if the light intensity is lower than a preset low light threshold, the weights of the red-blue warning light drive-off mode and the sound drive-off mode are increased; and if the noise intensity is higher than a preset noise threshold, the weights of the laser drive-off mode and the ultrasonic drive-off mode are increased.

9. The method according to claim 1, characterized in that, The method of adjusting the weights of the deportation mode combination used in this deportation in real time based on the effectiveness level of this deportation includes: After each bird deterrence event, if the effectiveness level of the deterrence is successful, the weight of the deterrence mode combination used in the deterrence event increases by 0.02; if the effectiveness level of the deterrence event is ineffective, the weight of the deterrence mode combination used in the deterrence event decreases by 0.03; the weight of each deterrence mode combination is limited to between 0.1 and 0.

4.

10. A bird deterrent device for preventing birds from developing adaptations, comprising a module for performing the bird deterrent method according to any one of claims 1 to 9.