Bird control system and control method based on wheel-legged robot
Patent Information
- Application Number
- CN202610991971.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-07-06
AI Technical Summary
其一,多数移动驱鸟平台采用普通轮式底盘,难以适应农田、果园中常见的田垄、沟渠、碎石、坡地及根际凸起等复杂地形,易出现打滑、托底或陷车;而纯足式平台虽通过性较好,但在大面积巡护时能耗高、效率不足
通过数据获取模块同步采集视觉图像、声学音频、地形点云及本体位姿,并结合风险评估模块执行小目标检测与声纹识别,以生成鸟类视觉置信度、目标位置、声学置信度及声源方位,进而融合得到鸟害风险总评分,并利用声源方位对目标位置进行校验或补充以生成精确的风险目标空间坐标,从而显著提升了复杂农业环境下对鸟类目标的多模态早期发现能力与定位准确性;驱离决策模块基于鸟害风险总评分与空间坐标计算各预设驱离策略的综合效用值,并自动选择效用最高且满足安全边界的策略生成包含驱离方式、强度及区域的驱离任务指令,实现了对鸟群适应行为的主动抑制与策略的理性调度;执行控制模块根据地形点云实时确定局部地形可通行度并自适应切换轮腿式机器人的运动模式,同时基于该运动模式根据驱离任务指令生成云台指向、声光激光驱动及底盘运动指令,使机器人能够灵活逼近风险目标并执行分级驱离动作。由此,本发明有效克服了固定式驱鸟设备覆盖范围小、单一驱离手段易被适应及复杂地形通过性差等缺陷,在保证作业安全的前提下大幅提升了单机覆盖效率、鸟情响应速度与长期驱鸟效果。
Smart Images

Figure CN122488829B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and more specifically, to a bird control system and method based on a wheeled-legged robot. Background Technology
[0002] In agricultural settings, bird pecking at ripe fruit, grains in the filling stage, and seedlings often causes significant yield reduction and quality decline. Traditional bird deterrence methods include scarecrows, reflective strips, blasting devices, ultrasonic devices, and fixed speakers. While these methods are low-cost, they suffer from limited coverage radius, poor directionality, and rapid attenuation of effectiveness due to bird adaptation.
[0003] In recent years, some studies have attempted to use mobile robot platforms for bird control tasks, but existing solutions still have the following technical shortcomings: Firstly, most mobile bird deterrent platforms use ordinary wheeled chassis, which are difficult to adapt to the complex terrain commonly found in farmland and orchards, such as ridges, ditches, gravel, slopes, and root protrusions, and are prone to slipping, bottoming out, or getting stuck. While pure-footed platforms have better passability, they consume a lot of energy and are not efficient enough when patrolling large areas.
[0004] Secondly, the perception layer relies heavily on a single visible light camera, which makes it easy to miss small target birds in environments with obstructed branches, backlighting, or low light, and it is difficult to provide early warnings before flocks of birds enter crop areas.
[0005] Third, the methods of repelling birds are usually fixed patterns of sound and light stimulation, which lack the ability to dynamically evaluate the repelling effect on different bird species and in different scenarios and to adapt strategies. After long-term use, bird flocks are prone to adapting, leading to the failure of repelling. Summary of the Invention
[0006] The problem solved by this invention is one or more of the aforementioned related technical problems.
[0007] To address the aforementioned problems, this invention provides a bird control system and method based on a wheeled-legged robot.
[0008] In a first aspect, the present invention provides a bird deterrence control system based on a wheeled-legged robot, comprising: The data acquisition module is used to collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area; The risk assessment module is used to perform small target detection on the visual image sequence to generate bird visual confidence and target location, perform voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source orientation, and perform fusion processing based on the visual confidence, the acoustic confidence, and the target location to generate a total bird damage risk score; and verify or supplement the target location based on the sound source orientation to generate the spatial coordinates of the risk target. The bird deterrence decision module is used to generate a comprehensive utility value for each bird deterrence strategy based on a preset deterrence strategy, according to the total bird damage risk score and the spatial coordinates of the risk target, select the deterrence strategy with the highest comprehensive utility value and that meets the preset safety boundary, and generate a deterrence task instruction, wherein the deterrence task instruction includes deterrence method, execution intensity and area of action. The execution control module is used to determine the local terrain accessibility based on the terrain point cloud and adaptively switch the motion mode of the wheeled robot; and based on the motion mode, it generates gimbal pointing instructions, acousto-optical-laser actuator drive signals and chassis motion instructions according to the expulsion task instructions, so as to control the wheeled robot to approach the risk target and perform graded expulsion actions.
[0009] Optionally, the movement modes include a wheeled mode, a walking obstacle-crossing mode, and a hybrid wheel-legged mode; the step of determining local terrain accessibility based on the terrain point cloud and adaptively switching the movement modes of the wheel-legged robot includes: The terrain point cloud is projected onto a horizontal plane and divided into grids, with each grid serving as a local terrain block. The local height variation standard deviation, local slope, and ground adhesion coefficient are estimated based on the point cloud data in each grid cell for the corresponding local terrain block. The local elevation fluctuation standard deviation, the local slope, and the ground adhesion coefficient estimate are fused together to obtain the local terrain accessibility corresponding to the local terrain block. When the terrain accessibility is higher than the first threshold, switch to wheel mode; When the terrain accessibility is between the first threshold and the second threshold, switch to the wheel-leg hybrid mode; When the terrain accessibility is below the second threshold, switch to walking obstacle crossing mode.
[0010] Optionally, performing voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source location includes: The multi-channel acoustic audio is framed, denoised, and subjected to short-time Fourier transform to extract the decibel power spectrum; Based on the Mel frequency cepstral coefficients, the voiceprint embedding features are extracted from the decibel power spectrum, input into the preset bird voiceprint recognition model, and the acoustic confidence of the bird is generated. Based on the arrival time difference of each sound source in the multi-channel acoustic audio, the spatial location of the chirping sound source is determined, and the sound source location information is generated.
[0011] Optionally, the step of generating a comprehensive utility value for each of the bird-repelling strategies based on a preset bird-damage risk score and the spatial coordinates of the risk target includes: Based on the total bird damage risk score and the spatial coordinates of the risk target, determine the expected deterrence effectiveness and safety disturbance cost of each deterrence strategy; Using Equation 1, the corresponding comprehensive utility value is generated based on the expected expulsion effectiveness, the security disturbance cost, the fitness estimate, and the energy consumption cost. ; in, Let be the combined utility value of the k-th expulsion strategy. to For adjustment coefficients, To predict the effectiveness of the eviction, This is an estimate of the fitness of the flock to the k-th dispersal strategy. The energy cost of implementing the k-th expulsion strategy, The safety disturbance cost of the k-th expulsion strategy.
[0012] Optionally, the fitness estimate is obtained in the following way: The duration of the flock's stay, return time, and approach speed after the expulsion were obtained to determine the real-time adaptation indicators at the current moment. Based on the fitness estimate and memory factor from the previous moment, the historical weighted value is obtained; based on the immediate adaptation index and the memory factor, the current feedback weighted value is obtained. The fitness estimate at the current moment is obtained based on the historical weighted value and the current feedback weighted value.
[0013] Optionally, the step of fusing the visual confidence score, the acoustic confidence score, and the target location to generate a total bird damage risk score includes: Based on a preset bird risk scoring model, the visual confidence level, acoustic confidence level, and target location are fused together to generate the total bird damage risk score. The preset bird risk scoring model includes: ; in, The overall score for the bird damage risk; The recognition confidence is obtained by weighting the visual confidence and the acoustic confidence. This is an occlusion correction term generated by cross-validation of visible light images and thermal infrared images in the visual image sequence; The relative motion term of the target position toward the crop core area; This is the population density term; This is a historically high-incidence item generated based on historical eviction records; to Configurable weights.
[0014] Optionally, the bird control system based on the wheeled-legged robot further includes a multi-robot collaboration module, which is used to receive shared information from other robots in the same working area. The shared information includes bird flock trajectory, heat map, remaining battery power and areas where passage is obstructed. When a robot is unable to cover adjacent high-risk areas due to insufficient power, terrain limitations, or being engaged in a high-level expulsion mission, it generates a replacement request to request a neighboring robot to take over.
[0015] Optionally, the bird control system based on the wheeled-legged robot further includes an edge-cloud collaboration module, which is used to upload sample data and receive updated model parameters and strategy knowledge graphs when communicating with the cloud.
[0016] Optionally, the execution control module generates the chassis motion commands, specifically including: Based on the expulsion mission command and the current motion mode, the desired body motion state is determined, including the desired body speed, the desired body posture, and the desired leg support force, and the desired body motion state is used as the content of the chassis motion command. Based on the whole-body control law, the desired body motion state in the chassis motion command is converted into the output torque vector of each actuator in the wheel-legged robot to drive the wheel drive wheel and the leg lifting mechanism to execute the chassis motion command; The whole-body control law is as follows: ; in, J is the output torque vector; J is the robot's Jacobian matrix. For expected support; and These represent the desired joint state and the current joint state, respectively. and These are the expected speed and the current speed, respectively. and This is the gain matrix.
[0017] Secondly, the present invention provides a bird deterrence control method based on a wheeled-legged robot, applied to the bird deterrence control system based on a wheeled-legged robot as described in the first aspect, wherein the bird deterrence control method based on a wheeled-legged robot includes: Collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area; Small target detection is performed on the visual image sequence to generate bird visual confidence and target location; voiceprint recognition and sound source localization are performed on the multi-channel acoustic audio to generate bird acoustic confidence and sound source orientation; and the visual confidence, acoustic confidence, and target location are fused to generate a total bird damage risk score; the target location is verified or supplemented based on the sound source orientation to generate the spatial coordinates of the risk target. Based on the preset bird deterrence strategy, a comprehensive utility value corresponding to each deterrence strategy is generated according to the total bird damage risk score and the spatial coordinates of the risk target. The deterrence strategy with the highest comprehensive utility value and that meets the preset safety boundary is selected, and a deterrence task instruction is generated. The deterrence task instruction includes deterrence method, execution intensity and area of action. The local terrain accessibility is determined based on the terrain point cloud, and the motion mode of the wheeled robot is adaptively switched. Based on the motion mode, gimbal pointing instructions, acousto-optical-laser actuator drive signals, and chassis motion instructions are generated according to the expulsion task instructions to control the wheeled robot to approach the risk target and perform graded expulsion actions.
[0018] The beneficial effects of the bird deterrence control system and method based on a wheeled-legged robot of the present invention are: The data acquisition module synchronously collects visual images, acoustic audio, terrain point clouds, and the robot's pose. Combined with the risk assessment module, it performs small target detection and voiceprint recognition to generate bird visual confidence, target location, acoustic confidence, and sound source orientation. These are then fused to obtain a total bird damage risk score. The sound source orientation is used to verify or supplement the target location to generate accurate spatial coordinates of the risk target, thus significantly improving the multimodal early detection capability and positioning accuracy of bird targets in complex agricultural environments. The repulsion decision module calculates the comprehensive utility value of each preset repulsion strategy based on the total bird damage risk score and spatial coordinates. It automatically selects the strategy with the highest utility that meets the safety boundary to generate repulsion task instructions that include repulsion method, intensity, and area. This achieves active suppression of bird flock adaptive behavior and rational scheduling of strategies. The execution control module determines the local terrain accessibility in real time based on the terrain point cloud and adaptively switches the motion mode of the wheeled robot. Based on this motion mode, it generates gimbal pointing, acoustic-optical-laser drive, and chassis motion instructions according to the repulsion task instructions, enabling the robot to flexibly approach the risk target and perform graded repulsion actions. Therefore, this invention effectively overcomes the shortcomings of fixed bird deterrence equipment, such as small coverage area, easy adaptation of single deterrence methods, and poor passability in complex terrain. Under the premise of ensuring operational safety, it greatly improves the coverage efficiency of a single machine, the speed of bird response, and the long-term bird deterrence effect. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a bird control system based on a wheeled-legged robot according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a bird control method based on a wheeled-legged robot according to an embodiment of the present invention. Detailed Implementation
[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0021] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0022] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0023] It should be noted that the one or more modifications mentioned in this invention are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as one or more.
[0024] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0025] Existing solutions generally lack safety identification and restraint mechanisms for personnel, livestock, vehicles, and sensitive facilities, posing safety risks when carrying out forceful expulsion actions (such as laser scanning and high-frequency sound waves) in agricultural production areas.
[0026] In stand-alone operation mode, the coverage and reliability are limited, making it difficult to achieve regional collaborative defense and continuous patrols. Furthermore, the system lacks the ability to update models and transfer knowledge through edge-cloud collaboration, which means that experience needs to be re-accumulated when deploying in different plots or seasons, resulting in a long cold start cycle.
[0027] Therefore, how to systematically integrate wheeled-legged hybrid mobility platforms, multimodal environmental perception, adaptive strategy scheduling, and safety constraint mechanisms to construct a bird control system capable of autonomously patrolling complex agricultural terrain, detecting bird activity early, continuously and effectively driving away birds, and ensuring operational safety is a technical problem that urgently needs to be solved in this field.
[0028] To solve the above problems, such as Figure 1 As shown in the figure, an embodiment of the present invention provides a bird deterrence control system based on a wheeled-legged robot, comprising: The data acquisition module is used to collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area.
[0029] Specifically, the data acquisition module is configured as follows: during robot patrol or stay, it synchronously acquires continuous frame images of the work area at a set frame rate using a visible light camera and a thermal infrared camera installed on the gimbal or the front of the robot body, forming the visual image sequence. The visible light images are used to capture the appearance of birds, the structure between crop rows, and the texture and color information of safety targets (personnel, vehicles). The thermal infrared images are used to enhance the thermal radiation signals of warm-blooded organisms (birds, livestock) under low light, shadow, or foliage obstruction conditions. The module also continuously acquires environmental data through an array of at least three microphones distributed in a preset geometric layout around the robot's casing. Multi-channel sound signals are used to form the multi-channel acoustic audio, which includes components such as bird calls, wing flapping, flocking movements, and environmental noise (wind noise, mechanical sounds, and human voices). The area around the robot is actively scanned using millimeter-wave radar or lidar to acquire three-dimensional spatial point cloud data, forming the terrain point cloud. This point cloud records the elevation information of the ground surface, obstacles, crops, and terrain undulations. The robot's real-time position, attitude, angular velocity, and linear velocity in the working coordinate system are jointly calculated using an inertial measurement unit, wheel speedometer, and an optional global navigation satellite system receiver, forming the robot's pose information. These four types of data are synchronized with a unified timestamp and output to subsequent modules, providing the raw information foundation for risk assessment and motion control.
[0030] The data acquisition module provides the robot with comprehensive perception dimensions, including vision, hearing, terrain, and its own state, by simultaneously collecting multimodal and multi-view environmental and ontological data. Visual image sequences (complementary visible light and thermal infrared) cover target detection needs under day and night conditions and occlusion. Multi-channel acoustic audio extends the passive listening capability for bird flocks at a distance or outside the field of view. Terrain point clouds endow the robot with geometric cognition capabilities for complex agricultural terrain (ridges, ditches, slopes), while ontological pose information ensures the accuracy of spatiotemporal alignment and motion feedback of sensor data. This fundamentally solves the problem of missing or unreliable information from a single sensor in unstructured environments such as orchards and farmland, significantly improving the quality of raw data for early bird detection, terrain adaptation, and safety monitoring, laying a solid foundation for subsequent efficient and safe bird control decisions.
[0031] The risk assessment module is used to perform small target detection on the visual image sequence to generate bird visual confidence and target location, perform voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source orientation, and perform fusion processing based on the visual confidence, the acoustic confidence, and the target location to generate a total bird damage risk score; and verify or supplement the target location based on the sound source orientation to generate the spatial coordinates of the risk target.
[0032] Specifically, the risk assessment module is configured as follows: First, the visible light frames and thermal infrared frames in the visual image sequence are spatiotemporally aligned, and a small target detection network based on deep learning (such as multi-scale feature fusion and attention mechanism) is used to identify individual birds or groups of birds in the image frame by frame, and output the visual confidence score of the bird corresponding to each detection box (indicating the probability value of the presence of birds in the area) and the target position of the bird in the image coordinate system (e.g., the coordinates of the center point or foot point of the bounding box, which can be combined with depth information or known camera models to convert it into a three-dimensional position in the robot coordinate system).
[0033] Secondly, the multi-channel acoustic audio is segmented, denoised, and feature-extracted (e.g., Mel-frequency cepstral coefficients). A pre-defined bird voiceprint recognition model (e.g., recurrent neural network or converter model) is used to determine whether the audio segment contains a specific bird call, outputting a bird acoustic confidence score (representing the probability that the acoustic event originates from a bird). Simultaneously, the time difference or phase difference of sound arrival between the microphone array channels is used to calculate the direction and pitch angles of the sound source relative to the robot coordinate system, obtaining the sound source's location. Then, the visual confidence score and the acoustic confidence score are used as multimodal evidence of bird activity. Combined with the spatial distribution of the target location (e.g., the distance and speed direction of the target entering the crop core area), a weighted fusion or Bayesian inference is performed to generate a comprehensive bird damage risk score, used to quantify the urgency of bird damage in the current work area. Finally, since visual detection may experience positional deviations or missed detections in low light, occlusion, or at long distances, while the sound source orientation can provide supplementary clues, this module further utilizes the sound source orientation to verify the target position (for example, if the difference between the sound source orientation and the visual target orientation is less than a threshold, the visual position is used; if the difference is large, the weighted center of the two is taken) or, when there is no visual target, only the sound source orientation is used as a preliminary spatial orientation, thereby obtaining the accurate spatial coordinates of the risk target (i.e., the three-dimensional coordinates of the target point that the robot believes is most likely to be driven away in the working coordinate system).
[0034] The risk assessment module achieves multimodal fusion of visual small target detection and acoustic voiceprint recognition results, compensating for the blind spots of single sensors in situations such as foliage obstruction, low light, or long distances. This significantly improves the recall rate and reliability of bird target identification. Simultaneously, it uses sound source orientation to verify and supplement the visual target position, effectively reducing positioning errors caused by changes in lighting, background interference, or temporary target loss in complex agricultural scenarios, ensuring the continuity and accuracy of the spatial coordinates of risk targets. The resulting bird damage risk score comprehensively reflects multiple dimensions such as the probability of bird presence, invasion direction, flock density, and crop vulnerability stage. This provides a scientific and quantitative basis for the rational selection of subsequent repelling strategies, greatly enhancing the system's situational awareness and response timeliness in real farmland environments.
[0035] The bird deterrence decision module is used to generate a comprehensive utility value for each bird deterrence strategy based on a preset deterrence strategy, according to the total bird damage risk score and the spatial coordinates of the risk target, select the deterrence strategy with the highest comprehensive utility value that meets the preset safety boundary, and generate a deterrence task instruction, wherein the deterrence task instruction includes deterrence method, execution intensity and area of effect.
[0036] Specifically, upon receiving the overall bird damage risk score and the spatial coordinates of the risk target, the module first retrieves multiple pre-defined graded repulsion strategies (e.g., low-disturbance approach strategy, directional acoustic strategy, strobe reflection strategy, restricted laser scanning strategy, etc.). Each strategy is associated with a set of predefined execution parameters (e.g., sound pressure level, strobe frequency, laser scanning angle and power, duration of action, etc.). For the current risk target, the module calculates a comprehensive utility value for each candidate strategy. This utility value reflects the expected benefits and costs of applying the strategy to the current scenario. The calculation comprehensively considers the following factors: the higher the overall bird damage risk score, the more effective the strategy needs to be, and the higher the expected repulsion effectiveness of the corresponding strategy; the closer the spatial coordinates of the risk target are to safety-sensitive areas (e.g., sidewalks, sheds, roads), the greater the safety disturbance cost of the high-disturbance strategy. The module also refers to the historical usage effects of each strategy (i.e., the degree to which the bird flock adapts to the strategy) and the estimated energy consumption of executing the strategy. By weighted fusion of the above multiple indicators, the comprehensive utility value of each strategy is obtained.
[0037] In some embodiments, the hierarchical expulsion strategy in the expulsion decision module includes at least four levels, L1 to L4: Level 1 employs a low-disturbance patrol strategy, using the robot's physical presence and proximity to drive away enemies. Level L2 employs a directional acoustic strategy, using a pan-tilt unit to drive directional loudspeakers to emit sound waves with variable frequency bands and modulation patterns for dispersal. Level L3 employs a strobe and reflection strategy, emitting visible light stimuli of variable frequency through strobe lights and reflective components; Level L4 is a restricted laser scanning strategy, which allows laser scanning to be performed only after authorization through safety boundary verification, and at limited angles and power. Subsequently, the module reads pre-loaded preset safety boundaries (e.g., electronic fences for no-fire zones, real-time detection frames for personnel / vehicles / livestock, laser up-tilt angle restrictions, etc.) and excludes any strategies that touch the safety boundaries (e.g., the presence of personnel or reflective surfaces in the laser scanning direction). Among the remaining strategies, the one with the highest overall utility value is selected as the current optimal strategy, and a structured deportation task instruction is generated accordingly. This instruction contains at least three parts: deportation method (specifying which actuator combination to invoke, such as speaker + strobe light), execution intensity (e.g., sound pressure level and modulation mode for acoustic deportation, power duty cycle and scanning speed for laser scanning), and area of effect (determining the gimbal pointing angle or chassis positioning target point based on the spatial coordinates of the risk target and the estimated escape direction). Finally, this instruction is sent to the execution control module and the safety monitoring module respectively to trigger specific deportation actions and ensure that the action process does not exceed the boundaries.
[0038] The bird control decision-making module takes the total bird damage risk score and the spatial coordinates of the risk target as input. By calculating the comprehensive utility value of each preset bird control strategy and filtering it in conjunction with real-time safety boundaries, it achieves rational strategy-level optimization for each bird incident. Compared with fixed logic or randomly switching bird control methods, this module can proactively upgrade the bird control intensity under high-urgency risk and prioritize low-disturbance strategies when the risk is low or near sensitive areas, thus ensuring both effectiveness and operational safety and environmental friendliness. In addition, by incorporating the historical fitness of strategies into the utility evaluation, the system can implicitly suppress the bird flock's tolerance to a single stimulus, delay strategy failure, and maintain long-term bird control effectiveness. The final generated structured instructions, including the bird control method, execution intensity, and area of effect, provide clear and quantifiable action guidelines for the underlying execution module, ensuring the accuracy and efficiency of the entire bird control closed loop.
[0039] The execution control module is used to determine the local terrain accessibility based on the terrain point cloud and adaptively switch the motion mode of the wheeled robot; and based on the motion mode, it generates gimbal pointing instructions, acousto-optical-laser actuator drive signals and chassis motion instructions according to the expulsion task instructions, so as to control the wheeled robot to approach the risk target and perform graded expulsion actions.
[0040] Specifically, firstly, after receiving the terrain point cloud, the point cloud is rasterized, and the local terrain accessibility (a comprehensive quantitative value reflecting the ability of the wheeled robot to safely and stably pass through the area in each local grid cell) is calculated. A higher value indicates greater suitability for wheeled or fast movement. Then, based on changes in accessibility, the wheeled robot's movement mode is adaptively switched. For example, a high-speed, low-energy wheeled mode is used on flat, hard surfaces; a hybrid wheeled mode is used on slightly rugged or soft soil to improve adhesion and posture stability; and a walking obstacle-crossing mode is switched when encountering high embankments, deep ditches, or dense obstacles. Secondly, based on the determined (or switched) motion mode, and in conjunction with the expulsion task instructions (including the expulsion method, execution intensity, and area of effect), the following three actions are performed: ① Generate gimbal pointing instructions to control the two-axis or three-axis gimbal mounted on the upper part of the robot to drive the expulsion actuators (speakers, strobe lights, lasers, etc.) to accurately align with the spatial coordinates of the risk target or the estimated escape direction; ② Generate acoustic-optical-laser actuator drive signals to drive each actuator to output corresponding stimuli in real time according to the expulsion method and execution intensity specified in the expulsion task instructions (such as the sound pressure level and modulation waveform of directional acoustics, the flashing frequency of strobe lights, and the power and sweeping range of laser scanning); ③ Generate chassis motion instructions to plan the robot's linear velocity, angular velocity, leg support posture, and path according to the current motion mode and expulsion task requirements (such as close approach or lateral encirclement), so that the robot can safely and smoothly approach the risk target and maintain body stability during the expulsion action. Through the aforementioned collaborative control, the robot can autonomously move to a suitable location and complete the bird-repelling task using graded, target-oriented sound, light, laser, and posture display methods.
[0041] The wheel-legged robot is an autonomous mobile platform that combines wheel drive with a leg-type lifting / obstacle-crossing mechanism on the same mobile chassis. It is specifically designed for complex agricultural terrain (such as field ridges, ditches, slopes, orchard root zones, muddy roads, etc.). Its specific structure and working principle are described below: Wheeled-legged robots mainly include: 1. Main frame of the vehicle body: It is a rigid load-bearing structure used to install batteries, computing units, sensing components, drive-away gimbals and other auxiliary equipment; the vehicle body is designed to have a low center of gravity to improve its anti-overturning ability.
[0042] 2. Wheel drive system: This system includes at least two active drive wheels (usually located on the sides or front and rear of the vehicle). Each drive wheel is driven by an independent hub motor or a servo motor with a reducer, responsible for achieving high-speed, low-energy cruising movement on flat, hard, and well-adhesive surfaces. The drive wheels are also equipped with encoders for speed feedback and odometer calculation.
[0043] 3. Leg-type lifting / obstacle-crossing mechanism: Contains at least two or more actively controllable leg units, typically arranged symmetrically on both sides of the vehicle body (e.g., one pair at the front and rear, or one pair on each side). Each leg unit includes: Lifting actuators (such as electric push rods, lead screw mechanisms, and linkage swing arms): can extend or lift the legs relative to the vehicle body, thereby changing the vehicle's ground clearance, pitch angle, and roll angle.
[0044] Foot end: The part that contacts the ground, usually designed with anti-slip and shock-absorbing structures (such as rubber pads or spherical hinges).
[0045] Or joint actuators: Some advanced leg units also have multiple degrees of freedom (such as hip and knee joints), enabling them to mimic quadrupedal or hexapedal gait.
[0046] 4. Suspension and buffer components: Elastic and damping elements that connect the wheels, legs and vehicle body to absorb ground impacts, improve ride comfort and sensor stability.
[0047] 5. Braking and emergency stop device: Used to keep the vehicle stationary on a slope or in an emergency.
[0048] Depending on the accessibility of different terrains, the robot can adaptively switch between the following three modes: Wheeled mode: The leg units are fully retracted (off the ground or only slightly in contact but not providing major support), and the robot moves solely by the drive wheels. In this mode, the robot moves quickly, consumes little energy, and is easy to control, making it suitable for paved roads, leveled field ridges, and flat passageways between rows.
[0049] Hybrid wheel-leg mode: The drive wheels remain the primary drive, but the leg units extend moderately to adjust the vehicle's ground clearance, increase the number of support points, or change the vehicle's posture (for example, in soft soil, the legs gently insert into the ground to increase traction; or on a side slope, the legs extend at unequal lengths to keep the vehicle level). This mode is suitable for transitional terrain such as slightly rugged terrain, root zones, and shallow ditches, balancing passability and efficiency.
[0050] Walking obstacle crossing mode: The drive wheels are locked or follow the movement, and the vehicle mainly relies on the alternating lifting, forward movement, and support of the leg units to achieve stepping movement (similar to a quadrupedal / sixth-legged gait). This mode is used in extreme terrains such as high embankments, deep ditches, steep slopes, and areas with high grass entanglement, enabling it to cross high obstacles and prevent getting stuck.
[0051] The triggering condition for mode switching is automatically determined by the execution control module based on the "local terrain traversability" score (obtained by fusing slope, undulation standard deviation, adhesion coefficient, etc. from terrain point cloud computing) and current task requirements (such as whether the target needs to be urgently driven away), and the stability of the switching process is ensured by the torque smooth transition algorithm.
[0052] Wheel-legged robots not only include a chassis and its motion mechanism, but also integrate the complete system required to complete bird-scaring tasks: Upper body platform: The gimbal-type drive-away actuator (directional speaker, strobe light, laser) is installed on the top of the vehicle body, and the horizontal and pitch directions are achieved through the gimbal.
[0053] Perception and Computing: Multimodal sensors (cameras, microphone arrays, millimeter-wave / LiDAR, IMU, etc.) and edge computing units are installed in appropriate locations on the vehicle body (such as the roof, front bumper, or dedicated brackets) to ensure visibility and stability.
[0054] Power and Communication: Battery packs, wireless communication modules, etc. support long-term autonomous operation and edge-cloud collaboration.
[0055] The wheel-legged robot used in this application combines the high speed and long endurance of wheeled robots with the powerful obstacle-crossing and posture adjustment capabilities of legged robots. Through adaptive switching of three motion modes, it can achieve a mobility strategy of "running fast on flat roads, moving slowly on rough roads, and crossing obstacles" in complex agricultural environments. Compared with a single wheeled chassis, it significantly reduces the risk of getting stuck, slipping, or bottoming out; compared with a purely legged robot, it significantly reduces energy consumption and improves coverage efficiency in conventional patrol scenarios. At the same time, the stable chassis and precise mode switching provide a stable working platform for the multimodal sensing components and bird-repelling gimbal mounted on the body, ensuring the continuity of sensing data and the accuracy of bird-repelling direction, thus enabling the entire bird-repelling system to operate autonomously and reliably in real-world operating scenarios such as orchards, farmland, and grain bases for a long time.
[0056] The execution control module integrates terrain perception and adaptive motion mode switching into the bird deterrence task execution process, enabling the robot to move flexibly and reliably in complex agricultural terrains such as orchard paths, field ridges, and muddy, slippery areas. This ensures both obstacle-crossing capability and patrol efficiency. Simultaneously, through the joint generation of gimbal pointing commands, acoustic, optical, and laser drive signals, and chassis motion commands, coordinated actions between the deterrence actuators and the robot base are achieved. Specifically, the gimbal maintains lock-on on the target, the chassis executes positioning or flanking maneuvers, and each actuator outputs appropriate intensity stimuli according to a tiered strategy. This tightly coupled control significantly improves the directionality and deterrent effect of the deterrence action, avoiding blind pursuit or ineffective deterrence, and ensuring operational safety and energy efficiency in complex scenarios.
[0057] In this embodiment, the bird control system based on a wheeled-legged robot: Simultaneously acquires visual images, acoustic audio, terrain point clouds, and the bird's pose through a data acquisition module, and combines this with a risk assessment module to perform small target detection and voiceprint recognition to generate bird visual confidence, target location, acoustic confidence, and sound source orientation. These are then fused to obtain a total bird damage risk score, and the sound source orientation is used to verify or supplement the target location to generate accurate spatial coordinates of the risk target. This significantly improves the multimodal early detection capability and positioning accuracy of bird targets in complex agricultural environments; the bird deterrence decision module is based on... The invention calculates the overall utility value of each preset bird deterrence strategy based on the bird risk score and spatial coordinates, and automatically selects the strategy with the highest utility that meets the safety boundary to generate deterrence task instructions including deterrence method, intensity, and area. This achieves proactive suppression of bird flock adaptive behavior and rational scheduling of strategies. The execution control module determines the local terrain accessibility in real time based on the terrain point cloud and adaptively switches the motion mode of the wheeled robot. Based on this motion mode, it generates gimbal pointing, acoustic-optical-laser drive, and chassis motion instructions according to the deterrence task instructions, enabling the robot to flexibly approach the risk target and execute graded deterrence actions. Thus, this invention effectively overcomes the shortcomings of fixed bird deterrence equipment, such as small coverage area, easy adaptation of single deterrence methods, and poor passability in complex terrain. While ensuring operational safety, it significantly improves the single-machine coverage efficiency, bird situation response speed, and long-term bird deterrence effect.
[0058] Optionally, the movement modes include a wheeled mode, a walking obstacle-crossing mode, and a hybrid wheel-legged mode; the step of determining local terrain accessibility based on the terrain point cloud and adaptively switching the movement modes of the wheel-legged robot includes: The terrain point cloud is projected onto a horizontal plane and divided into grids, with each grid serving as a local terrain block; The local height variation standard deviation, local slope, and ground adhesion coefficient are estimated based on the point cloud data in each grid cell for the corresponding local terrain block. The local elevation fluctuation standard deviation, the local slope, and the ground adhesion coefficient estimate are fused together to obtain the local terrain accessibility corresponding to the local terrain block. When the terrain accessibility is higher than the first threshold, switch to wheel mode; When the terrain accessibility is between the first threshold and the second threshold, switch to the wheel-leg hybrid mode; When the terrain accessibility is below the second threshold, switch to walking obstacle crossing mode.
[0059] Specifically, the execution control module first projects the acquired terrain point cloud (a set of three-dimensional spatial points obtained by millimeter-wave radar or lidar scanning) onto a horizontal ground, and divides it into regular grids according to a set resolution (e.g., 0.1m × 0.1m), with each grid corresponding to a local terrain block. For each local terrain block, the following three feature parameters are calculated from the point cloud data within the grid: Local height variation standard deviation : Statistically calculate the elevation values (relative height) of all points within a grid, and calculate their standard deviation to reflect the flatness of the ground in that area. The smaller the value, the flatter the ground (such as a paved road or a leveled field ridge). The larger the value, the more rugged the terrain (such as gravel ground, raised root zones, and tall grass accumulation).
[0060] Local slope ( ): Perform plane fitting on the point cloud within the grid to obtain the normal vector of the fitted plane, and then calculate the angle between the plane and the horizontal plane, that is, the tilt angle of the terrain. The larger the value, the steeper the slope.
[0061] Ground adhesion coefficient estimation ( ): By combining the slip detection information from the wheel speedometer and IMU with the classification of ground texture (such as mud, sand, grass, hard surface) by vision / point cloud, the maximum static friction coefficient between the wheel and the ground is estimated. The higher the value, the less slippery the ground (such as dry concrete), and the lower the value, the more slippery or soft the ground (such as mud or snow).
[0062] Subsequently, the module assigns the aforementioned feature values to preset weights (e.g., ...). Weighted fusion is performed, and a remaining power factor can be optionally added. The local terrain accessibility of the local terrain block is calculated. . It is a dimensionless score, usually ranging from 0 to 1. The higher the value, the more suitable it is for fast and low-energy passage.
[0063] In some embodiments, ; in: This refers to the accessibility of local terrain. For estimation of ground adhesion coefficient; The standard deviation of local height fluctuations; For local slope; The maximum slope that the robot can safely traverse; This is the current remaining battery level; Rated power; These are the weighting coefficients for each item.
[0064] Then the module will With two preset thresholds (first threshold) For example, 0.7; second threshold By comparing values such as 0.3, adaptive switching of motion modes can be achieved. when When the terrain is deemed to be good for travel (such as flat hard roads and dry short grass), it switches to wheel mode, moving quickly using only the drive wheels, with the leg units retracted, resulting in the lowest energy consumption and the fastest speed.
[0065] when When the terrain is determined to be medium (such as slightly undulating field ridges, soft but non-steep dirt roads, or shallow grass areas), switch to the wheel-leg hybrid mode. The drive wheels remain the main force, but the leg units extend appropriately to adjust the ground clearance, increase support points, or assist in attitude balance in order to improve adhesion and passability.
[0066] when If the terrain is deemed to be severe (such as deep ditches, high embankments, steep slopes, muddy pits, or areas entangled with tall grass), switch to walking obstacle crossing mode. The drive wheels are locked or follow the movement, and the vehicle relies entirely on the stepping motion of the leg units (similar to a quadrupedal gait) to cross obstacles and move, avoiding getting stuck or bottoming out.
[0067] This approach transforms complex terrain point cloud data into quantified local terrain accessibility scores and, based on two thresholds, subdivides movement modes into three states: wheeled, wheel-legged hybrid, and walking obstacle crossing. This achieves full-scene adaptive coverage from high-speed flat roads to extreme obstacles. Compared to a simple binary judgment of obstacle crossing, this solution can smoothly switch modes according to continuous terrain changes, avoiding energy waste and mechanical shock caused by frequent switching. Furthermore, by introducing three independent dimensions—local height fluctuation standard deviation, local slope, and ground adhesion coefficient estimation—into accessibility calculations, it fully reflects the ruggedness, slope, and adhesion characteristics of agricultural terrain, making mode switching decisions more scientific and robust. Ultimately, the robot can automatically select the most efficient and safest movement mode in unstructured environments such as orchards and farmland, significantly improving the coverage efficiency, terrain clearance, and endurance of bird control operations.
[0068] Optionally, performing voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source location includes: The multi-channel acoustic audio is framed, denoised, and subjected to short-time Fourier transform to extract the decibel power spectrum; Based on the Mel frequency cepstral coefficients, the voiceprint embedding features are extracted from the decibel power spectrum, input into the preset bird voiceprint recognition model, and the acoustic confidence of the bird is generated. Based on the arrival time difference of each sound source in the multi-channel acoustic audio, the spatial location of the chirping sound source is determined, and the sound source location information is generated.
[0069] Specifically, the multi-channel acoustic audio (synchronously acquired by a microphone array mounted on the robot body) undergoes preprocessing. Specifically, the audio signal is divided into frames of fixed duration, and spectral subtraction or Wiener filtering is used for noise reduction to suppress wind noise, mechanical noise, and human voice interference. Subsequently, a short-time Fourier transform is performed on each frame to obtain the frequency domain amplitude spectrum, which is then converted into a decibel power spectrum. Based on this, Mel-frequency cepstral coefficients (MFCCs) are extracted as voiceprint embedding features. This feature is input into a pre-trained bird voiceprint recognition model (e.g., a classifier based on a convolutional neural network or recurrent neural network). The model outputs a probability value, namely the bird acoustic confidence score, which indicates the likelihood that the current audio segment belongs to the call of the target bird.
[0070] Simultaneously, using at least three non-collinearly arranged microphone units in the microphone array, the spatial orientation of the call source is estimated using the Time Difference of Arrival (TDOA) algorithm. Specifically, the time difference between arrival times of the same call event at each microphone is detected (e.g., calculated using generalized cross-correlation). Then, based on the known geometric layout of the microphone array and the speed of sound propagation, the horizontal azimuth and pitch angles of the sound source in the robot coordinate system are calculated using the geometric intersection method, thus obtaining the sound source's orientation information. This process does not rely on visual images and can provide directional cues even before the bird enters the field of vision.
[0071] By segmenting, denoising, performing short-time Fourier transform, and extracting Mel-frequency cepstral coefficients from multi-channel acoustic audio, combined with a deep voiceprint recognition model, bird calls can be accurately identified and high-confidence probabilities can be generated even in strong environmental noise. This allows for early detection of bird activity even when visual detection is blurred or outside the field of view. Simultaneously, utilizing the time difference of arrival (TDOA) algorithm of the microphone array, the spatial location of the call source can be calculated in real time without additional sensors, providing the robot with vision-independent orientation information. The fusion of acoustic and visual confidence levels significantly improves the robustness of bird activity recognition, while the sound source location information can guide the gimbal or chassis to quickly turn towards potentially risky areas, compensating for blind spots and delays in visual perception, and achieving earlier and more reliable bird damage warnings and target localization.
[0072] Optionally, the step of generating a comprehensive utility value for each of the bird-repelling strategies based on a preset bird-damage risk score and the spatial coordinates of the risk target includes: Based on the total bird damage risk score and the spatial coordinates of the risk target, determine the expected deterrence effectiveness and safety disturbance cost of each deterrence strategy; Using Equation 1, the corresponding comprehensive utility value is generated based on the expected expulsion effectiveness, the security disturbance cost, the fitness estimate, and the energy consumption cost. ; in, Let be the combined utility value of the k-th expulsion strategy. to For adjustment coefficients, To predict the effectiveness of the eviction, This is an estimate of the fitness of the flock to the k-th dispersal strategy. The energy cost of implementing the k-th expulsion strategy, The safety disturbance cost of the k-th expulsion strategy.
[0073] Optionally, the fitness estimate is obtained in the following way: The duration of the flock's stay, return time, and approach speed after the expulsion were obtained to determine the real-time adaptation indicators at the current moment. Based on the fitness estimate and memory factor from the previous moment, the historical weighted value is obtained; based on the immediate adaptation index and the memory factor, the current feedback weighted value is obtained. The fitness estimate at the current moment is obtained based on the historical weighted value and the current feedback weighted value.
[0074] Specifically, when the total bird damage risk score is obtained... After determining the spatial coordinates of the risk target, the module first uses this information to determine the expected effectiveness of each preset expulsion strategy. and the cost of security disturbances . The higher the level or the closer the risk target is to the core crop area, the better the strategy. The higher the assigned value; Quantify the potential disturbance or damage risk of this strategy to people, livestock, vehicles, and sensitive facilities (such as greenhouses and roads). The closer the risk target is to the safety boundary, the higher the risk of the high-disturbance strategy. The larger the value, the better. Furthermore, the module also reads the current fitness estimate of the policy. (This reflects the degree to which the flock has become accustomed to the strategy; a higher value indicates greater ineffectiveness) and the energy cost required to execute the strategy. (Estimated from hardware rated power and estimated duration). Then, the overall utility value of the k-th strategy is calculated using Equation 1. .
[0075] To dynamically track the flock's adaptation to each strategy, the fitness estimation... The update is recursively performed as follows. After each expulsion task is completed, the module acquires three measurable physical quantities: dwell time (the time from the start of the expulsion action to the flock taking off and leaving), return time (the interval between the flock re-entering the monitored area after being expelled), and re-approach speed (the speed at which the flock approaches the core crop area upon returning). These quantities are normalized and weighted to obtain the current real-time adaptive index. The value ranges from 0 to 1, with a larger value indicating deeper adaptation. Then, the memory factor is used. ( (Controlling the weights of historical memory and current feedback) Update fitness estimates: ; in, For the fitness estimate of the k-th expulsion strategy at the current time (time t), This is a fitness estimate from the previous time step (after the last use of this strategy). When a certain strategy is used frequently in a short period of time and the duration of stay after each expulsion increases and the return rate accelerates, It will be relatively large, leading to An increase; conversely, if the strategy is not used for a long time or its effectiveness regresses, Will because It decreases slowly due to the decay effect. Updated The policy will be stored in the policy library and used to reduce the priority of the adapted policies when calculating the overall utility value next time, thereby achieving adaptive scheduling.
[0076] This bird control decision-making module takes the total bird damage risk score and the spatial coordinates of the risk target as input. It quantifies the expected bird control effectiveness, current fitness, energy cost, and safety disturbance cost of each strategy through a comprehensive utility value, and dynamically updates the fitness estimate, achieving adaptive suppression of bird flock behavior habits. Compared to fixed-sequence or randomly switching bird control methods, this solution proactively avoids strategies that have been adapted to by the flock, prioritizing the most effective, lowest-cost, and safest method in the current scenario. This significantly slows down the habituation process of birds to specific stimuli, maintaining long-term bird control effectiveness. Furthermore, because the fitness estimate incorporates objective feedback such as dwell time, return time, and reapproach speed, the system can continuously self-optimize during operation, adapting to changes in different plots, bird species, and seasons without human intervention, greatly improving the intelligence level and deployment flexibility of the bird control system.
[0077] Optionally, the step of fusing the visual confidence score, the acoustic confidence score, and the target location to generate a total bird damage risk score includes: Based on a preset bird risk scoring model, the visual confidence level, acoustic confidence level, and target location are fused together to generate the total bird damage risk score. The preset bird risk scoring model includes: ; in, The overall score for the bird damage risk; The recognition confidence is obtained by weighting the visual confidence and the acoustic confidence. This is an occlusion correction term generated by cross-validation of visible light images and thermal infrared images in the visual image sequence; The relative motion term of the target position toward the crop core area; This is the population density term; This is a historically high-incidence item generated based on historical eviction records; to Configurable weights.
[0078] Specifically, firstly, the visual confidence score (from small target detection results in visible light and thermal infrared images) and the acoustic confidence score (from probability values output by the voiceprint recognition model) are weighted and fused to obtain the recognition confidence score. This is used to characterize the overall probability of birds existing within the current spatiotemporal window. Secondly, cross-validation is performed using visible light images and thermal infrared images from the visual image sequence: if the target region in the visible light image is obscured by branches and leaves, resulting in a fragmented outline, while the thermal infrared image still shows a complete thermosphere shape, then the feature overlap rate and missing area of the two are calculated to generate an occlusion correction term. This is used to correct for the underestimation of confidence caused by occlusion. Simultaneously, based on the target position changes over multiple consecutive frames, the radial velocity or approach rate of the target towards the crop core area is calculated to obtain the relative motion term. ; Calculate the population density term based on the number of individual birds or the number of sound source clusters detected in the same area. Based on the historical eviction records of the plot, the time period, and the current crop stage (such as the frequency and severity of past bird-related incidents), historically high-incidence items were extracted. Finally, the five indicators mentioned above are multiplied by their respective configurable weights and then summed to obtain the overall bird damage risk score. .
[0079] This risk scoring model integrates multi-dimensional information such as visual confidence, acoustic confidence, and target location through five interpretable physical terms (identification confidence, occlusion correction, relative motion, population density, and historical high incidence). This not only quantifies the probability of bird presence but also considers the risk of missed detection due to occlusion, invasion direction and speed, population size, and long-term bird damage patterns in the plot, making the scoring more consistent with the actual threat level in agricultural scenarios. Compared to simple threshold judgments relying solely on single-frame visual detection, this approach can identify concealed targets in advance, distinguish between accidental passing and active intrusion, and adapt to the risk differences in plots at different crop stages. This significantly reduces false alarms and missed alarms, providing a robust and interpretable decision-making basis for the rational selection of subsequent repellency strategies.
[0080] Optionally, the bird control system based on the wheeled-legged robot further includes a multi-robot collaboration module, which is used to receive shared information from other robots in the same working area. The shared information includes bird flock trajectory, heat map, remaining battery power and areas where passage is obstructed. When a robot is unable to cover adjacent high-risk areas due to insufficient power, terrain limitations, or being engaged in a high-level expulsion mission, it generates a replacement request to request a neighboring robot to take over.
[0081] In some embodiments, the specific working process of the multi-machine collaborative module is as follows: In operational areas (such as large orchards or grain bases) where multiple wheeled-legged bird-repelling robots are deployed, each robot shares its operational status and environmental perception information in real time via edge stations (or direct point-to-point communication). This shared information includes: bird flock trajectories (the locations and movement paths of birds detected by each robot), heatmaps (the distribution of bird damage risk levels in different plots), remaining battery power (the current battery capacity percentage of each robot), and areas with obstructed passage (locations that are impassable due to terrain, water accumulation, crop stacking, etc.). A multi-robot collaborative module operates in each robot or edge station, continuously monitoring the shared information and evaluating its own operational capabilities. When the robot determines that one of the following situations occurs: ① the remaining battery power is below a preset threshold (e.g., 20%), it needs to return to base for charging; ② the terrain ahead is too impassable (e.g., deep ditches or steep slopes are detected), making safe entry impossible; ③ it is currently performing a high-level bird control task (e.g., laser scanning or continuous sound and light suppression) and the task is not yet completed, preventing it from attending to other areas, and at this time, it learns from shared information that there is a high-risk area adjacent to it (e.g., a flock of birds gathering in another mature crop area) but no other robot is covering it, the robot automatically generates a replacement request and sends it to the edge station or a designated coordination node, requesting that a nearby robot with available capacity and sufficient battery power be assigned to take over the bird control task in that area. The edge station performs comprehensive scheduling based on each robot's current location, remaining battery power, accessibility, and task priority, assigning the replacement task to the best candidate robot to achieve regional joint defense.
[0082] This multi-robot collaborative module achieves dynamic task allocation and vacancy takeover within large-scale planting areas through information sharing among robots (bird flock trajectories, heat map, remaining battery power, and areas with obstructed passage) and a replacement request mechanism. When a single robot is unable to cover adjacent high-risk areas due to insufficient battery power, terrain obstruction, or task saturation, the system can automatically dispatch other robots to fill the gap, avoiding bird control gaps caused by single-point failures or resource limitations. This mechanism significantly improves the robustness and coverage continuity of the entire bird control system, enabling multi-robot clusters to collaboratively respond to the spatiotemporal movement of bird flocks, making it particularly suitable for all-weather bird control in high-value scenarios such as open farmland and large orchards.
[0083] Optionally, the bird control system based on the wheeled-legged robot further includes an edge-cloud collaboration module, which is used to upload sample data and receive updated model parameters and strategy knowledge graphs when communicating with the cloud.
[0084] In some embodiments, the edge-cloud collaboration module operates as follows: Each wheel-legged bird-repelling robot has a built-in edge computing unit that processes perception data and executes repelling decisions in real time. Simultaneously, the edge-cloud collaboration module maintains communication with the cloud platform in the background (when network access is available). The module periodically packages and uploads local sample data (including newly acquired bird images, voiceprint fragments, records of successful or failed repelling events, and anomaly logs) to the cloud. The cloud server aggregates data from multiple robots and multiple sites, updating model parameters (e.g., weights of the small object detection network, feature extraction layers of the voiceprint recognition model, and weight coefficients in the risk scoring model) through offline training or federated learning. to and the adjustment coefficient in the utility function to (etc.) Simultaneously, statistical learning methods are used to construct a strategy knowledge graph, which records the most effective combinations of repelling strategies and their historical success rates for different plots, crop stages, bird species, and time periods. When the robot completes its patrol task in the current plot and moves to a new plot, or enters a new crop growth stage in the same plot, the edge-cloud collaboration module pulls updated model parameters and the strategy knowledge graph for the current scenario from the cloud and automatically loads them into the local decision base, thereby achieving cross-plot knowledge transfer and strategy preheating.
[0085] This edge-cloud collaborative module enables continuous evolution and cross-scenario migration of model parameters and policy knowledge graphs through bidirectional data synchronization between the robot's local machine and the cloud. On one hand, the rich sample data gathered in the cloud allows the deep learning model to become more accurate with use, continuously improving the accuracy of small target detection and voiceprint recognition. On the other hand, the policy knowledge graph supports newly deployed or transitioned robots to quickly gain mature experience, significantly shortening the cold start cycle. Simultaneously, even when the network is offline, the robot can still rely on local cached decisions, performing incremental synchronization after the network is restored, balancing real-time response and long-term optimization. Therefore, the entire bird-repelling system possesses self-learning, self-evolution, and cross-plot generalization capabilities, significantly reducing manual configuration and maintenance costs, and improving adaptability and bird-repelling efficiency in different agricultural scenarios.
[0086] Optionally, the execution control module generates the chassis motion commands, specifically including: Based on the expulsion mission command and the current motion mode, the desired body motion state is determined, including the desired body speed, the desired body posture, and the desired leg support force, and the desired body motion state is used as the content of the chassis motion command. Based on the whole body control law, the desired body motion state in the chassis motion command is converted into the output torque vector of each actuator in the wheel-legged robot to drive the wheel drive wheel and the leg lifting mechanism to execute the chassis motion command; The whole-body control law is as follows: ; in, J is the output torque vector; J is the robot's Jacobian matrix. For expected support; and These represent the desired joint state and the current joint state, respectively. and These are the expected speed and the current speed, respectively. and This is the gain matrix.
[0087] Specifically, upon receiving the obstacle removal task command (including the obstacle removal method, execution intensity, and area of effect) and determining the current motion mode (wheeled, wheel-legged hybrid, or walking obstacle crossing), the execution control module first plans a desired motion trajectory that enables the robot to safely and efficiently approach the risky target and perform the obstacle removal action. This trajectory is converted into the desired body motion state, specifically including: the desired body velocity (linear velocity and angular velocity, which determine the speed and direction of the robot's movement), the desired body attitude (pitch angle, roll angle, and yaw angle, ensuring stable gimbal perception and obstacle removal direction), and the desired leg support force (the ideal force value of each leg unit in contact with the ground to ensure anti-slip and balance). These three factors together constitute the content of the chassis motion command.
[0088] Subsequently, the module uses a whole-body control law to convert the desired motion state into the actual output torque vectors of each actuator (including the left / right drive wheel motors and the lifting / swinging motors of each leg joint). This drives the wheel drive wheels and the leg lifting mechanism to move in a coordinated manner, precisely executing the chassis movement commands.
[0089] The first of the laws of whole-body control The second item is responsible for achieving the desired distribution of ground support (e.g., increasing foreleg support on a slope to prevent slippage). Correcting joint position deviation, third item Compensating for speed deviations and increasing system damping ensures that the robot moves smoothly and accurately along the desired trajectory.
[0090] By uniformly planning the desired body motion states (velocity, posture, leg support force) and combining a whole-body control law based on Jacobian matrix and PID feedback, high-level bird deterrence task commands are transformed into precise torque distribution for each actuator at the lower level. Compared to decentralized methods that control wheels or legs individually, the whole-body control law can collaboratively optimize the force and motion of wheels and legs, ensuring that the robot maintains stability and precise gimbal pointing even on complex terrains (such as slopes, soft soil, and ditches), and achieves smooth torque transitions when switching motion modes. As a result, the execution of bird deterrence actions is more robust, and the visual and acoustic perception modules will not fail due to body vibration, thus significantly improving the success rate of bird deterrence tasks and the overall reliability and safety of the system.
[0091] In some embodiments, the expulsion decision module is also specifically used to: automatically decide to upgrade to the next level of strategy when the tolerance time of the flock exceeds a preset threshold or the return time is shortened after the current level strategy is executed; and to forcibly maintain or downgrade to a lower level strategy when the safety boundary is triggered or the flock is expelled in a very short time.
[0092] Specifically, the process by which the deterrence decision module upgrades or downgrades the deterrence strategy is as follows: After executing a certain level of strategy (such as L2 directional acoustics), the module continuously monitors two key indicators: tolerance time (the time from the start of the deterrence action to the actual take-off and departure of the bird flock from the core crop area) and return time (the interval between the bird flock re-entering the monitored area after being deterred). When the tolerance time exceeds a preset threshold (e.g., 3 seconds) or the return time is significantly shortened compared to the historical average (e.g., the current return time is only half of the previous one), it is determined that the deterrence effect of the current strategy on the bird flock has significantly decreased, and the bird flock is adapting. The module automatically decides to upgrade the deterrence intensity to the next level (e.g., from L2 to L3 strobe reflective strategy) to re-establish the deterrent effect. Conversely, when a safety boundary is triggered (e.g., a person or vehicle is suddenly detected entering the de-encroachment area during execution), the module immediately forces a downgrade to a low-disturbance strategy (e.g., downgrade from L3 to L1 low-disturbance patrol) or suspends the de-encroachment to ensure safety; or when a flock of birds is driven away by the current strategy in a very short time (e.g., less than 1 second), it indicates that the low-intensity strategy is effective enough, and the module will maintain the current level or even try to downgrade to a lower level in subsequent similar situations to save energy and reduce environmental disturbance.
[0093] This upgrade / downgrade mechanism uses tolerance time and return time as quantitative feedback to dynamically sense the flock's adaptation to the current strategy. It automatically upgrades the deterrence intensity when the effect diminishes and proactively downgrades when safety risks arise or low intensity is sufficient, achieving intelligent strategy scheduling. Compared to fixed-sequence or blindly cyclical deterrence methods, this mechanism effectively slows down the habituation process of the flock, ensuring deterrence effectiveness while avoiding excessive environmental disturbance and energy waste, and always prioritizing safety. This significantly improves the long-term effectiveness, environmental friendliness, and operational safety of the bird deterrence system.
[0094] In some embodiments, based on the above-described implementations, the system further includes the following optimization mechanisms, which can be implemented individually or in combination: (i) Precise calculation of occlusion correction items: In the risk assessment module, the occlusion correction items...
[0095] Cross-validation using visible light and thermal infrared images was employed: visible light and thermal infrared frames acquired at the same time were registered, and the bird target region was extracted separately; the overlap of feature edges within the two regions and the area ratio of the thermal infrared-specific region were calculated. If the effective texture area within the detection box in the visible light image is lower than the pre-defined area due to foliage occlusion, the cross-validation was performed. If a threshold is set (e.g., 50%), and the thermal infrared image still shows a complete thermosphere outline, then a model is generated based on the proportion of the missing area and the shape integrity of the significant thermal infrared region. Value (ranging from 0 to 1, with larger values indicating less occlusion); if the target regions of the two modalities highly overlap, then It is close to 0. This correction term is used in bird risk scoring to compensate for the underestimation of visual confidence caused by occlusion.
[0096] (ii) Quantification formula for immediate adaptation indicators: After each expulsion task is completed, the expulsion decision module collects three physical quantities: duration of stay. The time from the start of the dispersal action to the flock taking off and leaving the core crop area, and the return time. It refers to the number of seconds between the birds being driven away and re-entering the monitored area, as well as their speed when they approach again. This represents the speed at which the flock of birds returns towards the core area of the crop. After normalizing the above three quantities, the immediate adaptation index is calculated using the following formula. : ; in, The preset maximum allowable tolerance time (e.g., 5 seconds). This represents the historical average return time for bird flocks to this area. This represents the maximum approach speed of birds (e.g., 3 m / s). For weights, such as , , . The larger the value, the more deeply the flock has adapted to the current strategy.
[0097] (III) Dynamic setting of memory factors: the memory factors (0<ρ<1) Dynamically adjust based on the biological forgetting curve and historical invasion frequency of the target bird species: For bird species with strong memory and frequent visits in a short period of time (such as magpies and crows), set... This allows fitness to decay slowly, avoiding the blind repetition of failed strategies; for bird species with low-frequency, occasional visits, a set... This gives the strategy a chance to be retried; it can also be manually adjusted through the edge station configuration interface.
[0098] (iv) Dual safety interlocks for high-level expulsion such as lasers: For L4 level laser scanning strategies, the system performs dual safety interlock verification: Software layer verification: Before generating the drive-away task command, check the global no-fire zone electronic fence and confirm that there are no people, vehicles, livestock, or glass or greenhouse reflective surfaces in the laser scanning direction at the current moment in the visual recognition results; Hardware layer verification: The enable switch of the laser emission circuit is directly connected in series with the hardware signal of the gimbal's physical attitude sensor. The circuit is only allowed to be powered on when the gimbal's pitch angle is higher than the hard safety threshold (e.g., the horizontal plane angle > 15°) and the gimbal's horizontal rotation angle is within the allowable sector.
[0099] Light can only be emitted if both verifications pass; any unauthorized attempts will trigger an emergency electrical shutdown at the underlying level.
[0100] (v) Offline data synchronization and conflict resolution: The edge-cloud collaboration module enters offline autonomous navigation mode when the network is disconnected. High-value samples (new bird flock images, failed dispersal cases, security interception records) and policy fitness changes generated during the offline period are cached in local non-volatile memory. Synchronization is performed after the network is restored. For global configurations (map updates, addition of no-fire zones, model parameters), the latest version in the cloud will be overwritten locally. For locally collected patrol logs and strategy fitness increments, an incremental merging method is used to upload them to the cloud. After receiving them, the cloud server performs deduplication and weight fusion.
[0101] (vi) Time interpolation method in torque smooth transition: When switching between wheeled mode and walking obstacle-crossing mode, to avoid impact caused by sudden changes in joint torque, the expected support force in the whole-body control law is required. Expected joint status and expected speed A polynomial transition curve is used for constraint solving. (During the switching time window...) Within this process, the target values of the aforementioned variables are smoothly interpolated from the pre-switching state to the post-switching target state. The interpolation function uses a fifth-degree polynomial. The interpolated sequence is used as the input to the control law, ensuring the output torque of each actuator. This interpolation satisfies the continuity conditions for boundary position, velocity, and acceleration. Smooth and continuous.
[0102] (vii) Sandbox validation and negative migration avoidance for strategy migration: When a robot is deployed on a new plot of land and downloads a policy knowledge graph from the cloud, the edge-cloud collaboration module initiates a sandbox verification mode: it assigns a low initial execution weight to the migration strategy (e.g., the human factor in the overall utility value is only 30% of the normal value) and monitors the actual local eviction effect frequently (through tolerance time and relapse rate). If the strategy effect deviates significantly from the cloud record in three consecutive evictions (the local tolerance time is more than twice the standard deviation of the cloud record), the weight of the migration strategy is automatically reduced, and the focus shifts to the fitness parameters accumulated through local reinforcement learning, until the local statistical data reaches a confidence level before gradually releasing the use of the migration strategy.
[0103] (viii) Structural variations of the leg unit in different agricultural terrains: For different agricultural scenarios, the leg unit of the wheel-legged robot can be equipped with the following structural variations: Multi-degree-of-freedom triangular leg unit: Composed of vertical linear actuator I, horizontal linear actuator II, and damping spring, it is suitable for complex structural scenarios that require large-scale lifting across steps and ditches (such as greenhouse entrances and water channels). Arc-shaped connecting legs: In areas where fragile fruits such as strawberries and mulberries are grown, the leg units adopt an arc-shaped bending design, using their own elasticity to absorb impact and avoid damaging the fruit and the covering film. Independent steering leg unit: Each leg unit is equipped with an independent motor-driven steering column at its upper end, which can rotate around the vertical axis. In conjunction with the differential drive wheel, it can turn around on the spot and move laterally, improving the ability to get out of trouble in muddy and tall grass areas.
[0104] All of the above variants are compatible with the whole-body control law, requiring only modifications to the Jacobian matrix dimension and kinematic constraints.
[0105] (ix) Crop height threshold setting and amplitude limiting algorithm: To prevent the gimbal or its upper components from impacting the fruit and branches, the execution control module senses the crop height in real time: using semantic segmentation of lidar point clouds and visible light depth cameras, it dynamically estimates the average canopy height of the crops on both sides in front. and increase safety redundancy. (e.g., 10cm) to obtain the dynamic height threshold The threshold is mapped to the angle hard constraints of the gimbal joint and the upper frame swing arm through inverse kinematics. When generating gimbal pointing commands, if the planned pitch angle or extension attitude will cause a component to intrude into the height envelope, the limiting algorithm will automatically truncate the over-limit command and trigger an emergency stop log recording; at the same time, the chassis control can automatically reduce the driving speed or increase the vehicle's ground clearance according to the crop height.
[0106] like Figure 2 As shown, this embodiment of the invention provides a bird deterrence control method based on a wheeled-legged robot, applied to the bird deterrence control system based on a wheeled-legged robot as described above. The bird deterrence control method based on a wheeled-legged robot includes: Collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area; Small target detection is performed on the visual image sequence to generate bird visual confidence and target location; voiceprint recognition and sound source localization are performed on the multi-channel acoustic audio to generate bird acoustic confidence and sound source orientation; and the visual confidence, acoustic confidence, and target location are fused to generate a total bird damage risk score; the target location is verified or supplemented based on the sound source orientation to generate the spatial coordinates of the risk target. Based on the preset bird deterrence strategy, a comprehensive utility value corresponding to each deterrence strategy is generated according to the total bird damage risk score and the spatial coordinates of the risk target. The deterrence strategy with the highest comprehensive utility value and that meets the preset safety boundary is selected, and a deterrence task instruction is generated. The deterrence task instruction includes deterrence method, execution intensity and area of action. The local terrain accessibility is determined based on the terrain point cloud, and the motion mode of the wheeled robot is adaptively switched. Based on the motion mode, gimbal pointing instructions, acousto-optical-laser actuator drive signals, and chassis motion instructions are generated according to the expulsion task instructions to control the wheeled robot to approach the risk target and perform graded expulsion actions.
[0107] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A bird deterrence control system based on a wheeled-legged robot, characterized in that, include: The data acquisition module is used to collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area; The risk assessment module is used to perform small target detection on the visual image sequence to generate bird visual confidence and target location, perform voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source location, and perform fusion processing based on the visual confidence, the acoustic confidence and the target location to generate a total bird damage risk score. The target location is verified or supplemented based on the sound source location to generate the spatial coordinates of the risk target; The bird deterrence decision module is used to generate a comprehensive utility value for each bird deterrence strategy based on a preset deterrence strategy, according to the total bird damage risk score and the spatial coordinates of the risk target, select the deterrence strategy with the highest comprehensive utility value and that meets the preset safety boundary, and generate a deterrence task instruction, wherein the deterrence task instruction includes deterrence method, execution intensity and area of action. The execution control module is used to determine the local terrain accessibility based on the terrain point cloud and adaptively switch the motion mode of the wheeled robot; and based on the motion mode, it generates gimbal pointing instructions, acousto-optical-laser actuator drive signals and chassis motion instructions according to the expulsion task instructions, so as to control the wheeled robot to approach the risk target and perform graded expulsion actions. The method of generating a comprehensive utility value for each of the bird-repelling strategies based on the preset bird-damage risk score and the spatial coordinates of the risk target includes: Based on the total bird damage risk score and the spatial coordinates of the risk target, determine the expected deterrence effectiveness and safety disturbance cost of each deterrence strategy; The corresponding comprehensive utility value is generated based on the expected expulsion effectiveness, the security disturbance cost, the fitness estimate, and the energy consumption cost: ; in, Let be the combined utility value of the k-th expulsion strategy. to For adjustment coefficients, To predict the effectiveness of the eviction, This is an estimate of the fitness of the flock to the k-th dispersal strategy. The energy cost of implementing the k-th expulsion strategy, The safety disturbance cost of the k-th expulsion strategy.
2. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, The movement modes include wheeled mode, walking obstacle-crossing mode, and wheel-legged hybrid mode; the step of determining local terrain accessibility based on the terrain point cloud and adaptively switching the movement mode of the wheel-legged robot includes: The terrain point cloud is projected onto a horizontal plane and divided into grids, with each grid serving as a local terrain block. The local height variation standard deviation, local slope, and ground adhesion coefficient are estimated based on the point cloud data in each grid cell for the corresponding local terrain block. The local elevation fluctuation standard deviation, the local slope, and the ground adhesion coefficient estimate are fused together to obtain the local terrain accessibility corresponding to the local terrain block. When the terrain accessibility is higher than the first threshold, switch to wheel mode; When the terrain accessibility is between the first threshold and the second threshold, switch to the wheel-leg hybrid mode; When the terrain accessibility is below the second threshold, switch to walking obstacle crossing mode.
3. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, The step of performing voiceprint recognition and sound source localization on the multi-channel acoustic audio to generate bird acoustic confidence and sound source location includes: The multi-channel acoustic audio is framed, denoised, and subjected to short-time Fourier transform to extract the decibel power spectrum; Based on the Mel frequency cepstral coefficients, the voiceprint embedding features are extracted from the decibel power spectrum, input into the preset bird voiceprint recognition model, and the acoustic confidence of the bird is generated. Based on the arrival time difference of each sound source in the multi-channel acoustic audio, the spatial location of the chirping sound source is determined, and the sound source location information is generated.
4. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, The fitness estimate is obtained as follows: The duration of the flock's stay, return time, and approach speed after the expulsion were obtained to determine the real-time adaptation indicators at the current moment. Based on the fitness estimate and memory factor from the previous moment, the historical weighted value is obtained; based on the immediate adaptation index and the memory factor, the current feedback weighted value is obtained. The fitness estimate at the current moment is obtained based on the historical weighted value and the current feedback weighted value.
5. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, The process of fusing the visual confidence score, the acoustic confidence score, and the target location to generate a total bird damage risk score includes: Based on a preset bird risk scoring model, the visual confidence level, acoustic confidence level, and target location are fused together to generate the total bird damage risk score. The preset bird risk scoring model includes: ; in, The overall score for the bird damage risk; The recognition confidence is obtained by weighting the visual confidence and the acoustic confidence. This is an occlusion correction term generated by cross-validation of visible light images and thermal infrared images in the visual image sequence; The relative motion term of the target position toward the crop core area; This is the population density term; This is a historically high-incidence item generated based on historical eviction records; to Configurable weights.
6. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, It also includes a multi-robot collaboration module, which is used to receive shared information from other robots in the same working area. The shared information includes bird flock trajectories, heat map, remaining battery power, and areas where passage is obstructed. When a robot is unable to cover adjacent high-risk areas due to insufficient power, terrain limitations, or being engaged in a high-level expulsion mission, it generates a replacement request to request a neighboring robot to take over.
7. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, It also includes an edge-cloud collaboration module, which is used to upload sample data and receive updated model parameters and policy knowledge graphs when communicating with the cloud.
8. The bird control system based on a wheeled-legged robot according to claim 1, characterized in that, The execution control module generates the chassis motion commands, specifically including: Based on the expulsion mission command and the current motion mode, the desired body motion state is determined, including the desired body speed, the desired body posture, and the desired leg support force, and the desired body motion state is used as the content of the chassis motion command. Based on the whole-body control law, the desired body motion state in the chassis motion command is converted into the output torque vector of each actuator in the wheel-legged robot to drive the wheel drive wheel and the leg lifting mechanism to execute the chassis motion command; The whole-body control law is as follows: ; in, is the output torque vector; J is the Jacobian matrix of the wheel-legged robot; For expected support; and These represent the desired joint state and the current joint state, respectively. and These are the expected speed and the current speed, respectively. and This is the gain matrix.
9. A bird-deterrence control method based on a wheeled-legged robot, characterized in that, The bird deterrence control system based on a wheeled-legged robot as described in any one of claims 1 to 8, wherein the bird deterrence control method based on the wheeled-legged robot comprises: Collect visual image sequences, multi-channel acoustic audio, terrain point clouds, and robot pose information of the work area; Small target detection is performed on the visual image sequence to generate bird visual confidence and target location; voiceprint recognition and sound source localization are performed on the multi-channel acoustic audio to generate bird acoustic confidence and sound source orientation; and the visual confidence, acoustic confidence, and target location are fused to generate a total bird damage risk score; the target location is verified or supplemented based on the sound source orientation to generate the spatial coordinates of the risk target. Based on the preset bird deterrence strategy, a comprehensive utility value corresponding to each deterrence strategy is generated according to the total bird damage risk score and the spatial coordinates of the risk target. The deterrence strategy with the highest comprehensive utility value and that meets the preset safety boundary is selected, and a deterrence task instruction is generated. The deterrence task instruction includes deterrence method, execution intensity and area of action. The local terrain accessibility is determined based on the terrain point cloud, and the motion mode of the wheeled robot is adaptively switched. Based on the motion mode, gimbal pointing instructions, acousto-optical-laser actuator drive signals, and chassis motion instructions are generated according to the expulsion task instructions to control the wheeled robot to approach the risk target and perform graded expulsion actions.
Citation Information
Patent Citations
Active positioning bird repelling system based on vision and radar fusion
CN121955975A
Self-adaptive airport bird repelling method and system based on unmanned aerial vehicle and storage medium
CN122090238A