Endangered animal whereabouts monitoring device and method

Through infrared thermal imaging and multi-agent reinforcement learning, the drone inspection path is optimized, and the problem of insufficient intelligence and robustness of the drone monitoring system in the monitoring of endangered animals is solved, achieving efficient and stable monitoring effects.

CN120411175BActive Publication Date: 2025-09-05CHINA CRIMINAL POLICE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510906375.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-05
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing drone monitoring systems are not very intelligent in the monitoring of whereabouts of endangered animals, and are easily restricted by communication and energy during coordinated operations, resulting in insufficient monitoring efficiency and robustness, especially in long-term, large-scale and refined monitoring scenarios.

Method used

Infrared thermal imaging devices are used to identify animal targets, combine multi-agent reinforcement learning models to optimize patrol paths, adaptively adjust the coordinated flight actions of the drone, and perform fault-tolerant control and dynamic task reallocation when communication is interrupted or energy is exhausted, improving the intelligence and robustness of the monitoring system.

Benefits of technology

It improves the detection and identification capabilities and tracking stability of endangered animals, enhances the initiative of monitoring tasks and the ability to discover abnormal behaviors, ensures the long-term and continuity of monitoring data, and provides more reliable technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411175B_ABST
    Figure CN120411175B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision and artificial intelligence technology, and discloses an endangered animal whereabouts monitoring device and method, wherein a method for monitoring the whereabouts of endangered animals includes: collecting infrared thermal imaging data, tracking potential animal targets, and obtaining target tracking information; generating a collaborative patrol path optimization strategy for a group of drones based on real-time status perception information of each drone in a group and communication topology information between drones; adaptively adjusting the execution of the collaborative patrol path optimization strategy; when some drones in a group of drones fail due to communication interruption or energy exhaustion, adjusting the tasks of the remaining drones according to a preset fault-tolerant control logic and a dynamic task reallocation model; the present invention optimizes the patrol strategy through autonomous learning of an intelligent agent, thereby enhancing the initiative of the monitoring task and the ability to discover abnormal animal behavior and potential risk areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and more particularly, to an endangered animal whereabouts monitoring device and method. Background Art

[0002] Effectively monitoring the movements of endangered animals is fundamental to their population conservation and ecological research. Traditional monitoring methods, such as ground patrols and fixed-point camera traps, suffer from low coverage, high labor costs, and poor real-time performance when dealing with animals with wide ranges and secretive behaviors, especially those that are nocturnal.

[0003] In recent years, drone remote sensing technology has been increasingly used in wildlife monitoring due to its flexibility and accessibility. Drones equipped with infrared thermal imaging devices, in particular, are capable of detecting animal targets at night or under conditions obscured by vegetation. However, existing drone-based monitoring methods still have several shortcomings: First, most drone inspections rely on preset paths or simple manual planning, lacking the ability to intelligently adapt to environmental dynamics and animal activity patterns, resulting in low search efficiency and difficulty in proactively discovering new activity areas or abnormal behaviors. Second, when multiple drones are used to collaborate to improve coverage and efficiency, wireless communications in complex outdoor environments are often unstable, and coordination between drones is easily affected. At the same time, uneven energy consumption among different drones may cause some key units to exit the mission prematurely, affecting the continuity of monitoring.

[0004] These factors together make it difficult for existing drone monitoring systems to meet the high demands of endangered animal movement monitoring in terms of intelligence, robustness and mission continuity. Especially in scenarios that require long-term, large-scale and sophisticated monitoring, more advanced technical means are urgently needed to solve the above problems. Summary of the Invention

[0005] The present invention provides an endangered animal whereabouts monitoring device and method, which solves the technical problems in related technologies such as low intelligence level of drone monitoring, susceptibility to communication and energy limitations during collaborative operations, and insufficient monitoring efficiency and robustness.

[0006] The present invention provides a method for monitoring the whereabouts of endangered animals, comprising:

[0007] The infrared thermal imaging device carried by each drone in a group of drones deployed in the monitoring area collects infrared thermal imaging data, identifies potential animal targets based on the infrared thermal imaging data, tracks the potential animal targets, and obtains target tracking information;

[0008] Based on target tracking information, a multi-agent reinforcement learning model is used to generate a collaborative inspection path optimization strategy for a group of drones according to the real-time state perception information of each drone in a group and the communication topology information between drones;

[0009] Based on the collaborative inspection path optimization strategy, the execution of the collaborative inspection path optimization strategy is adaptively adjusted according to the real-time communication quality between each drone in a group and the real-time energy level of each drone;

[0010] When some drones in a group fail due to communication interruption or energy exhaustion, the tasks of the remaining drones are adjusted based on the real-time status perception results, according to the preset fault-tolerant control logic and dynamic task reallocation model.

[0011] Furthermore, the infrared thermal imaging data identifies potential animal targets, including:

[0012] An adaptive threshold segmentation algorithm is used to process infrared thermal imaging data and preliminarily extract potential animal heat source areas.

[0013] Furthermore, the potential animal target is tracked to obtain target tracking information, including:

[0014] A Kalman filter algorithm or a particle filter algorithm is used to estimate and predict the motion state of the initially extracted potential animal heat source area to achieve stable tracking, and to record the real-time position coordinate sequence and motion parameters of the animal target.

[0015] Furthermore, the multi-agent reinforcement learning model is a multi-agent deep deterministic policy gradient model;

[0016] The real-time state perception information includes the position, speed, and remaining energy of each drone; the local environment information includes at least the heat source target information within the current field of view;

[0017] The communication topology information between the drones is modeled through a graph neural network to determine the connection relationship and communication quality between the drones.

[0018] Furthermore, the generation of a collaborative inspection path optimization strategy for a group of drones includes:

[0019] Each agent in the MADDPG model corresponds to one drone in a group of drones, and each agent is configured with an Actor network and a Critic network;

[0020] The Actor network outputs the UAV’s flight actions based on the real-time state information, local environment information, and communication topology information of the intelligent agent;

[0021] The critic network evaluates the value of the joint state and joint actions of all agents;

[0022] The Actor network and Critic network are trained by maximizing a globally shared cumulative reward function, which comprehensively considers new target discovery, area coverage, energy consumption, and communication connection quality.

[0023] Furthermore, the adaptively adjusting the execution of the collaborative inspection path optimization strategy according to the real-time communication quality between each drone in a group of drones and the real-time energy level of each drone includes:

[0024] When it is monitored that the communication quality between drones drops below a preset threshold or the real-time energy level of a drone is lower than a preset energy threshold, the flight action of the drone or its role in the collaborative mission is adjusted to prioritize ensuring the stability of the communication link and avoiding excessive energy consumption.

[0025] Furthermore, when some UAVs in a group fail due to communication interruption or energy exhaustion, the tasks of the remaining UAVs are adjusted based on the real-time state perception results, according to the preset fault-tolerant control logic and dynamic task reallocation model, including:

[0026] When a drone experiencing communication interruption fails to restore communication within a preset time threshold, the drone automatically switches to an emergency flight mode based on local perception;

[0027] When a drone responsible for a critical subtask is detected to have failed, the dynamic task reallocation model is used to reallocate the failed drone's tasks to one or more remaining healthy drones based on the current global task priority and the status of the remaining healthy drones.

[0028] Furthermore, the dynamic task reallocation model is a pre-trained fast decision network, which takes the task information of the failed drone, the status information of all healthy drones and the environmental information as input and outputs a task succession plan.

[0029] Furthermore, the infrared thermal imaging data is data collected at night or under low light conditions.

[0030] The present invention provides an endangered animal whereabouts monitoring device, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned endangered animal whereabouts monitoring method.

[0031] The beneficial effects of the present invention are: improving the detection and identification capabilities and tracking stability of endangered animals at night and in complex environments;

[0032] Through autonomous learning, the intelligent agent optimizes patrol strategies, enhancing the proactiveness of monitoring tasks and the ability to detect abnormal animal behavior and potential risk areas;

[0033] Through adaptive management of communications and energy and effective fault-tolerant redistribution, the overall operational robustness and mission continuity of the drone swarm under harsh conditions in the wild are improved, ensuring the long-term and continuity of monitoring data, and providing more reliable and efficient technical support for the protection of endangered animals. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a flow chart of a method for monitoring the whereabouts of endangered animals in the present invention;

[0035] Figure 2 is a flow chart of step 1 in the present invention;

[0036] Figure 3 It is a flow chart of step 2 in the present invention;

[0037] Figure 4 It is a flow chart of step 3 in the present invention;

[0038] Figure 5 It is a flow chart of step 4 in the present invention. DETAILED DESCRIPTION

[0039] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0040] At least one embodiment of the present invention discloses a method for monitoring the whereabouts of endangered animals, such as Figures 1 to 5 Shown, including:

[0041] Step 1: The infrared thermal imaging device carried by each drone in a group of drones deployed in the monitoring area collects infrared thermal imaging data, identifies potential animal targets based on the infrared thermal imaging data, tracks the potential animal targets, and obtains target tracking information;

[0042] This step aims to effectively detect and continuously track endangered animal targets at night or in low-light conditions through infrared thermal imaging devices carried by drone clusters.

[0043] Step 1.1, capturing thermal radiation image data of the ground environment;

[0044] Deploy a fleet of drones, each equipped with an infrared thermal imaging module. The drones fly over a pre-set or dynamically adjusted inspection area, while the infrared thermal imaging modules capture real-time thermal radiation image data from the ground environment.

[0045] Input data types: inspection area geographic information, drone initial flight instructions;

[0046] Specific output results: original infrared image sequences containing ground thermal radiation information collected by each drone in real time.

[0047] Step 1.2, processing the infrared image sequences collected by each UAV;

[0048] An adaptive threshold segmentation algorithm is applied to the infrared image sequences collected by each drone to preliminarily extract potential animal heat source targets. It should be noted that the adaptive threshold segmentation algorithm can, for example, employ the Otsu algorithm (maximum inter-class variance method) or an adaptive threshold method based on local statistical features to effectively distinguish the target heat source from the background environment. Alternatively, in some embodiments, a segmentation strategy assisted by region growing or edge detection can be employed to accommodate infrared images with varying background complexity.

[0049] Input data type: raw infrared image sequence;

[0050] Specific output results: binary image or target mask of potential animal heat source areas marked in the infrared image.

[0051] Step 1.3, stably tracking the initially extracted potential animal heat source targets;

[0052] An optimized target tracking algorithm is used to stably track the initially extracted potential animal heat source targets. The target tracking algorithm can be implemented based on Kalman filtering through the following state prediction equation and observation update equation:

[0053] State prediction equation:

[0054] ;

[0055] in Indicates in The target state vector at time , Indicates in The target state vector at time t; is the state transition matrix (describing the target state from arrive Linear transformation relationship of ); is the control input matrix (the control input Mapping to state space); For The control input vector at time t; is the process noise (usually assumed to be Gaussian white noise with zero mean, reflecting system modeling errors and external disturbances);

[0056] Observation update equation:

[0057] ;

[0058] in For The observation vector at time instant; is the observation matrix (mapping the state vector to the observation space); is the observation noise (usually assumed to be Gaussian white noise with zero mean, reflecting sensor measurement error); Indicates in The target state vector at time t.

[0059] Alternatively, in some implementations, particle filtering (a Bayesian filtering method based on importance sampling) or deep learning-based target trackers (such as lightweight versions of the SiamRPN family of algorithms) can be employed. These are particularly suitable for scenarios with complex target motion patterns or frequent occlusions. These algorithms can record the position, morphology, and basic behavioral data (such as movement speed and direction) of animal targets.

[0060] Input data type: Binarized image or target mask marked with potential animal heat source areas, historical target status information;

[0061] Specific output results: identity identification (temporary), real-time position coordinate sequence, movement speed, movement direction and other structured data of the animal target being continuously tracked.

[0062] Step 2: Based on the target tracking information, a multi-agent reinforcement learning model is used to generate a collaborative inspection path optimization strategy for a group of drones according to the real-time state perception information of each drone in the group and the communication topology information between drones;

[0063] This step uses the target perception and tracking information output in step 1 as one of the inputs, and implements collaborative path optimization and task allocation of the drone swarm through the Multi-Agent Reinforcement Learning (MARL) model.

[0064] Step 2.1, build a multi-agent reinforcement learning environment;

[0065] In this step, a multi-agent reinforcement learning environment is constructed. Each drone is first defined as an agent. All agents share environmental information and make collaborative decisions.

[0066] Specifically, define the state space of the agent , action space and the reward function .

[0067] State Space It includes: the status of each drone itself (for example, three-dimensional spatial position, flight speed, remaining energy, sensor orientation), local environmental information perceived by the onboard infrared sensor (for example, the number, confidence and distribution of heat source targets in the current field of view), and the summary of the status of other intelligent agents and the current communication topology information perceived through the inter-drone communication network (for example, the connection relationship between drones and the communication quality of each connection).

[0068] In some embodiments, the state space It may also include historical inspection information (such as maps of covered areas, historical target sighting locations) or more detailed environmental parameters (such as terrain slope, vegetation cover type, etc., if available).

[0069] Action Space Includes: flight control instructions for each drone in the next time step (for example, three-dimensional velocity vector or heading angle and climb rate adjustments), and detection mode commands (such as sensor scanning range, focusing on a specific area).

[0070] The specific dimensions and value range of are determined by the UAV platform and mission requirements.

[0071] Reward Function It is a globally shared scalar reward signal, which takes into account multiple goals, such as: rewards for the number of newly discovered endangered animal targets; rewards for the area of ​​newly explored areas; penalties for the total energy consumption of the drone swarm; rewards for the quality of maintaining effective communication links; penalties for repeated inspections of areas; and rewards for actively identifying and inspecting areas where potential animal activity is concentrated.

[0072] Optionally, the weight factors for each component of the reward function Different configuration schemes can be dynamically adjusted or preset according to different monitoring stages (such as wide-area search stage, key area monitoring stage) or specific mission objectives (such as prioritizing coverage or prioritizing continuous tracking of known targets).

[0073] The formula is:

[0074] ;

[0075] in Indicates time The total reward, is the total number of reward items, For the The weight factor of the reward, For the Item at time The reward value, Express from arrive The sum of .

[0076] Step 2.2, train the MARL model;

[0077] Then, the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is used to train the MARL model.

[0078] Each agent has an actor network that outputs deterministic actions and a critic network that evaluates the joint state-action value. Typically, both the actor network and the critic network are composed of multi-layer perceptrons (MLPs), which contain several hidden layers and nonlinear activation functions (such as ReLU).

[0079] In some embodiments, for scenarios with complex state spaces or that require processing serialized inputs (e.g., decisions based on historical trajectories), the Actor or Critic network may also employ a recurrent neural network (RNN) or a structure that includes an attention mechanism.

[0080] Specifically, in the drone collaborative inspection scenario, the Actor network receives the The observed state of an agent As input, output The specific flight action of an intelligent agent (such as speed and direction adjustment instructions);

[0081] The Critic network receives all The joint state of the agents and all Joint actions of agents As input, the output is Evaluation value of the joint state action of the agents , this evaluation value is used to guide the learning of the Actor network.

[0082] in 、 Respectively represent the first and The actions of an agent, is the total number of agents participating in collaborative inspection, Indicates all The joint state of the agents.

[0083] The actor network is updated using a policy gradient method, aiming to maximize the expected cumulative reward. The critic network is updated by minimizing the temporal difference error (TD-error) to accurately evaluate the state-action-value function. During training, the agent generates experience through interaction with the environment. This experience is stored in a replay buffer and used to update the network.

[0084] Input data types: state information sequence of each drone, joint action sequence, and global reward signal sequence, which are collected from the interaction between the agent and the environment;

[0085] Specific output results: The trained MARL model is specifically manifested as a set of Actor network parameters and Critic network parameters that can guide the drone swarm to perform efficient collaborative inspections and path optimization.

[0086] Step 3: Based on the collaborative inspection path optimization strategy, the execution of the collaborative inspection path optimization strategy is adaptively adjusted according to the real-time communication quality between each drone in a group and the real-time energy level of each drone;

[0087] The communication quality and energy state information required in this step are important components of the state space of the MARL model, which directly affect the collaborative decision-making and path optimization of the UAV swarm.

[0088] Step 3.1, build the UAV communication quality perception module;

[0089] Each drone monitors communication signal strength, data transmission latency, and packet loss rate between itself and its neighbors in real time to quantify the quality of the current communication connection. This communication quality information becomes part of the state space of the agent in the MARL model. A commonly used signal strength metric is RSSI (Received Signal Strength Indicator), latency can be expressed using RTT (Round-Trip Time), and packet loss rate is the proportion of packets lost per unit time.

[0090] Input data type: wireless signal parameters between drones;

[0091] Specific output results: quantitative values ​​or status levels representing the quality of the communication links between UAVs.

[0092] Step 3.2, build the UAV energy status perception module;

[0093] Each drone monitors the remaining battery power percentage in real time and uses it as a key component of the agent's state space. The energy percentage is calculated as:

[0094] ;

[0095] in, Indicates the current remaining energy percentage. Indicates the current remaining power. Indicates the maximum battery capacity.

[0096] Input data type: voltage and current data provided by the drone battery management system;

[0097] Specific output results: the current remaining energy percentage of each drone .

[0098] Step 3.3, integrating communication perception information and energy perception information;

[0099] When generating actions, the agent's actor network comprehensively considers current communication quality and its own energy level. For example, when communication quality degrades or energy levels are low, the strategy will favor less energy-intensive flight maneuvers, avoid subtasks requiring high-bandwidth communication, or proactively seek communication relay points or return for recharge. When allocating tasks or assigning roles (such as focused area searches or target relay tracking), MARL's global optimization objective prioritizes subtasks requiring high energy consumption or high communication continuity to drones with sufficient energy and good communication connections.

[0100] Input data types: MARL model state input (including communication quality and energy state), trained Actor-Critic network;

[0101] Specific output results: UAV flight action instructions and task allocation tendencies dynamically adjusted according to real-time communication and energy status.

[0102] Step 4: When some UAVs in a group fail due to communication interruption or energy exhaustion, the tasks of the remaining UAVs are adjusted based on the real-time state perception results, the preset fault-tolerant control logic and dynamic task reallocation model;

[0103] This step relies on the aforementioned collaborative decision-making and state perception results to ensure the robustness and continuity of the overall monitoring task in abnormal situations such as communication interruption or energy exhaustion of the drone.

[0104] Step 4.1, build the UAV individual fault-tolerant flight module;

[0105] When a UAV is unable to receive collaborative instructions from the MARL model due to severe communication interruption for more than a preset time threshold ( Indicates the maximum allowed communication interruption duration threshold. When the threshold is exceeded, the emergency mode is triggered. The drone automatically activates the preset emergency flight mode.

[0106] In this mode, the drone makes autonomous decisions based on its last valid command, its local environmental perception (such as infrared sensor data and GPS positioning), and its built-in obstacle avoidance logic to either complete the currently assigned sub-mission objective or execute a safe return procedure. For example, optional emergency actions may include: continuing to fly in a straight line in the last commanded direction for a preset distance or time, hovering in place and attempting to reestablish communication, or directly initiating a return to a preset safe point.

[0107] Input data types: UAV current status, local sensor data, preset emergency rule library;

[0108] Specific output results: UAV autonomous flight action instructions in the event of communication interruption.

[0109] Step 4.2: Build a dynamic task reallocation model based on a fast decision network;

[0110] When a drone responsible for a critical subtask is detected to have failed due to energy exhaustion, unexpected malfunction, or loss of communication, the system triggers a dynamic task reallocation process, which is executed by a pre-trained fast decision-making network.

[0111] As an example, the fast decision network can be a small multilayer perceptron (MLP). Its input layer receives encoded information, including the failed UAV's mission information (such as mission type, importance, and original target), the current status of all healthy UAVs (such as location, remaining energy, current mission payload, and special sensor capabilities), and environmental information (such as the cost of traveling to the target area). The network's hidden layers process these input features through nonlinear transformations to extract key information for decision-making. The output layer then outputs one or a set of instructions to determine which healthy UAV (or UAVs) is most suitable to take over the failed mission, as well as possible mission adjustment parameters (such as new waypoints and mission priority). The MLP is pre-trained using supervised learning. The training data can be derived from a large number of task reassignment examples in a simulation environment and their corresponding expert solutions or optimal solutions. Optionally, in some embodiments, the training data can also include successful and failed task reassignment examples collected from actual operational history. Imitation learning or offline reinforcement learning methods can be used to further optimize the performance of the decision network.

[0112] The optimization goal of task reallocation is multidimensional. For example, it can include minimizing the overall task completion time increment caused by reallocation. , maximize the probability of successful task succession and minimize the interference with other existing subtasks;

[0113] The calculation formula is:

[0114] ;

[0115] in, It represents the estimated overall task completion time after task reallocation, Indicates the estimated completion time before reallocation, represents the increment of the overall task completion time due to task reallocation, The smaller it is, the less interference the reallocation strategy has on the task progress, and the stronger the system's response and recovery capabilities are.

[0116] Therefore, during the dynamic task reallocation process, one of the optimization goals is to minimize , to ensure efficient and continuous execution of monitoring tasks.

[0117] Input data types: failed drone information, status information of remaining healthy drones, current task list and priority, and pre-trained fast decision network parameters;

[0118] Specific output results: Reassignment instructions for failed drone missions, clearly specifying the replacement drone and adjusted mission parameters.

[0119] A device for monitoring the whereabouts of endangered animals includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned method for monitoring the whereabouts of endangered animals.

[0120] Here, the present invention provides an implementation example:

[0121] This implementation was applied to a national nature reserve to monitor the movements of a rare nocturnal musk deer and conduct a preliminary population assessment. The reserve covers approximately 200 square kilometers and features complex terrain, including extensive virgin forests and steep mountainous terrain. Traditional manual patrols are difficult to fully cover, especially at night. Musk deer are reclusive and have low population densities, making effective monitoring crucial for developing conservation strategies.

[0122] Examples of UAV swarm collaborative perception and target recognition:

[0123] A swarm of five small rotary-wing drones, each equipped with a high-resolution infrared thermal imaging camera, was deployed. During the nighttime hours (e.g., 8:00 PM to 4:00 AM), the swarm began patrolling within a pre-defined gridded inspection area (each grid cell was 2 km x 2 km). While flying over Area A (a patch of mixed coniferous and broad-leaved forest where musk deer have historically frequented forest musk deer), the thermal imaging camera of one of the drones (UAV-1) detected an unusual heat source. The onboard image processing unit immediately processed the infrared image using the adaptive Otsu threshold segmentation algorithm, successfully segmenting the heat source from the background and forming a clear target outline. Subsequently, a Kalman filter-based target tracking algorithm was activated. Based on the heat source's initial position, size (consistent with the size of an adult musk deer), and slight movement, it was identified as a highly suspected live musk deer. UAV-1 shared this target's preliminary identification information (temporary ID, GPS coordinates, and confidence score) with neighboring UAVs UAV-2 and UAV-3 via the inter-UAV ad hoc network. Over the next 15 minutes, UAV-1 continued to track the target steadily, recording its movement path and speed. During this period, the target entered a dense bush and part of its body was obscured, but the tracking algorithm was still able to maintain lock thanks to motion prediction and thermal signal characteristics.

[0124] Examples of MARL-based collaborative inspection, path optimization, and adaptive control:

[0125] After UAV-1 detected the target, the MADDPG model optimized the fleet's strategy based on the current status of all drones (UAV-1 tracking the target, UAV-2 and UAV-3 nearby, and UAV-4 and UAV-5 patrolling further away), the covered area, and a map of musk deer activity probability generated based on historical data. Specifically, UAV-2 was instructed to adjust its course and approach the target from the side, assisting UAV-1 in forming a multi-angle observation to prevent the target from being lost. UAV-3 was instructed to hover at a slightly longer distance, acting as a communication relay node to ensure smooth communication between UAV-1 and UAV-2 and the rear base station while monitoring the surrounding area for other potential targets. UAV-4 and UAV-5, driven by the reward function's incentive to explore uncovered areas, continued their search in other areas. However, their search intensity and direction were dynamically fine-tuned by the MARL model in the hope of discovering new musk deer individuals or groups. During this process, UAV-1 reported that its remaining energy had dropped to 30%, below the pre-set safety threshold of 40%. The MARL model senses this state and, based on the quality of the communication link (UAV-3 acts as a relay to ensure good communication), gradually and smoothly transfers the main tracking task from UAV-1 to UAV-2, which has more energy (70% remaining) and has reached a suitable observation position. At the same time, it instructs UAV-1 to choose the path with the lowest energy consumption to return to the base for recharging after completing the handover.

[0126] Fault-tolerant control and dynamic task redistribution examples:

[0127] During another patrol mission, UAV-4 was conducting a search mission in Area B (a canyon) when, due to signal obstruction caused by the complex terrain, communication with the rest of the fleet and the base station was suddenly interrupted for a period exceeding the preset 60-second threshold. UAV-4 automatically activated its individual fault-tolerant flight module. Based on the last received command (to search westward along the canyon) and its GPS and infrared sensing (no targets detected, open terrain ahead), it continued westward for a preset 500 meters before circling in place and attempting to reestablish communication. Simultaneously, the system detected that UAV-4 had lost contact and its Area B search submission had been aborted. The rapid decision network was activated, comprehensively assessing the current status of UAV-5 (closest to Area B, with 80% remaining power and a standard sensor configuration) and UAV-3 (further away but equipped with a wider-angle thermal imaging camera, with 60% remaining power). Considering the importance of Area B (historically frequented by musk deer but recently uninhabited), and the superior overall conditions of UAV-5, the rapid decision network issued instructions, directing UAV-5 to suspend its patrol mission in Area C and adjust its route to Area B, taking over UAV-4's search mission. Upon receiving the instructions, UAV-5 successfully arrived at Area B and commenced its search. Subsequently, UAV-4 successfully reestablished communication after circling for two minutes and, depending on the new instructions from the MARL model, either joined the coordinated search of Area B or returned to base.

[0128] By deploying the method described in this embodiment in the aforementioned application scenario and comparing it with the traditional single-UAV infrared inspection method based on fixed path planning, the effectiveness and superiority of this method were verified. The two most important technical performance indicators selected were "average target detection rate at night" and "number of mission interruptions due to communication or energy problems."

[0129] This method, through its MARL-optimized collaborative search strategy and wider coverage, enabled the drone swarm to detect suspected musk deer targets an average of 3.2 times per night during a one-month (30-night) test period, with an average of 1.8 of these targets subsequently confirmed as real musk deer. In comparison, a traditional single-drone fixed-path patrol method detected suspected targets an average of 1.1 times per night during the same test period, with an average of 0.5 confirmed targets. This demonstrates that this method can improve the probability of detecting cryptic endangered animals at night.

[0130] During a month-long test of this method, weak communication signals caused individual drones to briefly (less than 5 minutes) switch to emergency mode eight times. Adaptive control or dynamic task reallocation ensured mission continuity in all instances, preventing a complete disruption to the entire surveillance chain. Five times, a single drone returned home early due to power exhaustion, and its mission was successfully taken over. In contrast, a control group, employing simple formation flight (simply dispersed patrols along a pre-set path) with multiple drones but without advanced coordination and fault-tolerant control, experienced four instances within a month in which core communication node failures or unexpected power exhaustion by patrol aircraft in key areas resulted in large areas of unobstructed surveillance, thus failing to achieve mission objectives.

[0131] The data of technical effects are shown in Table 1:

[0132] Table 1: Comparison of technical effects

[0133]

[0134] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A method for monitoring the whereabouts of endangered animals, characterized in that: include: The infrared thermal imaging device carried by each drone in a group of drones deployed in the monitoring area collects infrared thermal imaging data, identifies potential animal targets based on the infrared thermal imaging data, tracks the potential animal targets, and obtains target tracking information; Based on target tracking information, a multi-agent reinforcement learning model is used. Each agent in the MADDPG model corresponds to one drone in a group of drones. Based on the real-time state perception information of each drone in the group and the communication topology information between drones, a collaborative inspection path optimization strategy for the group of drones is generated. The multi-agent reinforcement learning model is a multi-agent deep deterministic policy gradient model. The real-time status perception information includes the position, speed, and remaining energy of each drone; The communication topology information between the drones is modeled through a graph neural network to determine the connection relationship and communication quality between the drones; Based on the collaborative inspection path optimization strategy, the execution of the collaborative inspection path optimization strategy is adaptively adjusted according to the real-time communication quality between each drone in a group and the real-time energy level of each drone; When some drones in a group fail due to communication interruption or energy exhaustion, the tasks of the remaining drones are adjusted based on the real-time status perception results, according to the preset fault-tolerant control logic and dynamic task reallocation model.

2. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: The infrared thermal imaging data identifies potential animal targets, including: An adaptive threshold segmentation algorithm is used to process infrared thermal imaging data and preliminarily extract potential animal heat source areas.

3. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: Tracking potential animal targets and obtaining target tracking information includes: A Kalman filter algorithm or a particle filter algorithm is used to estimate and predict the motion state of the initially extracted potential animal heat source area to achieve stable tracking, and to record the real-time position coordinate sequence and motion parameters of the animal target.

4. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: The generation of a collaborative inspection path optimization strategy for a group of drones includes: Each agent is configured with an Actor network and a Critic network; The Actor network outputs the UAV's flight actions based on the real-time state information, local environment information, and communication topology information of the intelligent agent. The local environment information at least includes the heat source target information in the current field of view. The critic network evaluates the value of the joint state and joint actions of all agents; The Actor network and Critic network are trained by maximizing a globally shared cumulative reward function, which comprehensively considers new target discovery, area coverage, energy consumption, and communication connection quality.

5. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: Adaptively adjusting the execution of the collaborative inspection path optimization strategy based on the real-time communication quality between each drone in a group of drones and the real-time energy level of each drone includes: When it is monitored that the communication quality between drones drops below a preset threshold or the real-time energy level of a drone is lower than a preset energy threshold, the flight action of the drone or its role in the collaborative mission is adjusted to prioritize ensuring the stability of the communication link and avoiding excessive energy consumption.

6. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: When some UAVs in a group fail due to communication interruption or energy exhaustion, the tasks of the remaining UAVs are adjusted based on the real-time state perception results, the preset fault-tolerant control logic and the dynamic task reallocation model, including: When a drone experiencing communication interruption fails to restore communication within a preset time threshold, the drone automatically switches to an emergency flight mode based on local perception; When a drone responsible for a critical subtask is detected to have failed, the dynamic task reallocation model is used to reallocate the failed drone's tasks to one or more remaining healthy drones based on the current global task priority and the status of the remaining healthy drones.

7. The method for monitoring the whereabouts of endangered animals according to claim 6, characterized in that: The dynamic task reallocation model is a pre-trained fast decision network that takes the task information of the failed UAV, the status information of all healthy UAVs, and the environmental information as input and outputs a task succession plan.

8. The method for monitoring the whereabouts of endangered animals according to claim 1, characterized in that: Infrared thermal imaging data is collected at night or in low-light conditions.

9. An endangered animal whereabouts monitoring device, characterized in that: The device comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the device is used to implement a method for monitoring the whereabouts of endangered animals according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Aerial multi-agent distributed phase regulation and target tracking method under guidance of multiple closed paths

    CN114115347A

  • Unmanned aerial vehicle agricultural bird repelling method and system based on topological sorting reward mechanism

    CN117441701A