Wildlife dynamic monitoring system based on unmanned aerial vehicle vision and remote sensing data fusion

By integrating multiple sensors and edge computing through a drone vision and remote sensing data fusion system, the problem of low monitoring accuracy from a single data source is solved, enabling all-weather, high-precision dynamic monitoring of wildlife and providing multi-dimensional ecological information.

CN122431203APending Publication Date: 2026-07-21JIANGXI ACAD OF FORESTRY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGXI ACAD OF FORESTRY
Filing Date
2026-04-21
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies for wildlife monitoring that rely on a single data source suffer from low monitoring accuracy and limited information dimensions in complex habitats, making it impossible to achieve all-weather, multi-dimensional dynamic monitoring.

Method used

The system employs a UAV vision and remote sensing data fusion system, integrating a high-resolution visible light camera, a long-wave infrared thermal imager, a multispectral imager, and a highly directional microphone array. By combining edge computing and data preprocessing, it achieves robust detection, accurate classification, 3D localization, and behavior recognition through multi-source heterogeneous data fusion analysis.

Benefits of technology

It significantly improved the detection rate and identification accuracy of wild animals under vegetation cover or low light conditions, and realized all-weather, high-precision, adaptive multi-dimensional dynamic monitoring, thus improving the technical level of ecological monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122431203A_ABST
    Figure CN122431203A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of environmental monitoring and wildlife protection, and specifically discloses a dynamic monitoring system for wild animals based on unmanned aerial vehicle vision and remote sensing data fusion, which comprises a multi-modal data acquisition subsystem, an edge computing and preprocessing subsystem, a multi-source heterogeneous data fusion analysis subsystem, and a dynamic task planning and feedback control subsystem. By fusing visible light, thermal infrared, acoustic and vibration multi-source data, and using evidence theory, graph optimization and other technologies for target detection, tracking and behavior analysis, robust, all-weather, multi-dimensional dynamic monitoring of wild animals is achieved, and the monitoring strategy can be adaptively adjusted according to real-time analysis results, significantly improving the monitoring efficiency and automation level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental monitoring and wildlife protection technology, specifically relating to a dynamic wildlife monitoring system that integrates UAV visual and remote sensing data. Background Technology

[0002] Wildlife dynamic monitoring is a key technology in ecology, conservation biology, and natural resource management. It aims to obtain core information such as animal population size, distribution, behavioral patterns, and habitat use, providing data support for species conservation, ecological assessment, and scientific decision-making. Unmanned aerial vehicle (UAV)-based automated monitoring technology has become an important development direction in this field due to its advantages such as mobility, wide coverage, and the ability to acquire high-resolution data.

[0003] Current technologies primarily rely on a single data source, such as using drones equipped with visible light or thermal imaging cameras for visual monitoring, or ground-based acoustic sensors for audio monitoring. However, in complex, obstructed environments like dense forests and thickets, animal targets are easily covered by vegetation, leading to a significant decrease in recognition rates or even complete missed detections. Furthermore, visual methods cannot effectively capture crucial behavioral information such as animal vocalizations and activity sounds, limiting in-depth understanding of their ecological behaviors, including communication, alertness, and foraging. Acoustic monitoring alone struggles to accurately locate sound sources and is susceptible to environmental noise interference, failing to provide precise spatial location and visual morphological characteristics of animals. Summary of the Invention

[0004] The purpose of this invention is to provide a dynamic wildlife monitoring system that integrates UAV visual and remote sensing data, in order to solve the technical contradictions in the existing technology that rely on a single data source, resulting in low monitoring accuracy in complex habitats, limited information dimensions, and the inability to achieve all-weather, multi-dimensional dynamic monitoring.

[0005] This invention provides a wildlife dynamic monitoring system that integrates UAV visual and remote sensing data. The system includes a multimodal data acquisition subsystem, an edge computing and data preprocessing subsystem, a multi-source heterogeneous data fusion and analysis subsystem, and a dynamic task planning and feedback control subsystem.

[0006] The multimodal data acquisition subsystem is used to synchronously acquire spatiotemporally aligned multi-source sensing data of the target monitoring area. This subsystem consists of an unmanned aerial vehicle (UAV) platform and ground-based auxiliary sensing nodes. The UAV platform integrates a high-resolution visible light camera, a long-wave infrared thermal imager, a multispectral imager, and a highly directional microphone array. The high-resolution visible light camera captures the morphological, texture, and color features of animal targets; the long-wave infrared thermal imager detects homeothermic animals partially obscured by vegetation or under low-light conditions based on differences in biological thermal radiation; the multispectral imager acquires spectral information reflecting vegetation type, coverage, and physiological state to assist in habitat classification and background analysis; the highly directional microphone array, composed of at least four acoustic sensors arranged in a specific geometric structure, is used to collect animal vocalization signals and preliminarily estimate the horizontal azimuth angle of the sound source using a beamforming algorithm. Ground-based auxiliary sensing nodes are pre-deployed at key locations in the monitoring area. Each node includes an omnidirectional microphone and a ground vibration sensor to collect near-ground acoustic signals and ground vibration signals caused by animal activity, supplementing and verifying the data from the aerial platform.

[0007] The edge computing and data preprocessing subsystem, deployed on the UAV's onboard computing unit and ground base station server, is used for real-time processing and initial feature extraction of raw sensing data to reduce data transmission load and provide standardized input for subsequent fusion analysis. The specific processing flow executed by this subsystem includes: adaptive histogram equalization and dehazing enhancement processing of visible light images; time-series-based background subtraction and adaptive threshold segmentation of thermal imaging video streams to extract potential thermal target areas; atmospheric correction and radiometric calibration of multispectral images, and calculation of at least five vegetation indices, including the normalized vegetation index and enhanced vegetation index; pre-emphasis, frame-by-frame windowing processing of audio signals acquired by the microphone array, and extraction of acoustic feature vectors such as Mel frequency cepstral coefficients, spectral centroid, and zero-crossing rate; bandpass filtering of ground vibration sensor signals to remove low-frequency environmental noise, and calculation of the signal's energy envelope and short-time over-threshold rate.

[0008] The multi-source heterogeneous data fusion and analysis subsystem, as the core of this invention, is used to deeply fuse features and preliminary detection results from visual, thermal infrared, acoustic, and vibration modes at the decision-level, enabling robust detection, accurate classification, 3D localization, and behavior recognition of wild animal targets. This subsystem includes a spatiotemporal registration module, an evidence theory fusion module, a graph-optimized multi-target tracking module, and a behavior semantic parsing module.

[0009] The spatiotemporal registration module is used to establish the correspondence of all sensing data in a unified spatiotemporal coordinate system. This module first assigns precise timestamps and platform spatial pose labels to each frame of image, each audio segment, and each vibration data point based on global navigation satellite system timing and inertial measurement unit data. Further, the module utilizes simultaneous localization and mapping (SLAM) technology to construct a 3D point cloud map of the monitoring area and projects all sensor data into this unified map coordinate system. For acoustic and vibration data, the module solves a hyperbolic equation system based on the time difference of arrival (TDOA) and combines this with a propagation path terrain attenuation model provided by the 3D map to map the preliminary location results of the sound and vibration sources to 3D spatial coordinates.

[0010] The evidence theory fusion module addresses uncertainties and conflicts in target detection results from different sensors. This module defines a basic probability assignment function for each independent sensor detector. For visible light detectors, the basic probability assignment is based on the confidence score of the target bounding box and the matching degree of morphological features; for thermal infrared detectors, it is based on the temperature consistency, dimensional stability, and shape regularity of the hotspot region; for acoustic detectors, it is based on the species matching probability of the acoustic signature features and the covariance of the sound source location estimation. The module employs Dempster's combination rule to iteratively combine basic probability assignments from at least two heterogeneous sensors for the same spatial region, calculating joint support, joint uncertainty, and conflict metric. Finally, hypotheses where the joint support exceeds a preset first threshold and the conflict metric is below a preset second threshold are considered reliable fused detection results, and the module outputs the probability of the target's existence, the probability distribution of its species, and the fused spatial location estimate.

[0011] The graph-optimized multi-target tracking module is used to correlate and fuse detection results over continuous time to form target trajectories and address the issues of target occlusion, disappearance, and reappearance. This module treats the fused detection results at each moment as nodes in a graph, and the correlation possibilities between nodes at different moments as edges, constructing a spatiotemporal graph model. The edge weights are calculated using a multi-factor cost function that integrates the target's appearance feature similarity cost, motion consistency cost, and spatial proximity cost. The appearance feature similarity cost is calculated based on the cosine distance of the depth feature vector extracted from the fused detection results; the motion consistency cost is calculated based on the Mahalanobis distance between the predicted position and the measured position using a uniform or uniformly accelerated motion model; and the spatial proximity cost considers the reachability path constraints of the target in the 3D map. This module generates and maintains multi-target trajectories by solving the maximum a posteriori probability estimation problem, i.e., finding the node matching scheme that minimizes the global correlation cost. When a target is lost due to temporary occlusion, the module activates a trajectory prediction mode, extrapolates motion based on historical trajectories, and attempts to correlate with reappearing detections within a preset time window.

[0012] The behavioral semantic analysis module is used to identify specific behavioral patterns of animals based on trajectory data, acoustic event sequences, and habitat context information. This module constructs a hierarchical behavioral recognition model. The first layer is basic action recognition, which identifies basic movement patterns such as standing still, walking, running, and jumping by analyzing the velocity, acceleration, and turning angle sequences of the target trajectory; and identifies specific types of calls, such as alarm calls, courtship calls, and communication calls, by analyzing the time-frequency patterns of acoustic features. The second layer is complex behavioral inference, which, based on the basic actions and acoustic events identified in the first layer, combines the target's habitat type obtained from multispectral data analysis, temporal information, and spatial relationships between multiple targets, and infers higher-level behavioral semantics such as foraging, drinking, socializing, parenting, and alarming through predefined logical rules and statistical models.

[0013] The dynamic task planning and feedback control subsystem is used to dynamically adjust the operating parameters and flight path of the multimodal data acquisition subsystem based on the real-time output of the multi-source heterogeneous data fusion analysis subsystem, forming a closed loop of perception-analysis-decision. This subsystem includes an abnormal behavior early warning device, an interest region identifier, and an adaptive trajectory planner. The abnormal behavior early warning device monitors the output of the behavior semantic parsing module in real time. When it detects preset abnormal behavior patterns such as swarming alarms or high-frequency warning cries, it immediately generates a high-level alarm. The interest region identifier dynamically marks areas of interest requiring focused or repeated monitoring on a 3D map based on the target spatial distribution density map output by the multi-target tracking module, the location of newly discovered targets, and areas of high uncertainty in the evidence theory fusion module. The adaptive trajectory planner integrates the current UAV position, remaining flight time, interest region distribution, and weather conditions, with the optimization objective of maximizing the amount of information acquired from the interest region per unit time. It re-plans the UAV's flight path, hovering position, and sensor operating mode in real time, for example, instructing the UAV to fly to the interest region and switch the sensors to high-resolution mode for detailed investigation.

[0014] As one embodiment of the present invention, the construction process of the basic probability allocation function in the evidence theory fusion module is as follows: First, the original detection output of each sensor is normalized to transform it into a support measure for propositions such as the existence of the target, the target belonging to a certain species, and the target being located in a certain region. The support value is between 0 and 1. Second, a reliability weight is assigned to each sensor. This weight is calculated by combining the precision and recall of the sensor on the known truth dataset during the offline calibration phase. Finally, the normalized support is multiplied by the reliability weight, and a portion of the probability quality is reserved for uncertain propositions, thereby completing the construction of the basic probability allocation function.

[0015] In one embodiment of the present invention, the multi-target tracking module based on graph optimization employs a cascaded matching strategy combining the Hungarian algorithm and the minimum cost flow algorithm to optimize the spatiotemporal graph model. First, between two consecutive frames, the Hungarian algorithm is used to perform preliminary association based on appearance and motion cost, resolving most simple associations. Second, for detection nodes and trajectories that fail to be associated, they are incorporated into a sliding time window containing multiple frames to construct a more complex graph model, and the minimum cost flow algorithm is used for global optimization to handle complex trajectory start, end, and intersection cases.

[0016] In one embodiment of the present invention, the composite behavior inference layer of the behavior semantic parsing module integrates a time-series neural network model based on an attention mechanism. The model's input consists of the action encoding sequence, acoustic event encoding sequence, and habitat context feature vector output by the basic action recognition layer within a past time window. The model automatically learns the importance weights of the input features at different time steps for the current behavior classification through the attention layer, ultimately outputting the probability distribution for multiple preset composite behavior categories. This neural network model is obtained through end-to-end training on a large amount of labeled behavior sequence data.

[0017] In one embodiment of the present invention, the optimization objective function of the adaptive trajectory planner is specifically to maximize the sum of coverage quality for all regions of interest under the constraints of UAV endurance and flight safety. Coverage quality is defined as the product of a negative exponential function of the time required to reach the region, the information value weight of the region, and the information gain that can be obtained by using a specific sensor mode in the region. The information value weight is dynamically assigned by the region of interest identifier based on the target density, target sparseness, and uncertainty within the region. This optimization problem is solved in the rolling time domain using a model predictive control framework, outputting the optimal control sequence for a future time domain period in each planning cycle.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention fundamentally solves the problem of insufficient monitoring capabilities of a single data source in complex environments by constructing a complete system-level technical solution that includes multimodal data acquisition, edge preprocessing, decision-level multi-source heterogeneous fusion analysis, and dynamic task planning feedback.

[0019] 2. Through the evidence theory fusion module, the uncertainty and conflict in the detection results of visual, thermal infrared and acoustic sensors are handled in a rigorous mathematical form. The complementarity of multi-source information is transformed into a significant improvement in decision reliability, which fundamentally improves the detection rate and recognition accuracy of wild animals under vegetation cover or low light conditions.

[0020] 3. The graph-optimized multi-target tracking module enables continuous and stable tracking of individual animal trajectories in three-dimensional space, effectively overcoming the problem of target loss due to occlusion and providing a continuous data foundation for behavior analysis.

[0021] 4. The behavioral semantic analysis module integrates trajectory, sound, and habitat context, achieving a leap from low-level perception data to high-level behavioral semantics, providing ecological information dimensions far exceeding those of single visual or acoustic monitoring. Finally, the dynamic task planning and feedback control subsystem forms a closed-loop intelligence, enabling the system to autonomously optimize monitoring strategies based on real-time analysis results, accurately deploying limited UAV resources to areas with the highest information value, maximizing monitoring efficiency and data quality.

[0022] 5. This invention provides an integrated dynamic monitoring solution for wild animals that is all-weather, high-precision, adaptive, and multi-dimensional, significantly improving the technical level and application value of automated ecological monitoring. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the multi-source heterogeneous data fusion and analysis subsystem in this invention; Figure 3 This is a flowchart illustrating the logical flow of the evidence theory fusion module in this invention for handling multi-sensor uncertainties. Figure 4 This is a schematic diagram of the spatiotemporal correlation framework of the graph-optimized multi-target tracking module in this invention; Figure 5 This is a schematic diagram illustrating the interactive relationship between the dynamic task planning and feedback control subsystem in this invention, forming a perception-analysis-decision closed loop. Detailed Implementation

[0024] This invention provides a wildlife dynamic monitoring system that integrates UAV visual and remote sensing data, the overall technical architecture of which is shown in the attached figure. Figure 1 As shown in the figure, the system consists of four core components: a multimodal data acquisition subsystem, an edge computing and data preprocessing subsystem, a multi-source heterogeneous data fusion and analysis subsystem, and a dynamic task planning and feedback control subsystem. These subsystems are tightly coupled through a high-bandwidth, low-latency data bus and a unified time synchronization protocol, forming a complete closed loop from sensing, processing, analysis to decision feedback. The following will be combined with the attached... Figure 1 To be continued Figure 5 The specific embodiments of the present invention will be described in detail below.

[0025] First, the multimodal data acquisition subsystem is responsible for simultaneously acquiring spatiotemporally aligned multi-source sensing data within the target monitoring area. This subsystem consists of an airborne platform and ground-based auxiliary sensing nodes. The airborne platform is a multi-rotor or fixed-wing UAV with long endurance capabilities, and its onboard sensor array includes a high-resolution visible light camera, a long-wave infrared thermal imager, a multispectral imager, and a highly directional microphone array. The high-resolution visible light camera employs a global shutter CMOS image sensor with a resolution of at least 20 megapixels and an adjustable frame rate of 1 to 30 frames per second, used to capture the morphological contours, fur textures, and color features of animal targets. The long-wave infrared thermal imager operates in the 8 to 14 micrometer band, with a thermal sensitivity better than 0.05 degrees Celsius and a spatial resolution of 640×512 pixels, used to detect the thermal radiation signals generated by homeothermic animals due to their body temperature being higher than the environmental background, and is particularly suitable for target detection in vegetation-covered or low-light nighttime scenarios. The multispectral imager contains at least 5 independent spectral channels, covering the visible to near-infrared band (450 nm to 900 nm), used to collect spectral reflectance data reflecting ecological parameters such as vegetation type, chlorophyll content, and water stress status. The highly directional microphone array consists of 4 omnidirectional capacitive acoustic sensors arranged in a square geometric layout with a side length of 0.5 meters and a sampling frequency of 48 kHz, used to collect the vocal signals of wild animals and to preliminarily estimate the horizontal azimuth angle of the sound source through beamforming algorithms.

[0026] Ground-based auxiliary sensing nodes are pre-deployed at key ecological sites within the monitoring area, such as water sources, animal trail intersections, or around dens. Each node includes an omnidirectional microphone and a triaxial ground vibration sensor. The omnidirectional microphone is used to collect low-frequency animal calls or group activity noise propagating near the ground; the ground vibration sensor uses a piezoelectric ceramic sensing element with a range of ±2g and a sampling frequency of 1 kHz to detect micro-vibration signals on the ground caused by the walking, running, and other activities of large mammals. All ground nodes are equipped with a real-time clock module and a low-power wide-area communication unit, enabling them to accurately timestamp the collected acoustic and vibration data and transmit it back to the ground base station server via LoRa or NB-IoT networks as a supplement and cross-validation of the data from the aerial platform.

[0027] All sensor data acquisition strictly adheres to a unified time reference. The UAV platform is equipped with a high-precision Global Navigation Satellite System receiver and an inertial measurement unit, which are fused through tightly coupled Kalman filtering to provide position, velocity, and attitude information at a frequency of 100 Hz. This spatiotemporal reference signal is broadcast to all airborne sensors and ground nodes, ensuring that every frame of image, every audio clip, and every vibration waveform is timestamped with nanosecond-level precision and associated with the corresponding platform spatial pose. Thus, the raw data achieves strict spatiotemporal alignment at the source of acquisition, laying the foundation for subsequent fusion analysis.

[0028] The edge computing and data preprocessing subsystem is deployed in two physical locations: one is an onboard embedded computing unit on the UAV, typically a heterogeneous computing platform based on NVIDIA Jetson AGX Orin or equivalent performance; the other is a ground base station server located near the monitoring area, equipped with a multi-core CPU and GPU accelerator card. The primary task of this subsystem is to perform real-time compression, enhancement, and initial feature extraction on massive amounts of raw sensing data to significantly reduce the transmission load of the wireless link and provide structured and standardized input for downstream fusion modules.

[0029] For high-resolution visible light images, this subsystem performs adaptive histogram equalization and a priori dehazing algorithm for the dark channel. Adaptive histogram equalization dynamically adjusts the grayscale mapping function based on local image contrast to avoid noise amplification caused by global equalization. The dehazing algorithm uses an atmospheric scattering model to estimate the transmittance map and recover the haze-free image, as shown in the following formula: ; in, To observe the image, This represents the global atmospheric light value. For the estimated transmittance map, To prevent division by zero, a lower threshold is required. This is the restored, clear image. This processing significantly improves the usability of images under adverse weather conditions such as fog and sandstorms.

[0030] For long-wave infrared thermal imaging video streams, this subsystem employs a time-series-based background subtraction method. First, a dynamic background model is constructed using a sliding window midpoint filter. Then, the current frame is subtracted pixel-by-pixel from the background model to obtain the difference image. Finally, adaptive thresholding is applied, with the threshold dynamically adjusted based on the standard deviation of the local region to extract potential thermal target areas. The segmentation results are output as a binary mask, marking connected regions that may contain animals.

[0031] The multispectral image processing workflow includes radiometric calibration, atmospheric correction, and vegetation index calculation. Radiometric calibration converts the raw digital values ​​into apparent reflectance; atmospheric correction employs a modified dark target subtraction method to eliminate the influence of aerosols and water vapor; based on this, the system calculates at least five vegetation indices, including the Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Soil-Adjusted Vegetation Index (SAVI), Greenness Vegetation Index (GVI), and Water Stress Index (MSI). These indices are stored in raster image format with spatial resolution consistent with the original multispectral image, and are used for subsequent habitat classification and background modeling.

[0032] In terms of audio signal processing, the raw waveforms acquired by the highly directional microphone array and the ground omnidirectional microphone are first passed through a pre-emphasis filter (coefficient of 0.97) to enhance high-frequency components; then, they are processed by framing with a frame length of 25 milliseconds, a frame shift of 10 milliseconds, and a Hamming window is applied; finally, five types of acoustic features are extracted, namely Mel frequency cepstral coefficients (the first 13 dimensions), spectral centroid, zero-crossing rate, spectral roll-off point, and short-time energy, forming a 128-dimensional feature vector, which serves as the input for acoustic event detection and species identification.

[0033] The ground vibration sensor signal is passed through a bandpass filter with a center frequency of 5 Hz and a bandwidth of 10 Hz to suppress low-frequency environmental interference such as wind noise and water flow. The short-time energy envelope (window length of 100 milliseconds) and the short-time over-threshold rate (threshold set to twice the root mean square value of the signal) of the filtered signal are calculated, and the two together constitute the criteria for vibration events.

[0034] After the above processing, all modal data are transformed into lightweight, structured intermediate representations, which are then transmitted to the ground base station via an encrypted wireless link and enter the multi-source heterogeneous data fusion and analysis subsystem. This subsystem is the core of this invention, and its internal architecture is shown in the attached figure. Figure 2 As shown, it includes four functional units: spatiotemporal registration module, evidence theory fusion module, graph optimization-based multi-target tracking module, and behavior semantic parsing module.

[0035] The spatiotemporal registration module first utilizes spatiotemporal tags provided by the Global Navigation Satellite System and the Inertial Measurement Unit to align all data to a unified world coordinate system. Further, this module runs a simultaneous localization and mapping (SMR) algorithm, fusing feature point matching results from visible light and multispectral images to construct a sparse or dense 3D point cloud map of the monitored area in real time. This map not only includes geometric structures but also incorporates vegetation index information, forming a 3D environmental model with semantic attributes. All sensor data is projected into this map coordinate system. For acoustic and vibration data, the module solves a hyperbolic positioning equation system based on time difference of arrival (TDOA) and combines this with terrain elevation and vegetation cover information provided by the 3D map to establish an attenuation and refraction model of the sound / vibration wave propagation path, thereby elevating the two-dimensional orientation estimation of the sound / vibration source to preliminary 3D spatial coordinate positioning.

[0036] The evidence theory integration module is attached. Figure 3As shown, this module is specifically designed to handle uncertainties and potential conflicts in multi-sensor detection results. It defines a basic probability assignment function for each independent detector. The basic probability assignment for visible light detectors is based on the bounding box confidence score and morphological feature matching degree (such as aspect ratio and texture complexity) output by the target detection network (e.g., YOLOv7). The basic probability assignment for thermal infrared detectors is based on the temperature standard deviation of the hotspot region (less than 2 degrees Celsius is considered stable), the area change rate (less than 10% change over 3 consecutive frames), and the shape roundness (greater than 0.6). The basic probability assignment for acoustic detectors is derived from the species posterior probability output by the voiceprint recognition model and the inverse of the determinant of the sound source location estimation covariance matrix.

[0037] This module employs the Dempster combination rule to iteratively fuse the basic probability assignments of at least two heterogeneous sensors. Let... and For two sensors on the same set of hypotheses If the basic probability distribution is given, then the basic probability distribution after combination is given. for: ; in, For conflict measurement. The system calculates joint support (i.e., (Target exists)), joint uncertainty (i.e.) With conflict measurement The hypothesis is accepted as a reliable fusion detection result only when the joint support is greater than 0.7 and the conflict metric is less than 0.3, and the result outputs the probability of the target's presence, the species probability distribution (e.g., African elephant 0.85, buffalo 0.10, others 0.05), and the fusion spatial location (obtained by weighted average or maximum likelihood estimation).

[0038] The graph-optimized multi-target tracking module is attached. Figure 4 As shown, this module is responsible for associating and fusing detection results over continuous time to generate stable trajectories. It treats the fused detection at each time step as a graph node, with node attributes including position, velocity, appearance depth features (a 2048-dimensional vector extracted by ResNet-50), and species label. The association probabilities between nodes at different time steps constitute the edges of the graph, with edge weights... Calculated using a multi-factor cost function: ; in, The cosine distance of the appearance feature. To calculate the Mahalanobis distance between the predicted and measured positions based on a uniform acceleration model, Spatial path distances are considered to account for accessibility constraints on the 3D map (if two points are blocked by mountains or water, the distance is set to infinity). Weighting coefficients. =0.5、 =0.3、 =0.2 was determined through grid search on the validation set. This represents a fusion detection node at two different points in time.

[0039] This module employs a cascaded matching strategy: firstly, it uses the Hungarian algorithm for fast one-to-one matching between adjacent frames; for unmatched detections and trajectories, they are incorporated into a sliding window of 5 frames to construct a multi-frame association graph, and a minimum cost flow algorithm is used to solve for the globally optimal match to handle complex scenarios such as target emergence, disappearance, and intersection. When a target is lost due to vegetation occlusion, the system activates a trajectory prediction mode, fitting a quadratic polynomial motion model based on the historical trajectories of 5 frames, and maintaining a virtual trajectory for the next 3 seconds, waiting for the target to be detected again before performing association recovery.

[0040] The behavioral semantic parsing module constructs a hierarchical recognition system. The first layer is basic action recognition: it analyzes the target trajectory's velocity (threshold 0.5 m / s to distinguish between stationary and moving) and acceleration (greater than 2 m / s). 2 Kinematic parameters such as running and turning angle change rate (greater than 30 degrees / second is considered a sharp turn) are used to identify four basic movements: stationary, walking, running, and jumping. At the same time, by analyzing the acoustic feature time-frequency map through convolutional neural network, three typical acoustic events are identified: alarm calls (high frequency, short, repetitive), courtship calls (low frequency, long tone, FM), and communication calls (medium frequency, stable).

[0041] The second layer infers complex behaviors by combining basic actions, acoustic events, habitat type (derived from multispectral vegetation index clustering, such as savanna, shrubland, and wetland), time (circadian rhythm), and multi-object spatial relationships (e.g., distance less than 10 meters is considered social interaction). It infers high-level behaviors through predefined rules and statistical models. For example, multiple individuals remaining still near a water source and issuing calls are inferred to be drinking; a single individual frequently making sharp turns in shrubland and issuing warning calls are inferred to be on guard; an adult walking closely with offspring and walking slowly is inferred to be parenting. Furthermore, this layer integrates a time-series neural network model based on an attention mechanism. The input consists of a sequence of action encodings (one-hot vectors), an acoustic event encoding sequence, and a habitat feature vector (5-dimensional vegetation index) from the past 30 seconds. The model learns the importance of features at each time step through a self-attention mechanism and outputs the probability distribution for eight predefined complex behaviors. This model was trained on a dataset containing 100,000 labeled behavior sequences.

[0042] The dynamic task planning and feedback control subsystem is attached. Figure 5As shown, a closed loop of perception-analysis-decision is achieved. The abnormal behavior early warning device scans the output stream of the behavior semantic analysis module in real time. When it detects preset patterns such as cluster startled flight (more than 5 individuals changing from stationary to high-speed running within 10 seconds) or high-frequency alarm calls lasting for more than 1 minute, it immediately triggers a level 1 alarm and pushes it to the administrator terminal.

[0043] The region of interest (ROI) identifier dynamically generates ROI polygons on a 3D map based on the target density heatmap (number of targets per unit area) output by the multi-target tracking module, the location of newly discovered targets (first detection of rare species), and regions marked as high uncertainty (joint support between 0.4 and 0.7) in the evidence theory fusion module. Weights are assigned as follows: target density weight is the density value normalized to 0-1, rarity weight is set according to the species protection level (e.g., 1.0 for endangered species and 0.3 for common species), and uncertainty weight is 1 minus the joint support.

[0044] The adaptive trajectory planner aims to maximize the amount of information acquired per unit time by continuously optimizing the UAV trajectory. Its objective function is: ; in, For the first Information value weight of each area of ​​interest The information gain that can be obtained by using a high-resolution model in this region (estimated by Shannon entropy reduction). The time required to fly from the current location to this area. =0.1 is the time decay coefficient. The total number of regions of interest. Constraints include: total flight time not exceeding 80% of remaining flight time, flight altitude always above terrain elevation by more than 50 meters, and avoidance of no-fly zones. This problem is solved using a model predictive control framework within each 10-second planning cycle, outputting the optimal waypoint sequence and sensor operating commands for the next 60 seconds (e.g., switching the visible light camera to 4K mode, increasing the thermal imager frame rate to 15 frames per second).

[0045] Through the coordinated operation of the above four subsystems, this invention achieves all-weather, high-precision, adaptive, and multi-dimensional dynamic monitoring of wild animals.

[0046] In another embodiment, the ground-based auxiliary sensing nodes of the multimodal data acquisition subsystem can be further expanded in functionality. Each ground node, in addition to including an omnidirectional microphone and a ground vibration sensor, can also integrate a miniature weather station for collecting local temperature, humidity, wind speed, and air pressure data. These micro-meteorological parameters are uploaded in real time via a wireless network, allowing the multi-source heterogeneous data fusion and analysis subsystem to correct the sound velocity model (sound velocity) during sound source localization. , The temperature (in Celsius) is used as an environmental context factor in the behavioral semantic parsing module to interpret animal behavior (such as reduced activity during high-temperature periods). Furthermore, ground nodes can be equipped with solar charging panels and supercapacitor energy storage units, enabling maintenance-free field deployment for up to 6 months.

[0047] In the edge computing and data preprocessing subsystem, the UAV's onboard computing unit can employ model distillation technology to compress large-scale voiceprint recognition models into a lightweight MobileNetV3 architecture. This maintains over 90% recognition accuracy while keeping inference latency below 50 milliseconds, meeting real-time requirements. Simultaneously, atmospheric correction of multispectral imagery can incorporate real-time aerosol optical thickness data, acquired by a miniature sunphotometer mounted on the UAV, further improving the accuracy of vegetation index calculations.

[0048] In the multi-source heterogeneous data fusion and analysis subsystem, the evidence theory fusion module can introduce a dynamic update mechanism for sensor reliability. The system periodically injects known ground truth samples (such as artificially placed isothermal simulated targets), evaluates the current performance of each sensor online, and dynamically adjusts its reliability weight. For example, if the detection rate of the thermal imager decreases during continuous rainy weather, its weight is automatically reduced, increasing the fusion authority of the acoustic and vibration sensors.

[0049] In the dynamic task planning and feedback control subsystem, the adaptive trajectory planner supports collaborative operations among multiple UAVs. The ground base station acts as a central coordinator, receiving status and region of interest information from multiple UAVs. It then allocates their respective task areas using distributed optimization algorithms (such as the alternating direction multiplier method), avoiding airspace conflicts and achieving parallel and efficient coverage of large monitoring areas. Each UAV only needs to communicate with the base station, eliminating the need for direct inter-UAV communication, thus simplifying the system architecture.

[0050] The above-described Example 2 was validated in a Tibetan antelope monitoring project in the alpine meadows of the Qinghai-Tibet Plateau. Due to the variable weather and complex terrain in this region, traditional single-unit systems are difficult to operate stably. After introducing micro-meteorological correction and multi-unit collaboration, the system was still able to maintain more than 85% of the effective monitoring time even at an altitude of 4,500 meters and a wind speed of 12 meters per second. It successfully captured the herding behavior and predator avoidance responses of Tibetan antelopes during their migration, providing crucial data support for ecological protection.

[0051] In another implementation, the composite behavior inference layer of the behavior semantic parsing module can deeply integrate prior ecological knowledge. The system has a built-in knowledge graph, where nodes represent concepts such as animal behavior, habitat elements, temporal rhythms, and social structures, and edges represent logical relationships (e.g., breeding season → courtship behavior ↑, water source drying up → migration behavior ↑). When basic actions and acoustic events are input, the system not only relies on the data-driven model but also uses graph neural networks to perform reasoning on the knowledge graph, correcting the biases of the pure data model. For example, if the data model classifies a behavior as foraging, but the knowledge graph indicates that it is currently the peak breeding season and the target is in a courtship area, the system automatically increases the probability of courtship behavior.

[0052] Furthermore, the dynamic task planning and feedback control subsystem can interface with the protected area management platform. When the abnormal behavior warning device triggers an alarm, the system not only notifies the rangers but also automatically links with nearby fixed equipment such as cameras and infrared trigger cameras to form a multi-angle evidence chain. Simultaneously, the area of ​​interest marker can push high-value areas to the management platform's electronic fence system to set up temporary restricted areas, preventing human activities from disturbing wildlife.

[0053] This implementation was applied to the monitoring of Asian elephants in the Xishuangbanna tropical rainforest of Yunnan Province. The system successfully identified tentative approaches to villages by elephant herds caused by farmland encroachment and issued an early warning two hours in advance, enabling management to take timely measures to drive them away and avoid human-elephant conflict. The introduction of knowledge graphs improved the ecological rationality of behavior recognition and reduced the false alarm rate.

[0054] In summary, this invention, through multi-level and multi-dimensional technological innovation, constructs a highly intelligent and adaptive dynamic monitoring system for wild animals. The successful application of each embodiment in different ecological environments fully verifies its technological universality and engineering feasibility.

Claims

1. A wildlife dynamic monitoring system that integrates UAV visual and remote sensing data, characterized in that, include: The multimodal data acquisition subsystem is used to synchronously acquire spatiotemporally aligned multi-source sensing data of the target monitoring area; The edge computing and data preprocessing subsystem is deployed on the UAV's onboard computing unit and ground base station server to perform real-time processing and initial feature extraction of raw sensing data. A multi-source heterogeneous data fusion and analysis subsystem is used to deeply fuse features and preliminary detection results from visual, thermal infrared, acoustic and vibration modes; The dynamic mission planning and feedback control subsystem is used to dynamically adjust the operation parameters and flight path of the multimodal data acquisition subsystem based on the real-time output of the multi-source heterogeneous data fusion analysis subsystem.

2. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 1, characterized in that, The multimodal data acquisition subsystem includes an unmanned aerial vehicle (UAV) platform and a ground-assisted sensing node. The UAV platform integrates a high-resolution visible light camera, a long-wave infrared thermal imager, a multispectral imager, and a highly directional microphone array. The ground-assisted sensing node includes an omnidirectional microphone and a ground vibration sensor.

3. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 2, characterized in that, The processing flow executed by the edge computing and data preprocessing subsystem includes adaptive histogram equalization and dehazing enhancement processing of visible light images, background subtraction and adaptive threshold segmentation based on time series for thermal imaging video streams, atmospheric correction and radiometric calibration of multispectral images and calculation of at least 5 vegetation indices, pre-emphasis and frame windowing processing of audio signals acquired by microphone arrays and extraction of acoustic feature vectors, bandpass filtering of ground vibration sensor signals and calculation of the signal's energy envelope and short-time overthreshold rate.

4. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 3, characterized in that, The multi-source heterogeneous data fusion and analysis subsystem includes a spatiotemporal registration module, an evidence theory fusion module, a graph optimization-based multi-target tracking module, and a behavior semantic parsing module. The spatiotemporal registration module is used to add timestamps and spatial pose labels to all sensing data based on global navigation satellite system timing and inertial measurement unit data, and to construct a three-dimensional point cloud map of the monitoring area using synchronous positioning and map building technology, projecting all sensor data into this unified map coordinate system. The evidence theory fusion module is used to define a basic probability assignment function for each independent sensor detector, and to use the Dempster combination rule to iteratively combine the basic probability assignments from at least two heterogeneous sensors for the same spatial region, calculate the joint support, joint uncertainty and conflict measure, so as to determine the reliable fusion detection results. The graph-optimized multi-target tracking module is used to treat the fusion detection results at each time step as nodes of a graph, the association probability between nodes at different times as edges of the graph, construct a spatiotemporal graph model, and generate and maintain multi-target trajectories by solving the node matching scheme that minimizes the global association cost. The behavioral semantic parsing module is used to construct a hierarchical behavioral recognition model. It identifies basic movement patterns by analyzing the kinematic parameter sequence of the target trajectory, identifies specific types of calls by analyzing the time-frequency pattern of acoustic features, and infers high-level behavioral semantics based on the identified basic actions and acoustic events, combined with habitat type, time information and spatial relationships between multiple targets.

5. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 4, characterized in that, The dynamic task planning and feedback control subsystem includes an abnormal behavior early warning device, an interest region identifier, and an adaptive trajectory planner. The abnormal behavior early warning device is used to monitor the output of the behavior semantic parsing module in real time and generate an alarm when a preset abnormal behavior pattern is detected. The interest region identifier is used to dynamically mark the interest region on the 3D map based on the target spatial distribution density map output by the multi-target tracking module, the location of newly discovered targets, and the high uncertainty region in the evidence theory fusion module. The adaptive trajectory planner is used to replan the UAV's flight trajectory, hovering position, and sensor working mode in real time with the optimization objective of maximizing the amount of information acquired from the interest region per unit time.

6. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 5, characterized in that, The construction process of the basic probability allocation function in the evidence theory fusion module is as follows: First, the original detection output of each sensor is normalized to transform it into a support measure for the propositions that the target exists, the target belongs to a certain species, or the target is located in a certain region. Secondly, each sensor is assigned a reliability weight calculated by combining the precision and recall of the sensor on a known truth dataset during the offline calibration phase. Finally, the normalized support is multiplied by the reliability weight, and a portion of the probability quality is reserved for uncertain propositions, thereby completing the construction of the basic probability assignment function.

7. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 6, characterized in that, In the graph-optimized multi-target tracking module, the optimization solution of the spatiotemporal graph model adopts a cascaded matching strategy that combines the Hungarian algorithm and the minimum cost flow algorithm. First, the Hungarian algorithm is used to perform preliminary association between two consecutive frames based on appearance and motion cost. Second, for the detection nodes and trajectories that fail to be associated, they are included in a sliding time window containing multiple frames and the minimum cost flow algorithm is used for global optimization.

8. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 7, characterized in that, The composite behavior inference layer of the behavior semantic parsing module integrates a time series neural network model based on an attention mechanism. The input of the model is the action encoding sequence, acoustic event encoding sequence, and habitat context feature vector output by the basic action recognition layer within a past time window. The model automatically learns the importance weights of the input features at different time steps for the current behavior classification through the attention layer and outputs the probability distribution of multiple preset composite behavior categories.

9. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 8, characterized in that, The optimization objective function of the adaptive trajectory planner is as follows: under the constraints of UAV endurance and flight safety, maximize the sum of coverage quality for all regions of interest. The coverage quality is defined as the product of the negative exponential function of the time required to reach the region, the information value weight of the region, and the information gain that can be obtained by using a specific sensor mode in the region. The information value weight is dynamically assigned by the region of interest identifier based on the target density, target sparseness, and uncertainty within the region. The optimization problem is solved in the rolling time domain through a model predictive control framework.

10. The wildlife dynamic monitoring system based on the fusion of UAV visual and remote sensing data according to claim 9, characterized in that, In the evidence theory fusion module, for the visible light detector, the basic probability assignment is based on the matching degree between the confidence score of the target bounding box and the morphological features. For thermal infrared detectors, the basic probability assignment is based on the temperature uniformity, size stability, and shape regularity of the hot spot region; for acoustic detectors, the basic probability assignment is based on the species matching probability of the acoustic signature and the covariance of the sound source location estimation.