Intelligent monitoring method and system for forest wild animals based on multi-source data fusion
By combining ground-based infrared cameras with aerial drones to perform multi-source data fusion monitoring, the problems of limited monitoring range and delayed response in existing technologies have been solved, enabling rapid and comprehensive wildlife monitoring and scientific management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YANAN UNIV
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies for wildlife monitoring suffer from limitations in monitoring scope, delayed response, and insufficient data integration, resulting in inadequate timeliness and scientific rigor in conservation management.
A smart monitoring method for forest wildlife based on multi-source data fusion is adopted, which combines a ground infrared camera network with aerial drone inspections. Through comprehensive analysis of multimodal image data and geographic environment data, species identification and habitat assessment are carried out using cross-modal attention networks and deep learning models to generate decision support information.
It enables rapid monitoring with minute-level response, expands coverage from single points to regions, enhances the scientific nature of monitoring and the foresight of management, and provides a complete closed loop from data perception to management decision-making.
Smart Images

Figure CN121660505B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wildlife monitoring and ecological protection technology, specifically relating to an intelligent monitoring method and system for forest wildlife based on multi-source data fusion. Background Technology
[0002] Currently, wildlife conservation has become a crucial issue for global ecological sustainability. Accurately understanding wildlife population dynamics, distribution ranges, and habitat conditions is fundamental to developing effective conservation strategies. With the rapid development of remote sensing technology, the Internet of Things, and artificial intelligence, the field of ecological monitoring is evolving towards automation, intelligence, and integration, providing significant technological supplements and transformative opportunities for traditional monitoring methods that rely on manpower and single-method approaches.
[0003] In current technologies, wildlife monitoring primarily relies on manual ground surveys, automatic imaging with fixed-point infrared cameras, and transect methods. Infrared cameras, triggered by thermal sensing, can record mammal and bird activity unattended and are currently one of the standardized tools for terrestrial wildlife monitoring. Meanwhile, drone technology, due to its flexibility, maneuverability, and wide field of view, is beginning to be applied to animal surveys in open areas or sparse woodlands, carrying visible light or thermal imaging sensors for aerial patrols. Furthermore, satellite remote sensing technology can provide extensive habitat background information such as land cover and vegetation indices. These technologies are used individually or in combination in practice.
[0004] However, existing technologies still have significant problems. Ground-based infrared cameras have extremely limited monitoring range, failing to cover canopy species, and data retrieval is delayed, making it difficult to detect abnormal events in a timely manner. Drone monitoring is mostly conducted independently, lacking real-time coordination with ground-based sensor networks, making it impossible to achieve trigger-based precise surveys, and analysis of single data sources is prone to missed detections or misjudgments. The data generated by various technologies differ greatly in spatiotemporal scale, format, and resolution, lacking effective fusion analysis methods, making it difficult to form a holistic and dynamic assessment of species distribution, population status, and habitat quality. This situation of low monitoring efficiency, incomplete spatial coverage, slow system response, and insufficient data fusion seriously restricts the timeliness and scientific nature of conservation management. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention provides an intelligent monitoring method and system for forest wildlife based on multi-source data fusion. The technical problem to be solved by this invention is achieved through the following technical solution:
[0006] This invention provides an intelligent monitoring method for forest wildlife based on multi-source data fusion, comprising:
[0007] Step 1: Acquire multi-source data of the forest area, including ground infrared image data, multimodal image data, multispectral image data and geographic environment data, including visible light image data and thermal imaging image data;
[0008] Step 2: Based on the ground infrared image data, when a key protected species or abnormal activity event is detected, an emergency response command is generated and a drone is dispatched to fly to the incident location to track and photograph, and acquire real-time multimodal image data;
[0009] Step 3: After performing pixel-level preliminary spatiotemporal alignment on the ground infrared image data, the multimodal image data, and the real-time multimodal image data, input them into the pre-trained target detection model for species identification and behavior classification;
[0010] Step 4: After performing preliminary spatiotemporal alignment of the multispectral image data and the geographic environment data at the grid level, input them into the pre-trained habitat analysis model to conduct habitat suitability assessment;
[0011] Step 5: Based on the results of species identification and behavioral classification and habitat suitability assessment, conduct a comprehensive correlation analysis to generate decision support information;
[0012] The target detection model includes a feature fusion module and a target detection module;
[0013] The feature fusion module adopts a dual-branch cross-modal attention network structure, including: an implicit spatiotemporal alignment module, a first feature extraction module, a second feature extraction module, and a cross-modal cross attention module;
[0014] The output of the implicit spatiotemporal alignment module is connected to the input of the first feature extraction module and the second feature extraction module, respectively. The outputs of the first feature extraction module and the second feature extraction module are both connected to the input of the cross-modal cross-attention module. The output of the cross-modal cross-attention module is connected to the input of the target detection module.
[0015] The pre-aligned visible light image data, thermal imaging image data, and terrestrial infrared image data are input into the implicit spatiotemporal alignment module, which uses deformable convolutional layers to perform spatiotemporal alignment on the input image data. The first feature extraction module uses ResNet to extract visual features from the visible light image and the terrestrial infrared image. The second feature extraction module extracts thermal features from the thermal imaging image based on MobileNetV3 and a temperature attention mechanism. The cross-modal cross-attention module uses a bidirectional attention mechanism to enhance the interaction between the visual features and the thermal features, and uses an adaptive gating unit to fuse the enhanced visual features and thermal features to generate fused features.
[0016] The target detection module includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to extract multi-scale features from the received fused features to obtain multi-scale features. The neck network is used to extract and fuse shallow detail features and deep semantic features from the multi-scale features to obtain enhanced multi-scale features. The detection head is used to predict and output species identification and behavior classification results based on the enhanced multi-scale features.
[0017] The present invention also provides an intelligent monitoring system for forest wildlife based on multi-source data fusion, wherein the intelligent monitoring method for forest wildlife based on multi-source data fusion described in any of the above embodiments includes:
[0018] The data acquisition module is used to acquire multi-source data of the forest area. The multi-source data includes ground infrared image data, multimodal image data, multispectral image data, and geographic environment data. The multimodal image data includes visible light image data and thermal imaging image data.
[0019] Edge computing nodes are used to receive and preprocess the ground infrared image data, and generate emergency response instructions when key protected species or abnormal activity events are detected based on the ground infrared image data.
[0020] The central data processing module is communicatively connected to the edge computing node; the central data processing module includes:
[0021] The drone dispatch unit is used to dispatch drones to the incident location for tracking and filming according to the emergency response instructions, so as to obtain real-time multimodal image data.
[0022] The data processing unit is used to perform pixel-level preliminary spatiotemporal alignment on the ground infrared image data, the multimodal image data, and the real-time multimodal image data, and then perform species identification and behavior classification based on the embedded pre-trained target detection model; it is also used to perform grid-level preliminary spatiotemporal alignment on the multispectral image data and the geographic environment data, and then perform habitat suitability assessment based on the embedded pre-trained habitat analysis model.
[0023] The analysis and decision-making unit is used to generate decision support information by conducting comprehensive correlation analysis based on species identification and behavioral classification results and habitat suitability assessment results.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] 1. This invention presents an intelligent monitoring method for forest wildlife based on multi-source data fusion. It constructs an integrated air-ground three-dimensional observation system by combining a network of fixed-point infrared cameras on the ground with mobile unmanned aerial vehicle (UAV) inspection units in the air. This not only inherits the advantages of infrared cameras in nighttime and concealed environments but also utilizes UAVs to compensate for blind spots in canopy and open areas. Furthermore, it achieves a shift from passive recording to active response through an event-triggered mechanism. The synergy of multi-source data expands the monitoring range from single points to entire regions, and reduces response time from days or weeks to minutes, significantly overcoming the fundamental shortcomings of traditional monitoring methods, such as limited coverage and delayed response.
[0026] 2. The intelligent monitoring method for forest wildlife based on multi-source data fusion of the present invention goes beyond data collection and identification. It further integrates discrete species observation data with continuous habitat environmental data through comprehensive correlation analysis, and generates conservation decision-making suggestions based on the analysis results, providing a complete closed loop from data perception to management decision-making, thereby enhancing the scientific nature and forward-looking nature of conservation management.
[0027] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of an intelligent monitoring method for forest wildlife based on multi-source data fusion provided in an embodiment of the present invention;
[0029] Figure 2 This is a schematic diagram of the structure of a target detection model provided in an embodiment of the present invention;
[0030] Figure 3 This is a structural block diagram of an intelligent monitoring system for forest wildlife based on multi-source data fusion, provided in an embodiment of the present invention.
[0031] Figure 4 This is a graph showing the mean accuracy of each method on a test set of common species;
[0032] Figure 5 This is a simulation example diagram of the target detection model of the present invention, using the North China leopard as an example;
[0033] Figure 6 This is a habitat suitability probability map generated by the habitat assessment model of the present invention, using the North China leopard as an example. Detailed Implementation
[0034] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following describes in detail, with reference to the accompanying drawings and specific embodiments, a method and system for intelligent monitoring of forest wildlife based on multi-source data fusion proposed according to the present invention.
[0035] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and concrete understanding can be gained of the technical means and effects adopted by the present invention to achieve its intended purpose. However, the accompanying drawings are for reference and illustration only and are not intended to limit the technical solutions of the present invention.
[0036] In a first aspect, embodiments of the present invention provide an intelligent monitoring method for forest wildlife based on multi-source data fusion. Please refer to [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of an intelligent monitoring method for forest wildlife based on multi-source data fusion, provided by an embodiment of the present invention. Figure 1 As shown in the figure, the intelligent monitoring method for forest wildlife based on multi-source data fusion in this embodiment includes the following steps:
[0037] Step 1: Acquire multi-source data of the forest area. Multi-source data includes ground infrared image data, multimodal image data, multispectral image data, and geographic environment data. Multimodal image data includes visible light image data and thermal imaging image data.
[0038] In an optional embodiment, step 1 includes:
[0039] Step 1.1: Acquire ground infrared image data using a network of infrared cameras deployed on the forest floor;
[0040] Step 1.2: Using the visible light sensor, thermal infrared sensor and multispectral sensor on the UAV, acquire multimodal image data and multispectral image data of the forest area through periodic inspections;
[0041] Step 1.3: Obtain the geographical environment data of the forest area using satellite remote sensing data acquired by the satellite remote sensing platform.
[0042] Optionally, geographic environmental data, including land cover and digital elevation models of the monitored area, can provide information on topography, location, and human disturbance. Multispectral image data can provide information on vegetation, moisture, and surface temperature.
[0043] Step 2: Based on ground infrared image data, when a key protected species or abnormal activity event is detected, an emergency response command is generated and a drone is dispatched to fly to the incident location to track and photograph, and acquire real-time multimodal image data.
[0044] In an optional embodiment, step 2 includes:
[0045] Step 2.1: Use edge computing nodes to perform real-time target detection on ground infrared image data. When a preset key protected species or abnormal activity event is identified, an emergency response command is generated.
[0046] Optionally, edge computing nodes can be deployed near the monitoring area to process the transmitted ground infrared image data in real time.
[0047] Since infrared cameras have basic motion detection capabilities, they can trigger shooting when they sense heat and movement, generating images or short video clips. The ground infrared image data captured by the infrared camera can be pushed to edge computing nodes in real time via 4G or LoRa (Long Range Radio) wireless networks.
[0048] Optionally, a lightweight deep learning model can be run on the edge computing node to perform real-time lightweight target detection on ground infrared image data transmitted from infrared cameras. For example, the lightweight deep learning model could be a pruned and quantized YOLOv5s model specifically optimized for training on common protected species in forest areas, such as the North China leopard and brown eared pheasant, and unusual targets such as humans, vehicles, and fire sources. When a preset protected species (such as the North China leopard) or unusual activity (such as signs of poaching) is detected, an emergency response command is generated.
[0049] Step 2.2: Based on the emergency response instructions and the real-time status information of all drones, select drones from the schedulable drone cluster to perform tracking and shooting tasks, and obtain real-time multimodal image data of the incident location.
[0050] In this embodiment, the real-time status information of the drone includes the drone's geographical location, remaining battery power, current task load, and communication link quality.
[0051] In this embodiment, the monitoring system can dispatch the nearest drone to the incident location based on the drone's real-time location, battery level, and other information. Upon arrival, the drone first uses a thermal imaging sensor for wide-area scanning and positioning, then switches to a visible light zoom camera for detailed investigation and tracking, and transmits real-time multimodal image data back. This process achieves a minute-level closed loop from ground-based perception to aerial response, significantly improving the proactive detection and rapid response capabilities for critical events.
[0052] Step 3: After performing pixel-level preliminary spatiotemporal alignment of ground infrared image data, multimodal image data, and real-time multimodal image data, input them into the pre-trained target detection model for species identification and behavior classification.
[0053] In an optional embodiment, pixel-level preliminary spatiotemporal alignment is performed on the terrestrial infrared image data, multimodal image data, and real-time multimodal image data, including:
[0054] S1: Based on the trigger time of the ground infrared camera, select the visible light image data and thermal imaging image data of the UAV within the preset time window before and after the trigger time and match them with the ground infrared image data.
[0055] Optionally, all UAV data within a time window of [T0-30s, T0+30s] can be selected, using the infrared camera trigger time T0 as a reference. For video data, the corresponding frame number can be calculated based on the frame rate.
[0056] S2: Perform orthorectification on the selected UAV's visible light image data and thermal imaging image data to eliminate perspective differences and geometric distortions caused by terrain.
[0057] Understandably, infrared cameras have a fixed ground-based viewpoint, resulting in perspective distortion. Drones, on the other hand, have a 45° oblique viewpoint in the air, providing orthographic projection. This difference in perspective causes the same object to appear in different positions and shapes in the images. Therefore, orthographic correction is needed to eliminate the perspective difference and the geometric distortion caused by terrain.
[0058] S3: Feature points are detected and matched in the overlapping area of ground infrared image data, corrected visible light image data, and thermal imaging image data. The homography matrix is estimated by the RANSAC (Random Sample Consensus) algorithm to achieve sub-pixel-level spatial alignment of the image.
[0059] It should be noted that, in this embodiment, basic hardware alignment is required beforehand during the data acquisition phase. Hardware time synchronization and spatial reference unification are performed on all monitoring devices. For example, time synchronization includes: connecting all monitoring devices to a GPS (Global Positioning System) timing module, using PPS (Pulse Per Second) for hardware-level synchronization, and employing NTP (Network Time Protocol) / PTP (Precision Time Protocol) network time protocols to maintain system clock consistency; each data frame is accompanied by a UTC (Coordinated Universal Time) timestamp accurate to milliseconds. Spatial reference unification includes: all monitoring devices using the CGCS2000 national geodetic coordinate system; infrared cameras using RTK (Real-timekinematic) to measure coordinates (accuracy ±2cm); and UAVs recording POS (Position and Orientation System) data, including position and attitude angles.
[0060] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a target detection model provided in an embodiment of the present invention, as shown below. Figure 2 As shown, in this embodiment, the target detection model includes a feature fusion module and a target detection module. The feature fusion module adopts a dual-branch cross-modal attention network structure, including: an implicit spatiotemporal alignment module, a first feature extraction module, a second feature extraction module, and a cross-modal cross-attention module. The output of the implicit spatiotemporal alignment module is connected to the inputs of the first and second feature extraction modules, respectively. The outputs of both the first and second feature extraction modules are connected to the input of the cross-modal cross-attention module, and the output of the cross-modal cross-attention module is connected to the input of the target detection module.
[0061] Specifically, the pre-aligned visible light image data, thermal imaging image data, and terrestrial infrared image data are input into an implicit spatiotemporal alignment module, which uses deformable convolutional layers to perform spatiotemporal alignment on the input image data. The first feature extraction module uses ResNet to extract visual features from the visible light image and the terrestrial infrared image. The second feature extraction module extracts thermal features from the thermal imaging image based on MobileNetV3 and a temperature attention mechanism. The cross-modal cross-attention module achieves interactive enhancement of visual features and thermal features through a bidirectional attention mechanism, and generates fused features by fusing the enhanced visual features and thermal features through an adaptive gating unit.
[0062] In this embodiment, an implicit spatiotemporal alignment module is added before the traditional dual-branch network. This module uses deformable convolutional layers to learn and compensate for residual geometric deformations between different sensor data, so as to further align the initially aligned visible light image data, thermal imaging image data and ground infrared image data.
[0063] For example, the first feature extraction module can use the first three layers of ResNet-34 to extract visual features. The second feature extraction module can use a lightweight MobileNetV3 structure, with a temperature attention gating module inserted after its first bottleneck block. This temperature attention gating module generates an attention map by calculating local temperature statistics to enhance the animal's heat source and suppress environmental thermal noise. The cross-modal attention module can use an 8-head attention mechanism to achieve bidirectional interaction between visual and thermal features.
[0064] In this embodiment, the target detection module includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to extract multi-scale features from the received fused features to obtain multi-scale features; the neck network is used to extract and fuse shallow detail features and deep semantic features from the multi-scale features to obtain enhanced multi-scale features; the detection head is used to predict and output species identification and behavior classification results based on the enhanced multi-scale features.
[0065] For example, the object detection module can be an improved YOLOv7-tiny network. In the backbone network, the 6×6 standard convolution in the first layer of the original YOLOv7-tiny network is replaced with a Focus module. This module reassembles adjacent pixels in the spatial dimension of the input image into the channel dimension through slicing operations, achieving parameter-free downsampling, reducing computational cost and avoiding early feature loss. The CSPDarknet53-tiny structure is adopted as the backbone core. This structure uses Cross Stage Partial Connections (CSP) design to divide the feature flow into main branches and residual branches, effectively splitting gradients, alleviating the gradient vanishing problem and improving training efficiency. Three feature extraction stages are constructed, each outputting feature maps at different scales, specifically for capturing individuals with significant size differences in wild animal targets.
[0066] In the neck network portion, a spatial pyramid pooling module is added at the end of the deep feature path. This module uses three different scales of max pooling layers (5×5, 9×9, and 13×13) in parallel, concatenates the results with the original features, and then processes them through a CSP structure. This significantly expands the model's receptive field without significantly increasing computational complexity, enabling it to better adapt to contextual information perception of targets of different sizes, from large mammals to small birds. An enhanced path aggregation network is adopted, which combines a bidirectional fusion mechanism of feature pyramid network (FPN) and path aggregation network (PAN). This not only achieves top-down semantic information propagation (upsampling of deep features and fusion with shallow features) but also bottom-up detail information supplementation (downsampling of shallow features and fusion with deep features). This bidirectional multi-scale feature interaction fully integrates high-resolution details and rich semantic information, greatly improving the detection capability of wild animals in complex forest environments, especially those that are occluded or small in size.
[0067] In the detection head section, the coupled detection head in the original YOLOv7-tiny network is replaced with a decoupled detection head, separating the classification task from the bounding box regression task. A lightweight Transformer encoder layer (containing 4 attention heads) is added before each detection head to enhance the long-range contextual association capability of features, thereby improving detection performance under occlusion conditions. A new parallel behavior classification branch is added, which receives features from the neck network and outputs the probabilities of 6 preset behaviors, such as foraging, reproduction, resting, moving, vigilance, and socializing, through a three-layer fully connected network.
[0068] In this embodiment, the target detection model ultimately outputs the species category, bounding box coordinates, confidence score, and behavior classification result for each detected target.
[0069] Step 4: After performing preliminary spatiotemporal alignment of the multispectral image data and geographic environment data at the grid level, input them into the pre-trained habitat analysis model to conduct habitat suitability assessment.
[0070] In an optional embodiment, preliminary spatiotemporal alignment of multispectral image data and geographic environment data at the grid level is performed, including:
[0071] Step a: Perform temporal synthesis, spatial resampling and gridding, spatial range cropping and spectral normalization on the multispectral image data in sequence to obtain a subset of multispectral data.
[0072] Optionally, based on a preset analysis time window, the maximum value synthesis method can be used to generate a synthetic image product representing that time window from the multispectral image data of multiple time phases; the synthetic multispectral image data can be resampled from the original sensor resolution to a preset standard spatial grid scale, such as resampled to a 10-meter grid, and bilinear interpolation can be used to maintain spectral continuity; the resampled data can be cropped according to the boundary of the study area to ensure spatial consistency; based on the physical reflectance range of the multispectral bands, the data of each band can be normalized to generate the final multispectral data subset.
[0073] Step b: Perform time attribute labeling, spatial resampling and gridding, spatial range clipping and variable standardization on the geographic environment data in sequence to obtain a subset of geographic environment data.
[0074] Optionally, the time attributes of each layer of geographic environmental data can be uniformly labeled according to their temporal characteristics. For dynamically changing data, the actual observation time can be labeled, and for static or slowly changing data, the long-term average state can be labeled. The geographic environmental data of each layer can be resampled from their original spatial resolution to the same standard spatial grid scale as the multispectral data subset according to the data type. For continuous variables, the bilinear interpolation method is used, and for categorical variables, the nearest neighbor interpolation method is used. The resampled data of each layer is cropped according to the same study area boundary. The geographic environmental variables of each layer are standardized based on statistical distribution to eliminate dimensional differences and generate the final geographic environmental data subset.
[0075] In this embodiment, the multispectral data subset and the geographic environment data subset are aligned in terms of spatial extent, grid scale, and temporal representativeness.
[0076] Understandably, before performing initial alignment, it is necessary to convert the coordinate system of the multispectral image data and geographic environment data into a preset geodetic coordinate system.
[0077] In this embodiment, the habitat analysis model includes a multi-source environmental variable extraction module and a habitat analysis module. The multi-source environmental variable extraction module includes a spectral feature encoding module, a geographic feature encoding module, and a feature fusion layer. The spectral feature encoding module is used to extract spectral texture features from a subset of multispectral data to obtain a spectral feature map; the geographic feature encoding module is used to extract geographic environmental features from a subset of geographic environmental data to obtain a geographic feature map; the feature fusion layer is used to stitch and convolve the spectral feature map and the geographic feature map to obtain a fused feature map containing multi-dimensional environmental information; the fused feature map serves as the input to the habitat analysis module.
[0078] For example, the spectral feature encoding module can use a simple CNN containing three convolutional layers; the geographic feature encoding module can use two fully connected layers; and the feature fusion layer can concatenate the outputs of the two and then fuse them through a 1×1 convolutional layer.
[0079] In this embodiment, the habitat analysis module adopts a deep learning model with a U-Net++ architecture, including an encoder, a decoder, and an output layer. The encoder includes four levels of downsampling units, each containing two convolutional layers and a max-pooling layer connected in sequence, used to progressively extract multi-scale environmental features from the fused feature map; the decoder includes four levels of upsampling units, each restoring spatial resolution through transposed convolution and fusing with the features of the corresponding level of the encoder through skip connections, outputting a reconstructed feature map; the output layer processes the reconstructed feature map with convolution and a sigmoid activation function to obtain a habitat suitability probability map as the habitat suitability assessment result.
[0080] It is understandable that, for the U-Net++ network structure in this embodiment, the original simple concatenation operation in the skip connection between the encoder and decoder can be replaced with a skip connection module that includes a channel attention mechanism. This module first performs global average pooling on the encoder features, generates channel weights through two fully connected layers, reweights the features, and then concatenates them with the decoder features, thereby more effectively fusing multi-scale environmental features. The output layer generates a habitat suitability probability map through the Sigmoid activation function, where each pixel value p∈[0, 1] represents the probability or degree of habitat suitability for the target species (such as the North China leopard) at that geographic coordinate point (corresponding to a specific location in the real world).
[0081] In this embodiment, (0.8-1.0] represents core suitable habitat; (0.6-0.8] represents relatively suitable habitat; (0.4-0.6] represents generally suitable habitat; (0.2-0.4] represents marginal habitat; and (0.0-0.2] represents unsuitable habitat.
[0082] In this embodiment, a dedicated deep learning model was designed to address the challenges of complex forest environments, diverse targets, and heterogeneous multi-source data. For target detection, an improved bi-branch cross-modal attention network effectively fuses visible light texture features and thermal imaging temperature features. Furthermore, implicit alignment and Transformer enhancement modules significantly improve the detection capabilities for small targets, occluded targets, and intermodal target associations. For habitat analysis, the improved U-Net++ model deeply fuses spectral and geographic environmental information to accurately quantify habitat suitability.
[0083] Step 5: Based on the results of species identification and behavioral classification and habitat suitability assessment, conduct a comprehensive correlation analysis to generate decision support information.
[0084] In an optional embodiment, step 5 includes: conducting a spatial joint analysis of species distribution and habitat quality, as well as a behavioral pattern and environmental correlation analysis, based on the species identification and behavioral classification results and habitat suitability assessment results, and generating a forest area spatial optimization and management monitoring plan based on the analysis results.
[0085] Optionally, the species identification points output by the target detection model can be spatially overlaid with the habitat suitability probability map for analysis. Specifically, this can include: extracting habitat suitability probability values at the actual locations of species occurrences; calculating the average suitability probability corresponding to high-frequency occurrence points to verify the accuracy of the habitat analysis model's predictions; calculating the density of actual species occurrence points within each suitability level area to identify whether species are fully utilizing high-suitability areas and whether they are forced to utilize low-suitability areas; and identifying areas with high habitat suitability probabilities but where species have not appeared for a long time, and analyzing possible reasons.
[0086] Optionally, the species behavior classification results output by the target detection model can be spatially overlaid with the habitat suitability probability map to analyze the environmental background of specific behaviors. Specifically, this can include: statistically analyzing the environmental characteristics of different behaviors such as foraging, reproduction, and resting to establish behavioral preference environmental patterns; analyzing the habitat suitability and surrounding environmental stability of locations marked as breeding behaviors; and comparing the differences in species behavior between disturbed and undisturbed areas to quantify the degree of disturbance on behavioral patterns.
[0087] Optionally, conservation effectiveness can be assessed based on the spatial pattern of habitat suitability probability maps and the actual distribution of species. Specifically, this can include: overlaying existing protected area boundaries with the distribution of highly suitable habitats to identify highly suitable areas not covered by protected areas; calculating the habitat fragmentation index based on the spatial connectivity of highly suitable areas to identify key gaps in ecological corridor construction; and overlaying species occurrence points with layers of human activity intensity to quantify the impact radius of different disturbance sources on species distribution.
[0088] Optionally, the results of multiple analyses can be integrated to generate trend-based decision support information. Specifically, this may include: comparing habitat suitability probability maps of different periods to identify areas where suitability has increased or decreased, and analyzing the driving factors of change; analyzing the spatiotemporal correlation between changes in species distribution range and population size and changes in habitat quality; and generating tiered early warning information when the area of highly suitable habitat patches shrinks beyond a threshold, key breeding grounds are invaded, or species are forced to spread to less suitable areas.
[0089] In this embodiment, structured decision support information can be output based on the above analysis results, specifically including: proposing a boundary adjustment plan for the protected area, suggestions for adding new protected areas, and specific planning routes for ecological corridors.
[0090] For example, areas with p ≥ 0.8 are included in the core area for protection zone boundary optimization; highly suitable patches (p ≥ 0.6) are connected to form ecological corridor planning; restoration is prioritized for areas with p ≤ 0.4 and p < 0.6; and construction is avoided in highly suitable areas (p ≥ 0.7).
[0091] In this embodiment, based on the above analysis results, differentiated patrol frequencies, intensity of human activity control, and habitat restoration measures can be proposed for different regions; or, based on new discoveries on species distribution and habitat use, plans for adjusting the deployment of infrared camera networks and key areas and time periods for drone inspections can be proposed.
[0092] This invention presents an intelligent monitoring method for forest wildlife based on multi-source data fusion. By combining a network of fixed-point infrared cameras on the ground with mobile unmanned aerial vehicle (UAV) inspection units in the air, it constructs an integrated sky-ground three-dimensional observation system. This not only inherits the advantages of infrared cameras in nighttime and concealed environments but also utilizes UAVs to compensate for blind spots in canopy and open areas. Furthermore, it achieves a shift from passive recording to proactive response through an event-triggered mechanism. The synergy of multi-source data expands the monitoring range from single points to entire regions, and reduces response time from days or weeks to minutes, significantly overcoming the fundamental shortcomings of traditional monitoring methods, such as limited coverage and delayed response.
[0093] Secondly, embodiments of the present invention provide an intelligent monitoring system for forest wildlife based on multi-source data fusion, applicable to the intelligent monitoring method for forest wildlife based on multi-source data fusion provided in the first aspect. Please refer to... Figure 3 , Figure 3 This is a structural block diagram of an intelligent monitoring system for forest wildlife based on multi-source data fusion, as provided in an embodiment of the present invention. Figure 3 As shown, the intelligent monitoring system for forest wildlife based on multi-source data fusion in this embodiment includes: a data acquisition module, an edge computing node, and a central data processing module.
[0094] The data acquisition module is used to acquire multi-source data of the forest area, including ground infrared image data, multimodal image data, multispectral image data and geographic environment data. The multimodal image data includes visible light image data and thermal imaging image data.
[0095] In this embodiment, the data acquisition module includes: an infrared camera network deployed on the forest floor, a drone equipped with a visible light sensor, a thermal infrared sensor and a multispectral sensor, and an interface module for accessing satellite remote sensing data.
[0096] Edge computing nodes are used to receive and preprocess ground infrared image data, and generate emergency response commands when key protected species or abnormal activity events are detected based on the ground infrared image data.
[0097] In this embodiment, edge computing nodes can be deployed near the monitoring area. A lightweight deep learning model runs on these nodes to perform real-time, lightweight target detection on ground infrared image data transmitted from infrared cameras. This lightweight deep learning model can be a pruned and quantized YOLOv5s.
[0098] The central data processing module is connected to the edge computing nodes; the central data processing module includes: a drone scheduling unit, a data processing unit, and an analysis and decision-making unit.
[0099] In this embodiment, the UAV scheduling unit is used to schedule UAVs to fly to the incident location for tracking and shooting according to emergency response instructions to obtain real-time multimodal image data; the data processing unit is used to perform pixel-level preliminary spatiotemporal alignment of ground infrared image data, multimodal image data and real-time multimodal image data, and then perform species identification and behavior classification according to the embedded pre-trained target detection model; the data processing unit is used to perform grid-level preliminary spatiotemporal alignment of multispectral image data and geographic environment data, and then perform habitat suitability assessment according to the embedded pre-trained habitat analysis model; the analysis and decision unit is used to perform comprehensive correlation analysis to generate decision support information based on the species identification and behavior classification results and the habitat suitability assessment results.
[0100] In this embodiment, the data processing unit is equipped with a high-performance GPU (Graphics Processing Unit) to run target detection models and habitat analysis models, completing species and behavior identification and habitat assessment. The analysis and decision-making unit performs spatial correlation analysis based on a GIS (Geographic Information System) platform and integrates a decision support knowledge base, which can automatically generate structured monitoring reports and management recommendations.
[0101] This embodiment of the intelligent monitoring system for forest wildlife, which integrates multi-source data, utilizes edge computing nodes for real-time event detection and rapid response at the monitoring front end, meeting the stringent low-latency requirements of scenarios such as poaching early warning and fire detection. Big data storage, model training, and deep analysis are performed on a central cloud platform, solving the problem of limited computing power on edge devices. This collaborative architecture reduces network transmission load and allows the system to flexibly connect to new data sources or analytical models. It is easily customized and expanded according to different protected area needs and budget levels, demonstrating promising application prospects and widespread value.
[0102] For details regarding the intelligent monitoring system for forest wildlife based on multi-source data fusion and its corresponding beneficial effects, please refer to the relevant content on the intelligent monitoring method for forest wildlife based on multi-source data fusion provided in the first aspect, which will not be repeated here.
[0103] Furthermore, the effectiveness of the intelligent monitoring method for forest wildlife based on multi-source data fusion in this embodiment is verified and illustrated through simulation experiments.
[0104] 1. Experimental Setup
[0105] Study area: The core area of a nature reserve on the Loess Plateau, covering an area of approximately 200 square kilometers, including various habitats such as woodland, shrubland, and grassland.
[0106] Data preparation: 50 infrared cameras were deployed to acquire ground monitoring data for 6 months; UAVs conducted grid-based inspections twice a month; and Sentinel-2 multispectral data and ALOS 12.5-meter resolution digital elevation model data were acquired during the same period.
[0107] Comparison Method: The method proposed in this invention is compared with the following methods, referred to as this invention: 1) The method using only infrared camera data, referred to as IR-Only; 2) The method using infrared camera and UAV visible light data, referred to as IR+RGB; 3) The method using the data of this invention but without improving the network model, referred to as Ours-Base, that is, using standard YOLOv7-tiny and U-Net.
[0108] 2. Species identification performance verification
[0109] The methods were validated on a test set including 10 common species such as the North China leopard, brown eared pheasant, and wild boar. Please see [link to relevant documentation]. Figure 4 , Figure 4 This is a graph showing the mean accuracy of each method on a test set of common species. Table 1 shows the mean accuracy of each method on the test set of common species.
[0110] Table 1
[0111]
[0112] As shown in Table 1, multi-source data fusion significantly improved performance, while targeted network structure improvements further tapped into the potential of multi-source data. Ultimately, the mAP of the target recognition model of this invention reached 0.892, an improvement of over 20% compared to a single infrared data baseline, verifying the effectiveness of the method for accurate species identification in complex forest environments. Please refer to [link to relevant documentation]. Figure 5 , Figure 5This is a simulation example of the target detection model of the present invention, using the North China leopard as an example. From left to right, the figure shows an infrared image, a visible light image, a thermal imaging image, and a target detection result image. The target detection result image shows the species category, bounding box coordinates, confidence level, and behavior classification results of the target.
[0113] 3. Validation of Habitat Assessment Accuracy
[0114] Using the North China leopard as an indicator species, 120 actual sighting locations recorded by its GPS (Global Positioning System) collar were used as validation data. The coverage rate of these actual locations for the high-fitness areas (probability > 0.7) predicted by each method is shown in Table 2.
[0115] Table 2
[0116]
[0117] As shown in Table 2, the habitat assessment model of this invention has the highest prediction coverage rate for actual species occurrence points, reaching 91%. This indicates that, compared with traditional methods and standard deep learning models, the improved U-Net++ structure adopted in this invention can more fully utilize multi-source remote sensing and geographic information data to learn more accurate species-environment relationships, thereby generating higher-quality habitat suitability probability maps that better reflect the actual distribution patterns of species, providing a reliable basis for subsequent conservation spatial planning. The habitat suitability probability map generated by the habitat assessment model of this invention, using the North China leopard as an example, is shown below. Figure 6 As shown.
[0118] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device comprising said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. The orientations or positional relationships indicated by terms such as "upper," "lower," "left," and "right" are based on the orientations or positional relationships shown in the accompanying drawings and are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0119] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0120] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for intelligent monitoring of forest wildlife based on multi-source data fusion, characterized in that, include: Step 1: Acquire multi-source data of the forest area, including ground infrared image data, multimodal image data, multispectral image data and geographic environment data, including visible light image data and thermal imaging image data; Step 2: Based on the ground infrared image data, when a key protected species or abnormal activity event is detected, an emergency response command is generated and a drone is dispatched to fly to the incident location to track and photograph, and acquire real-time multimodal image data; Step 3: After performing pixel-level preliminary spatiotemporal alignment on the ground infrared image data, the multimodal image data, and the real-time multimodal image data, input them into the pre-trained target detection model for species identification and behavior classification; Step 4: After performing preliminary spatiotemporal alignment of the multispectral image data and the geographic environment data at the grid level, input them into the pre-trained habitat analysis model to conduct habitat suitability assessment; Step 5: Based on the results of species identification and behavioral classification and habitat suitability assessment, conduct a comprehensive correlation analysis to generate decision support information; The target detection model includes a feature fusion module and a target detection module; The feature fusion module adopts a dual-branch cross-modal attention network structure, including: an implicit spatiotemporal alignment module, a first feature extraction module, a second feature extraction module, and a cross-modal cross attention module; The output of the implicit spatiotemporal alignment module is connected to the input of the first feature extraction module and the second feature extraction module, respectively. The outputs of the first feature extraction module and the second feature extraction module are both connected to the input of the cross-modal cross-attention module. The output of the cross-modal cross-attention module is connected to the input of the target detection module. The pre-aligned visible light image data, thermal imaging image data, and terrestrial infrared image data are input into the implicit spatiotemporal alignment module, which uses deformable convolutional layers to perform spatiotemporal alignment on the input image data. The first feature extraction module uses ResNet to extract visual features from the visible light image and the terrestrial infrared image. The second feature extraction module extracts thermal features from the thermal imaging image based on MobileNetV3 and a temperature attention mechanism. The cross-modal cross-attention module uses a bidirectional attention mechanism to enhance the interaction between the visual features and the thermal features, and uses an adaptive gating unit to fuse the enhanced visual features and thermal features to generate fused features. The target detection module includes a backbone network, a neck network, and a detection head connected in sequence. The backbone network is used to extract multi-scale features from the received fused features to obtain multi-scale features. The neck network is used to extract and fuse shallow detail features and deep semantic features from the multi-scale features to obtain enhanced multi-scale features. The detection head is used to predict and output species identification and behavior classification results based on the enhanced multi-scale features.
2. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 1, characterized in that, Step 1 includes: Step 1.1: Acquire the ground infrared image data using an infrared camera network deployed on the forest floor; Step 1.2: Using the visible light sensor, thermal infrared sensor and multispectral sensor mounted on the UAV, the multimodal image data and multispectral image data of the forest area are acquired through periodic inspections; Step 1.3: Obtain the geographic environment data of the forest area using satellite remote sensing data acquired by the satellite remote sensing platform.
3. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 1, characterized in that, Step 2 includes: Step 2.1: Use edge computing nodes to perform real-time target detection on the ground infrared image data. When a preset key protected species or abnormal activity event is identified, generate an emergency response command. Step 2.2: Based on the emergency response instructions and the real-time status information of all UAVs, select UAVs from the schedulable UAV cluster to perform tracking and shooting tasks, and obtain real-time multimodal image data of the incident location.
4. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 1, characterized in that, Performing pixel-level preliminary spatiotemporal alignment of the ground infrared image data, the multimodal image data, and the real-time multimodal image data includes: S1: Based on the triggering time of the ground infrared camera, select the visible light image data and thermal imaging image data of the UAV within the preset time window before and after the triggering time and match them with the ground infrared image data. S2: Perform orthorectification on the selected UAV's visible light image data and thermal imaging image data to eliminate perspective differences and geometric distortions caused by terrain. S3: Detect and match feature points in the overlapping area of the ground infrared image data, the corrected visible light image data, and the thermal imaging image data. Estimate the homography matrix using the RANSAC algorithm to achieve sub-pixel level spatial alignment of the image.
5. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 1, characterized in that, Performing preliminary spatiotemporal alignment of the multispectral image data and the geographic environment data at the grid level, including: Step a: Perform temporal synthesis, spatial resampling and gridding, spatial range cropping and spectral normalization on the multispectral image data in sequence to obtain a subset of multispectral data; Step b: The geographic environment data is sequentially processed by time attribute labeling, spatial resampling and gridding, spatial range clipping and variable standardization to obtain a subset of geographic environment data; The multispectral data subset is aligned with the geographic environment data subset in terms of spatial extent, grid scale, and temporal representativeness.
6. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 5, characterized in that, The habitat analysis model includes a multi-source environmental variable extraction module and a habitat analysis module; The multi-source environmental variable extraction module includes: a spectral feature encoding module, a geographic feature encoding module, and a feature fusion layer; The spectral feature encoding module is used to extract spectral texture features from the multispectral data subset to obtain a spectral feature map; the geographic feature encoding module is used to extract geographic environmental features from the geographic environment data subset to obtain a geographic feature map; the feature fusion layer is used to concatenate and fusion the spectral feature map and the geographic feature map to obtain a fused feature map containing multi-dimensional environmental information; the fused feature map serves as the input to the habitat analysis module; The habitat analysis module adopts a deep learning model with a U-Net++ architecture, including an encoder, decoder, and output layer; The encoder includes four levels of downsampling units, each containing two convolutional layers and a max-pooling layer connected in sequence, used to progressively extract multi-scale environmental features from the fused feature map; the decoder includes four levels of upsampling units, each restoring spatial resolution through transposed convolution and fusing with the features of the corresponding level of the encoder via skip connections to output a reconstructed feature map; the output layer processes the reconstructed feature map using convolution and a sigmoid activation function to obtain a habitat suitability probability map as the habitat suitability assessment result.
7. The intelligent monitoring method for forest wildlife based on multi-source data fusion according to claim 1, characterized in that, Step 5 includes: Based on the results of species identification and behavioral classification and habitat suitability assessment, a spatial joint analysis of species distribution and habitat quality, as well as an analysis of the correlation between behavioral patterns and the environment, is conducted. Based on the analysis results, a spatial optimization and management monitoring plan for forest areas is generated.
8. An intelligent monitoring system for forest wildlife based on multi-source data fusion, characterized in that, The intelligent monitoring method for forest wildlife based on multi-source data fusion, applicable to any one of claims 1-7, includes: The data acquisition module is used to acquire multi-source data of the forest area. The multi-source data includes ground infrared image data, multimodal image data, multispectral image data, and geographic environment data. The multimodal image data includes visible light image data and thermal imaging image data. Edge computing nodes are used to receive and preprocess the ground infrared image data, and generate emergency response instructions when key protected species or abnormal activity events are detected based on the ground infrared image data. The central data processing module is communicatively connected to the edge computing node; the central data processing module includes: The drone dispatch unit is used to dispatch drones to the incident location for tracking and filming according to the emergency response instructions, so as to obtain real-time multimodal image data. The data processing unit is used to perform pixel-level preliminary spatiotemporal alignment on the ground infrared image data, the multimodal image data, and the real-time multimodal image data, and then perform species identification and behavior classification based on the embedded pre-trained target detection model; it is also used to perform grid-level preliminary spatiotemporal alignment on the multispectral image data and the geographic environment data, and then perform habitat suitability assessment based on the embedded pre-trained habitat analysis model. The analysis and decision-making unit is used to generate decision support information by conducting comprehensive correlation analysis based on species identification and behavioral classification results and habitat suitability assessment results.
9. The intelligent monitoring system for forest wildlife based on multi-source data fusion according to claim 8, characterized in that, The data acquisition module includes: an infrared camera network deployed on the forest floor, a drone equipped with a visible light sensor, a thermal infrared sensor, and a multispectral sensor, and an interface module for accessing satellite remote sensing data.
Citation Information
Patent Citations
Intelligent security video monitoring system for preventing animal attack behaviors
CN120236243A
Super-resolution binocular image generation method and system based on geometric structure consistency
CN120997049A