Intelligent wildlife monitoring equipment and methods based on multimodal fusion and self-organizing networks

CN120876978BActive Publication Date: 2026-08-14BEIJING FORESTRY UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

即使系统集成了多种传感器,也往往停留在各自独立工作、数据简单叠加的阶段,未能充分利用不同模态数据之间的互补性进行深度融合分析,例如利用声音信息辅助视觉定位,或利用视觉信息辅助声音事件分类

Benefits of technology

[0035]与现有技术相比,本发明提供了基于多模态融合与自组网的野生动物智能监测设备,具备以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876978B_ABST
    Figure CN120876978B_ABST
Patent Text Reader

Abstract

This invention relates to intelligent wildlife monitoring equipment and methods based on multimodal fusion and self-organizing networks, belonging to the field of wildlife monitoring technology. By fusing binocular infrared vision with microphone array hearing, this invention can acquire more comprehensive and accurate information on wildlife activity, overcoming the limitations of single sensors. The WS-YOLO algorithm at the device end enables intelligent false trigger screening and species identification, reducing manual intervention and improving data processing efficiency. Self-organizing network technology allows the monitoring system to cover a wider area without network coverage, making equipment deployment more flexible and unrestricted by network infrastructure. By reducing false triggers, implementing local intelligent processing, and achieving efficient data transmission, the costs of data storage, transmission, and subsequent manual analysis are reduced. This invention provides a more powerful technical means for wildlife population dynamic monitoring, behavioral ecology research, and biodiversity conservation, possessing significant ecological protection and scientific research value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wildlife monitoring technology, specifically to an intelligent wildlife monitoring device and method based on multimodal fusion and self-organizing networks. Background Technology

[0002] Existing wildlife monitoring technologies primarily rely on various sensors, including infrared sensors, sound sensors, environmental sensors, and visible light cameras. These technologies aim to understand the species, population size, behavior, and distribution of wildlife by capturing signs of animal activity, such as changes in body temperature, sound signals, or direct images. For example, infrared trigger cameras are among the most widely used tools, triggering capture by sensing the difference between an animal's body temperature and the ambient temperature. Sound monitoring technologies analyze acoustic signals such as animal calls and activity sounds to identify species, assess population density, or study animal behavior. Furthermore, some systems integrate environmental sensors (such as temperature, humidity, and light intensity sensors) to obtain background information about the monitoring site. In recent years, with the development of wireless communication and network technologies, some monitoring systems have begun to possess remote data transmission and networking capabilities, enabling researchers to acquire monitoring data in real-time or near real-time. However, these technologies often focus on acquiring data from a single source or lack in-depth data fusion and collaborative analysis when using multiple sensors.

[0003] Despite the progress made in existing wildlife monitoring technologies, numerous challenges and shortcomings remain in practical applications. First, the limitations of single-sensor monitoring are significant. For example, relying solely on infrared trigger cameras is susceptible to environmental interference (such as direct sunlight or wind rustling through grass), leading to numerous false triggers, wasting storage space and battery power, and increasing the workload of subsequent data filtering. Furthermore, infrared cameras may fail to detect animals effectively when obscured by dense vegetation or when they are stationary. While sound monitoring can compensate for the shortcomings of visual monitoring, it is difficult to accurately locate sound sources when used alone, and it is easily affected by environmental noise, resulting in poor monitoring of animals that are silent or have weak vocalizations. Second, the hardware integration of existing technologies is low. Many multi-sensor monitoring systems simply combine sensor modules with different functions, lacking an integrated hardware design. This not only increases the size and power consumption of the equipment but may also lead to difficulties in accurately synchronizing data from different sensors in time and space, affecting the subsequent data fusion effect. Third, the degree of multimodal data fusion is low. Even when systems integrate multiple sensors, they often remain at the stage of independent operation and simple data overlay, failing to fully utilize the complementarity between different modalities for deep fusion analysis. For example, they cannot utilize sound information to assist visual positioning or visual information to assist sound event classification. Finally, the level of intelligence and networking capabilities need improvement. Many monitoring devices lack the ability to perform intelligent data processing (such as target detection, false trigger screening, and species identification) at the device end, resulting in a large amount of raw data requiring manual processing or transmission back to the cloud for processing, which is inefficient and costly. In remote areas with weak network coverage, data transmission is also a major challenge; existing systems often struggle to achieve efficient and reliable data transmission and collaborative operation between devices.

[0004] To address the aforementioned issues, this invention proposes an intelligent wildlife monitoring device and method based on multimodal fusion and self-organizing networks. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] The purpose of this invention is to propose an intelligent wildlife monitoring device based on multimodal fusion and self-organizing networks to solve the following problems:

[0007] 1. The challenge of integrating and coordinating multimodal sensors. In existing multimodal monitoring systems, different types of sensors, such as visual and auditory sensors, are typically installed separately in a modular fashion, lacking deep integration at the structural, temporal, and power consumption levels. This results in a lack of spatiotemporal consistency among data, making information fusion difficult and hindering the performance of multimodal monitoring. This invention proposes an integrated solution for a binocular infrared trigger camera and microphone array. Through structural optimization, the optical axis and acoustic center are aligned; through hardware synchronization, the temporal alignment of visual and auditory data is achieved; and through system design, overall power consumption is reduced, thereby significantly improving the acquisition efficiency and correlation of multimodal data.

[0008] 2. The problem of stable transmission and collaborative processing of multimodal data in complex field environments. In wildlife habitats lacking communication infrastructure, ad hoc network communication systems suffer from poor transmission stability, high power consumption, and inflexible topology, making it difficult to meet the real-time transmission requirements of multimodal, high-frequency, and large-volume data. This invention addresses this problem by integrating optimized multimodal monitoring equipment with an ad hoc network system to construct a low-power, long-distance communication network supporting dynamic topology for field applications. It also enables collaborative information processing among multiple nodes, supporting intelligent monitoring functions such as multi-point joint sound source localization and cross-node target tracking, thus solving the challenges of data acquisition and network communication in complex environments.

[0009] 3. The problem of intelligent wildlife image recognition and data redundancy filtering on embedded devices. Traditional wildlife image acquisition systems often generate a large number of invalid images due to environmental interference, resulting in data redundancy, wasted transmission resources, and a burden on post-processing. To address this issue, this invention deploys a lightweight and optimized WS-YOLO target detection algorithm on the device side, which can identify animal targets in images in real time on edge devices, automatically filter out falsely triggered images, and complete species classification and identification. This significantly reduces the reliance on backend storage and manual analysis, and improves the overall intelligence and response efficiency of the monitoring system.

[0010] (II) Technical Solution

[0011] To achieve the above objectives, the present invention proposes the following:

[0012] Intelligent wildlife monitoring equipment based on multimodal fusion and self-organizing networks includes:

[0013] The outer casing is used to house and protect the internal modules;

[0014] Solar panels, fixed to the top of the casing, are used to convert solar energy into electrical energy;

[0015] A binocular infrared trigger camera module is mounted side-by-side on the front panel of the housing to collect stereoscopic visual information of wild animals;

[0016] A microphone array opening is provided, with the array located on the front panel of the housing. A microphone array module is installed inside the microphone array opening. The microphone array module is used to collect ambient sounds and acoustic signals emitted by animals, and to locate the sound source.

[0017] The outer casing also houses an active processing unit, a communication and self-organizing network module, and a power management module;

[0018] The main control and processing unit is connected to the binocular infrared trigger camera module and the microphone array module, respectively, to control data acquisition, realize the synchronization of visual and auditory signals, and run target detection and recognition algorithms;

[0019] The communication and self-organizing network module is connected to the main control and processing unit to realize wireless networking and data transmission between devices, and transmits data through multi-hop relay in areas without public network coverage.

[0020] The power management module, connected to the solar panel, is used to power the various functional modules.

[0021] Preferably, the binocular infrared trigger camera module and the microphone array module adopt a hardware synchronization mechanism to ensure that visual data and auditory data are time-aligned during acquisition.

[0022] Preferably, the main control and processing unit is equipped with the WS-YOLO algorithm, which is used to perform false triggering filtering and wildlife species detection and classification on the acquired image data.

[0023] Preferably, the WS-YOLO algorithm is optimized by performing model quantization, pruning, and hardware acceleration on the device side to adapt to the resource limitations of the embedded platform.

[0024] Preferably, the communication and self-organizing network module supports multiple wireless communication protocols and automatically selects or switches communication modes according to the network environment; the wireless communication protocols include LoRa, ZigBee, and Wi-Fi Mesh.

[0025] Preferably, the power management module includes a rechargeable lithium battery and a solar charging controller, and employs an intelligent power consumption management strategy to extend the device's battery life.

[0026] Preferably, the main control and processing unit is also used to extract depth information from the images acquired by the binocular infrared camera, and combine the sound source localization results of the microphone array to perform multimodal data fusion in order to obtain richer target information.

[0027] A wildlife intelligent monitoring system is disclosed, wherein the system connects a binocular infrared trigger camera module, a microphone array module, a main control and processing unit, a communication and self-organizing network module, and a power management module through a self-organizing network to collaboratively complete wildlife monitoring tasks in large-scale areas without network coverage.

[0028] A smart monitoring method for wild animals includes the following:

[0029] Data acquisition is performed by synchronously activating a binocular infrared camera and microphone array via infrared triggering.

[0030] The collected visual and auditory data are timestamped and aligned.

[0031] On the device side, the WS-YOLO algorithm is used to perform false trigger filtering and species identification on image data;

[0032] Effective monitoring data can be stored locally or transmitted to the data center via an ad hoc network.

[0033] Preferably, the method further includes fusing visual analysis results with acoustic information to form a comprehensive monitoring event report.

[0034] (III) Beneficial Effects

[0035] Compared with existing technologies, this invention provides an intelligent wildlife monitoring device based on multimodal fusion and self-organizing networks, which has the following beneficial effects:

[0036] (1) Improved monitoring accuracy and information richness: By combining binocular infrared vision with microphone array hearing, more comprehensive and accurate information on wildlife activities can be obtained, overcoming the limitations of a single sensor.

[0037] (2) Improved intelligence level: The WS-YOLO algorithm on the device side realizes intelligent false trigger screening and species identification, reducing manual intervention and improving data processing efficiency.

[0038] (3) Expanded monitoring range and enhanced deployment flexibility: Self-organizing network technology enables the monitoring system to cover a wider area without network, and the equipment deployment is more flexible and not limited by network infrastructure.

[0039] (4) Reduced operating costs: By reducing false triggers, local intelligent processing and efficient data transmission, the costs of data storage, transmission and subsequent manual analysis are reduced.

[0040] (5) Ecological protection and scientific research value: It provides more powerful technical means for monitoring the dynamics of wild animal populations, behavioral ecology research, biodiversity protection and early warning of human-wildlife conflict, and has important ecological protection and scientific research value. Attached Figure Description

[0041] The present invention is described with reference to the following figures:

[0042] Figure 1 This is a structural diagram of the intelligent wildlife monitoring device based on multimodal fusion and self-organizing network proposed in this invention;

[0043] Figure 2 This is a structural diagram of the WS-YOLO algorithm proposed in Embodiment 2 of the present invention;

[0044] Figure 3 This is a flowchart of the WTConv module processing proposed in Embodiment 2 of the present invention;

[0045] Figure 4 This is a flowchart of the SCSA module processing proposed in Embodiment 2 of the present invention;

[0046] Figure 5 This is a flowchart of the device triggering process proposed in Embodiment 2 of the present invention;

[0047] Figure 6 The image shows the detection results of WS-YOLO proposed in Embodiment 2 of this invention. (a) represents the original image, with the white box indicating a local magnification and the real name of the species in the upper right corner; (b) represents the detection result of the Faster-RCNN algorithm; (c) represents the detection result of the Swin Transformer; (d) represents the detection result of YOLOv11n; and (e) represents the detection result of WS-YOLO.

[0048] Figure 7 The results of input size and speed, input size and accuracy on NVIDIA Jetson Nano as proposed in Embodiment 2 of the present invention are shown below. (a) represents the relationship between input size and detection speed (FPS); (b) represents the experimental results on the African Wildlife dataset; (c) represents the experimental results on the Amur Tiger Re-identification in the Wild dataset; (d) represents the experimental results on the Northeast Tiger and Leopard National Park dataset; and (e) represents the experimental results on the Visual Object Classes dataset.

[0049] Explanation of the labels in the diagram:

[0050] 1. Solar panel; 2. Binocular infrared trigger camera module; 3. Microphone array opening; 4. Housing. Detailed Implementation

[0051] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.

[0052] The following description, with reference to the accompanying drawings, outlines an intelligent wildlife monitoring device and method based on multimodal fusion and self-organizing networks, in accordance with the present invention.

[0053] Example 1:

[0054] Please see Figure 1 This invention proposes an intelligent wildlife monitoring device based on multimodal fusion and self-organizing networks, comprising:

[0055] Outer shell 4, used to house and protect the internal modules;

[0056] Solar panel 1, fixed to the top of the outer casing 4, is used to convert solar energy into electrical energy to provide auxiliary power and extend the equipment's outdoor endurance.

[0057] The binocular infrared trigger camera module 2 is mounted in parallel on the front panel of the housing 4 and is used to collect stereoscopic visual information and depth perception of wild animals.

[0058] Microphone array opening 3, the array is opened on the front panel of the housing 4; a microphone array module is installed inside the microphone array opening 3, the microphone array module is used to collect ambient sound and acoustic signals emitted by animals, and to locate the sound source.

[0059] The outer casing 4 also includes an active processing unit, a communication and self-organizing network module, and a power management module.

[0060] The main control and processing unit is connected to the binocular infrared trigger camera module and the microphone array module, respectively, to control data acquisition, synchronize visual and auditory signals, and run target detection and recognition algorithms. The main control and processing unit is equipped with the WS-YOLO algorithm, which is used to filter false triggers and detect and classify wild animal species from the acquired image data. The WS-YOLO algorithm is optimized by model quantization, pruning and hardware acceleration on the device side to adapt to the resource limitations of the embedded platform.

[0061] The communication and self-organizing network module, connected to the main control and processing unit, is used to realize wireless networking and data transmission between devices. In areas without public network coverage, it transmits data through multi-hop relay. The communication and self-organizing network module supports multiple wireless communication protocols and automatically selects or switches communication modes according to the network environment. Wireless communication protocols include, but are not limited to, LoRa, ZigBee, and Wi-Fi Mesh.

[0062] The power management module, connected to the solar panel, is used to power the various functional modules. The power management module includes a rechargeable lithium battery and a solar charging controller, and adopts an intelligent power consumption management strategy to extend the device's battery life.

[0063] The aforementioned binocular infrared trigger camera module and microphone array module employ a hardware synchronization mechanism to ensure that visual and auditory data remain time-aligned during acquisition.

[0064] The main control and processing unit is also used to extract depth information from the images acquired by the binocular infrared camera and combine it with the sound source localization results of the microphone array to perform multimodal data fusion in order to obtain richer target information.

[0065] This invention also proposes a multimodal monitoring method for wild animals based on the aforementioned equipment. This method uses the WS-YOLO algorithm as its core, and optimizes the acquisition, processing and transmission of multimodal data by fusing binocular infrared visual information and microphone array auditory information, thereby improving the intelligence level and operational efficiency of wild animal monitoring.

[0066] Specifically, to address the problems of insufficient fusion of visual and auditory information, low hardware integration, limited intelligence, and difficulties in data transmission in areas without network coverage in existing wildlife monitoring equipment, this invention makes the following innovative design:

[0067] ① Integrated binocular infrared trigger camera and microphone array:

[0068] Deep hardware integration of the binocular infrared trigger camera module and microphone array module ensures precise temporal synchronization of visual and auditory signals. Optimized structural design ensures good spatial matching between the optical axis of the binocular camera and the acoustic center of the microphone array, providing a hardware foundation for multimodal data fusion.

[0069] Binocular infrared cameras acquire depth information of targets by calculating parallax, enhancing the ability to detect and locate wildlife targets; microphone arrays assist cameras in target search and triggering through sound source localization technology, improving monitoring efficiency.

[0070] ② WS-YOLO Algorithm Optimization and Deployment:

[0071] An improved YOLOv11n network was introduced as the basic architecture, and the network was optimized to suit the characteristics of wildlife monitoring scenarios. The WS-YOLO algorithm has the functions of false trigger filtering and species detection and classification, which can effectively identify and filter out false triggers caused by non-target factors, while identifying the species of detected animal targets.

[0072] The algorithm incorporates a waveform convolution (WTConv) module, which extracts spatial and frequency domain information through discrete wavelet decomposition, enhancing its ability to extract features from low-contrast, small targets, or partially occluded animals. WTConv replaces some traditional convolutional layers in the YOLOv11n network, significantly improving feature richness and detection accuracy.

[0073] A Spatial and Channel Co-Attention (SCSA) module is introduced to jointly model spatial and channel attention, enabling the network to better focus on the target region and suppress background noise. The SCSA module plays a role in the feature fusion stage, and performs particularly well in complex environments and occluded scenes.

[0074] An adaptive sliding gradient loss function (SlideLoss) is introduced to address the imbalance between foreground targets and background regions in wildlife detection tasks. The sensitivity of the loss function is dynamically adjusted to enhance the model's ability to handle difficult samples and blurred boundaries.

[0075] The WS-YOLO model is optimized through quantization, pruning, and hardware acceleration to enable efficient operation on embedded devices. Model quantization replaces floating-point operations with low-order fixed-point operations, significantly reducing model storage requirements and computational complexity; hardware accelerators (such as NPUs) further enhance the algorithm's efficiency.

[0076] ③ Multimodal data fusion and self-organizing network transmission:

[0077] The main control and processing unit implements timestamp alignment and feature-level fusion of visual and auditory data. By splicing or weighted fusion of sound features and image features, richer target description information is formed, improving the accuracy and information dimensionality of monitoring data.

[0078] The communication and ad-hoc networking module supports multiple wireless communication protocols and can automatically establish an ad-hoc network in remote areas without public network coverage, enabling multi-hop relay transmission between devices. Through ad-hoc networking technology, monitoring data can be reliably transmitted back to the data center or designated nodes, expanding the monitoring range and improving the system's robustness.

[0079] Through the above-mentioned optimized design, the present invention significantly improves the hardware integration, multimodal data fusion, intelligence level and data transmission capability of wildlife monitoring equipment, adapts to the monitoring needs of complex field environments, and improves the overall operating efficiency of the monitoring system.

[0080] In its implementation, this invention relies on an integrated hardware design and an optimized WS-YOLO algorithm to integrate multimodal data acquisition, processing, and transmission into a single device. It achieves collaborative operation between devices through self-organizing network technology, solving the data transmission problem in areas without network coverage. Simultaneously, algorithm optimization and hardware acceleration ensure efficient operation of the device under low power consumption conditions.

[0081] The multimodal monitoring method for wild animals proposed in this invention includes the following steps:

[0082] Signal Acquisition and Synchronization: Visual and auditory data are acquired synchronously through a binocular infrared trigger camera module and a microphone array module, and the data is precisely aligned in time through a hardware synchronization mechanism.

[0083] Multimodal data fusion: In the main control and processing unit, the collected visual and auditory data are time-stamp aligned and feature-level fused to form comprehensive target description information.

[0084] Intelligent analysis and screening: The WS-YOLO algorithm is used to screen for false triggers and classify species in the fused data, marking valid target data.

[0085] Data Management and Transmission: Filtered valid data is stored locally or transmitted to the data center via communication and self-organizing network modules. Self-organizing network technology enables multi-hop relay transmission to ensure reliable data return.

[0086] Low power optimization: When no target is detected, the device enters a low power sleep mode, keeping only the core sensors running to reduce power consumption and extend device battery life.

[0087] This invention, by combining integrated hardware design, WS-YOLO algorithm optimization, and self-organizing network technology, achieves efficient acquisition, processing, and transmission of multimodal wildlife data. It reduces the storage and transmission of invalid data, lowers system power consumption, and improves the intelligence and operational efficiency of the equipment. This invention is suitable for field monitoring scenarios requiring long-term low-power operation, providing an efficient and reliable solution for real-time wildlife monitoring and ecological protection.

[0088] Example 2:

[0089] Based on Embodiment 1, but with some differences, the present invention will be further described below with reference to the accompanying drawings. The specific details are as follows.

[0090] (1) Hardware structure design

[0091] The wildlife multimodal monitoring device of the present invention is mainly composed of the following hardware modules, which work together to achieve efficient and intelligent monitoring functions.

[0092] ① Binocular infrared trigger camera module

[0093] This module is the core for acquiring visual information, containing two side-by-side infrared camera lenses and image sensors, as well as an infrared trigger unit. The infrared trigger unit typically uses a passive infrared sensor (PIR), which triggers the camera to capture images by detecting the temperature difference between the animal's body temperature and the ambient background. The binocular structure can acquire stereoscopic information of the scene; by calculating parallax, the distance and three-dimensional shape of the target can be estimated, providing richer data for subsequent behavior analysis and target recognition. The lenses need to have a certain wide angle to cover a wider area and consider night vision capabilities to ensure clear imaging even in low-light or no-light conditions. The image sensor should be a low-power, high-sensitivity model to meet the needs of long-term field operation.

[0094] ② Microphone array module

[0095] This module is responsible for acquiring sound signals from the environment and consists of multiple microphones arranged in a specific geometric structure (such as linear or circular). The advantage of the microphone array lies in its ability to perform sound source localization and beamforming, thereby suppressing environmental noise and enhancing the target sound source signal. By analyzing the phase and amplitude differences of the sound signals received by different microphones, the direction and approximate distance of the sound source can be estimated. The sound data acquired by this module will be fused with visual data acquired by the binocular camera. For example, when the microphone array detects an animal call from a specific direction, it can guide the binocular camera to adjust in that direction or trigger a recording.

[0096] ③Main control and processing unit

[0097] The main control and processing unit is the brain of the device, typically employing a low-power, high-performance embedded processor (such as the ARM Cortex-A series) or a dedicated AI processing chip. This unit is responsible for controlling the coordinated operation of the binocular infrared camera module and microphone array module, including triggering logic, data acquisition, and signal synchronization. More importantly, it is responsible for running intelligent algorithms such as WS-YOLO to perform localized processing on the acquired image and sound data, such as target detection, species identification, and false trigger filtering. The processing results (such as images with species tags and key sound clips) will be stored or transmitted via the communication module.

[0098] ④ Communication and Ad Hoc Network Module

[0099] The communication module is responsible for data exchange between the device and external networks or other devices, typically using wireless communication technology. Depending on the actual network coverage of the monitoring area, multiple communication methods can be integrated, such as 4G / 5G modules (used when there is public network coverage), LoRa modules (for long-range low-power communication), and Wi-Fi modules (for short-range data transmission or device configuration). Self-organizing networking is one of the core features; the device has a built-in self-organizing network protocol stack, enabling it to automatically form a network with other nearby devices in areas without public network coverage, achieving multi-hop relay transmission of data. This ensures that monitoring data can be effectively collected and transmitted even in remote areas.

[0100] ⑤ Power Management Module

[0101] The power management module is responsible for providing a stable and efficient power supply to the entire device, maximizing its runtime in the field. This module typically includes a rechargeable lithium battery, a solar charge controller (if equipped with solar panels), and a high-efficiency DC-DC voltage conversion circuit. Power management strategies are crucial; for example, intelligently scheduling the operating states of each module (such as sleep and wake-up) and dynamically adjusting the device's power consumption based on battery level and solar input can enable long-term unattended operation.

[0102] ⑥ Outer shell and protective design

[0103] The equipment casing needs to have good protective performance to cope with the complex and ever-changing environmental conditions in the wild. The casing material should be high-strength, corrosion-resistant, and UV-resistant engineering plastics or metals. The design must consider waterproofing, dustproofing, moisture resistance, and insect and termite protection to ensure the safety of internal electronic components. At the same time, the casing design should also consider the equipment's concealment, using camouflage or colors similar to the environment to avoid disturbing wild animals. The installation method should be flexible, facilitating fixation in different locations such as trees and supports.

[0104] (2) Signal synchronization and data fusion methods

[0105] For effective multimodal monitoring to be achieved, precise synchronization of visual and auditory signals and deep data fusion are essential.

[0106] ① Hardware synchronization mechanism for visual and auditory signals

[0107] At the hardware level, it is necessary to ensure that the acquisition clocks of the binocular infrared camera and the microphone array are synchronized. This can be achieved by sharing the same clock source or using a precise clock synchronization signal (such as a PPS pulse). When the infrared trigger unit detects a target and triggers the camera to capture an image, a synchronization pulse should be sent to the microphone array module simultaneously to initiate or mark the acquisition of a segment of audio data. This hardware-level synchronization mechanism can minimize the time deviation between visual and auditory data, providing an accurate time reference for subsequent fusion analysis.

[0108] ② Timestamp alignment of multimodal data

[0109] Even with hardware synchronization, slight time differences may still exist between data from different modalities when they arrive at the fusion unit due to data processing and transmission delays. Therefore, precise timestamp alignment is required at the software level. Each data packet (such as an image frame or audio clip) should carry high-precision timestamp information. When processing data, the fusion algorithm matches and aligns visual and auditory data belonging to the same time window based on these timestamps, ensuring that the analysis focuses on the performance of the same event in different modalities.

[0110] ③ Data-level and feature-level fusion strategy

[0111] Data fusion can be carried out at different levels.

[0112] Data-level fusion (early fusion) refers to directly splicing or combining raw or pre-processed modal data and inputting it into a subsequent analysis model. For example, the spectrogram of sound can be spliced ​​with image data along the channel dimension to form a multi-channel input.

[0113] Feature-level fusion (intermediate fusion) refers to first extracting high-level features from data of different modalities (e.g., extracting visual features from images and acoustic features from sounds), and then concatenating, weighting, or fusing these feature vectors through more complex fusion networks (such as attention mechanisms).

[0114] Decision-level fusion (late-stage fusion) refers to analyzing different modalities of data separately to draw preliminary conclusions, and then making a comprehensive judgment based on these conclusions. In this invention, depending on the specific application scenario and algorithm requirements, data-level or feature-level fusion strategies can be adopted to achieve the best monitoring results.

[0115] (3) Detailed Explanation of the WS-YOLO Algorithm

[0116] The WS-YOLO proposed in this invention is a lightweight target detection algorithm suitable for deployment on embedded platforms in the field. It is an improvement upon the YOLOv11n framework, introducing three key modules to enhance performance in false trigger filtering and wildlife target detection: Wavelet Convolution (WTConv) module; Spatial and Channel Synergistic Attention (SCSA) module; and Adaptive SlideLoss module. The overall structure of the algorithm is as follows: Figure 2 As shown.

[0117] ① Wavelet Convolution Module (WTConv)

[0118] In natural wildlife monitoring scenarios, animal targets often appear in areas with low contrast, small scale, or partial occlusion. Traditional convolutional neural networks operate only in the spatial domain, which may make it difficult to extract sufficient discriminative features in such cases. To address this issue, this invention introduces wavelet convolution (WTConv) into the backbone and neck networks of WS-YOLO, achieving joint extraction of spatial and frequency domain information through discrete wavelet decomposition. Figure 3 The processing flow of WTConv is demonstrated.

[0119] The WTConv module introduces a two-dimensional discrete wavelet transform (DWT) to transform the input feature map. Decomposed into four sub-bands:

[0120] DWT(X) = {LL,LH,HL,HH}

[0121] Where LL is the low-frequency approximation component, and LH, HL, and HH are the high-frequency components in the horizontal, vertical, and diagonal directions, respectively. All sub-band dimensions are...

[0122] Each subband is processed independently using standard convolutional blocks (convolution, batch normalization, and activation), denoted as:

[0123] F S =BN(σ(Conv) 3×3 (S)),S∈{LL,LH,HL,HH}

[0124] Where σ is an activation function, such as SiLU or ReLU.

[0125] To align with the original spatial dimensions, each feature map is upsampled back to H×W using nearest neighbor interpolation:

[0126]

[0127] Simultaneously, the original input X undergoes spatial domain convolution:

[0128] F std =BN(σ(Conv) 3×3 The output of (X)))WTConv is obtained by concatenating all five processed branches:

[0129]

[0130] To compress the generated feature volume and achieve efficient downstream computation, a 1×1 convolution was applied:

[0131] F out =Conv 1×1 (F WTConv )

[0132] The above design allows the model to benefit from complementary representations: the global context provided by low-frequency signals, and the texture and edge cues provided by high-frequency components. In practice, this significantly enhances the feature richness of shallow and mid-level layers, where small objects and fine-grained environmental textures are most critical.

[0133] In the WS-YOLO architecture, WTConv replaces two traditional convolutional layers in the early stages of the backbone network, where the feature resolution is highest. This strategic integration effectively preserves and enhances object boundaries, contours, and frequency-based discriminators, laying a solid foundation for downstream attention modeling and detection refinement.

[0134] ② Spatial and Channel Synergistic Attention (SCSA) module

[0135] To enhance feature representations in challenging wildlife detection scenarios (e.g., occlusion, small size, low contrast), this invention integrates a Spatial and Channel Synergistic Attention (SCSA) module into WS-YOLO. Unlike traditional attention mechanisms that process spatial and channel dimensions independently, SCSA models spatial and channel attention sequentially in a mutually reinforcing manner. The structure diagram of SCSA is shown below. Figure 4 As shown.

[0136] The SCSA module enhances the network's ability to focus on wildlife target regions by cascading spatial attention and channel attention. The input feature map is:

[0137]

[0138] Where B represents the batch size, C represents the number of channels, and H×W represents the spatial resolution.

[0139] The spatial attention module can capture long-range dependencies along both height and width dimensions. First, average pooling is applied to the width and height to obtain the spatial context descriptor:

[0140]

[0141] Each descriptor is divided into four groups along the channel dimension:

[0142]

[0143] Each group is processed using one-dimensional convolutions of different depths, with kernel sizes k∈{3,5,7,9}, followed by group normalization and activation by a sigmoid function.

[0144]

[0145] Attention Map and Reshaped to adjust the original feature map:

[0146] X'=X⊙A h ⊙A w

[0147] To capture global inter-channel dependencies, we first obtain the downsampled image of X' through spatial pooling:

[0148]

[0149] The downsampled features are normalized and used to compute self-attention along the channel dimension. Queries, keys, and values ​​are generated through grouped 1×1 convolutions:

[0150] Q = Conv 1×1 (Y)

[0151] K = Conv 1×1 (Y)

[0152] V = Conv 1×1 (Y)

[0153] After reshaping through multi-head attention, the scaled dot product attention is calculated:

[0154]

[0155] The features of interest are pooled in the spatial dimension to generate a channel attention map:

[0156]

[0157] Finally, the output of the SCSA module is obtained through element-wise multiplication and residual join:

[0158] X”=X'⊙M C

[0159] Y out =X”+X

[0160] The SCSA module is integrated into the neck of the network, particularly in the feature fusion stage. By jointly modeling spatial and channel attention in a cascaded manner, SCSA enables the network to better focus on discriminative target regions while suppressing irrelevant background noise. This is especially beneficial in complex environments with occluded or camouflaged wildlife targets.

[0161] ③ Adaptive sliding gradient loss function (SlideLoss)

[0162] SlideLoss is used in the classification and confidence prediction branches to improve the model's ability to handle ambiguous and boundary samples. Standard loss functions such as Binary Cross Entropy (BCE) or Focus Loss treat each sample independently and assigns fixed weights, which is ineffective in long-tailed or partially observable cases. Although Focus Loss aims to alleviate class imbalance, it still treats all hard samples uniformly without considering their spatial or semantic context. Inspired by boundary-based learning and dynamic modulation principles, SlideLoss introduces a sliding boundary mechanism, enabling the model to dynamically focus based on the flexible sensitivity of uncertain samples.

[0163] Let p∈[0,1] be the predicted target score of the sample, and y∈{0,1} be the true label. The base loss is derived from the standard binary cross-entropy loss:

[0164] L BCE (p,y)=-[ylog(p)+(1-y)log(1-p)]

[0165] Introduce a sliding boundary δ and a modulation function Φ(p,y,δ) to scale the gradient according to sample uncertainty:

[0166]

[0167] Here, γ>0 is a focusing parameter (similar to focus loss), and δ∈(0,0.5) is an adaptive sliding boundary.

[0168] The complete SlideLoss expression is:

[0169] L Slide (p,y)=Φ(p,y,δ)·L BCE (p,y)

[0170] This formulaic setting allows the loss to emphasize: false negatives: when y = 1 but p is much lower than 1; false positives: when y = 0 but p is still too high; and to reduce the weight of simple examples when the predicted value is within an acceptable confidence interval.

[0171] Unlike fixed-margin methods, SlideLoss dynamically adjusts δ based on batch-level uncertainty statistics during training. Specifically:

[0172] δ=α·StdDev(p),α∈[0.1,0.5]

[0173] Here, StdDev(p) represents the standard deviation of the prediction confidence scores for all positive and negative anchors in a mini-batch. This adaptive mechanism allows SlideLoss to apply a larger margin in the early stages of training—when the prediction variance is high—and gradually reduce the margin as the model converges, thus focusing on finer error correction.

[0174] In WS-YOLO, SlideLoss is applied to both target prediction and classification branches. For bounding box regression, CIoU loss is preserved due to its spatial interpretability. SlideLoss, as a plug-in module, replaces the standard BCE / Focal loss in both classification and confidence outputs without requiring architectural modifications. Its ability to adaptively emphasize difficult samples significantly improves the network's robustness in detecting partially visible, small, or camouflaged animals under complex environmental conditions.

[0175] (4) Ad hoc networks and data transmission

[0176] Self-organizing network technology is key to enabling large-scale wildlife monitoring in areas without network coverage.

[0177] ① Application of self-organizing network technology

[0178] Based on ZigBee, LoRa, Wi-Fi Mesh, or other proprietary protocols, automatic discovery, route establishment, and data relay between monitoring devices are achieved. The choice of ad hoc networking technology depends on the specific needs of the monitoring area, such as transmission distance, data rate, power consumption, and network scale. The core idea is to integrate these existing ad hoc networking capabilities into the multimodal monitoring device of this invention, enabling it to work collaboratively in a networked manner.

[0179] ② Data transfer mechanism in areas without network coverage

[0180] In remote areas lacking public network (such as 4G / 5G) coverage, the data collected by monitoring equipment (valid data after WS-YOLO processing) will be transmitted via a self-organizing network. The basic mechanism is as follows:

[0181] Neighbor discovery and route establishment: After the device starts up, it will automatically search for other monitoring devices in the vicinity and establish multi-hop paths to the data sink node or gateway device through ad hoc network routing protocols (such as AODV, OLSR, or simpler flooding / flooding methods). The sink node is usually located in a location with public network coverage or where it is easy to retrieve data manually.

[0182] Data relay: When a device collects data, if it cannot directly connect to the aggregation point, it will send the data packets to its neighboring nodes that are closer to the aggregation point or have better link quality. The neighboring node then continues to forward the data in the same way until the data reaches the aggregation point.

[0183] Network maintenance and repair: Ad hoc network protocols need to be able to handle dynamic changes in network topology (such as equipment failure, power depletion, addition of new devices, etc.), automatically repair broken routes, and ensure the reliability of data transmission.

[0184] This jump transmission mechanism allows the monitoring network to be flexibly expanded to cover vast areas without network coverage.

[0185] (5) Workflow

[0186] The workflow of the wildlife multimodal monitoring device of the present invention can be summarized in the following steps, such as: Figure 5 As shown:

[0187] Triggering and Data Acquisition: After the infrared triggering unit detects the target, it synchronously triggers the binocular infrared camera to capture an image sequence and simultaneously activates the microphone array to acquire sound signals.

[0188] Signal synchronization and preprocessing: The acquired visual and auditory data are accurately timestamped and undergo preliminary preprocessing (such as noise reduction and format conversion).

[0189] WS-YOLO Intelligent Analysis: The preprocessed data is fed into the WS-YOLO algorithm mounted on the device. The algorithm first performs false trigger filtering to exclude non-target events. If an animal target is confirmed, further species detection and classification are performed.

[0190] Data fusion (optional): Depending on the needs, visual analysis results (such as animal location and species) can be fused with acoustic information (such as call type and sound source direction) to form a richer description of the monitoring event.

[0191] Data storage and transmission: Valid data that has undergone intelligent analysis and filtering (such as images of the target animal, key sound clips, species information, time and location metadata, etc.) is stored in local storage. Simultaneously, the device transmits data to an aggregation point or data center via a self-organizing network module. In areas without a public network, data is transmitted via multi-hop relays.

[0192] Power Management and Status Monitoring: The power management module continuously monitors battery level and device status, and adjusts the operating mode of each module according to preset strategies to optimize power consumption. The device periodically reports its own status (such as battery level, signal strength, and whether it is working properly) to the monitoring center.

[0193] Through the above process, automated, intelligent, and networked monitoring of wildlife activities has been achieved.

[0194] Example 3:

[0195] Based on Embodiments 1-2, but with some differences, the present invention will now be described in conjunction with specific examples of the intelligent wildlife monitoring device and method based on multimodal fusion and self-organizing network proposed in this invention. The specific content is as follows.

[0196] In a typical implementation, the wildlife multimodal monitoring device of this invention was deployed in a large nature reserve to monitor the activity patterns and population size of rare species. The reserve has complex terrain, dense vegetation, and most areas lack public network signal coverage.

[0197] Equipment Deployment: Researchers deployed dozens of monitoring devices in key areas (such as near water sources, animal trails, and dens) based on animal activity traces and habitat characteristics. The devices were fixed to tree trunks with brackets, and their height and angle were carefully adjusted to ensure optimal monitoring visibility and acoustic coverage.

[0198] Parameter configuration: The parameters of each device can be set through near-field communication (such as Bluetooth) or configuration commands within the self-organizing network, including infrared trigger sensitivity, shooting mode (photo / video), sound acquisition threshold, WS-YOLO model selection (optimized for target species), data upload interval, etc.

[0199] Data Acquisition and Processing: When a target animal enters the monitoring area, the infrared-triggered camera activates, capturing high-definition images and videos. Simultaneously, a microphone array records the animal's calls or activity sounds. The WS-YOLO algorithm on the device analyzes the image data in real time, filtering out false triggers such as rustling grass, and identifying the animal species. For example, the system successfully identified the flagship species A in the area and recorded its activity time, location, number of individuals, and related behavioral fragments (such as foraging and calls).

[0200] Ad-hoc network data transmission: Since there is no public network in the deployment area, the collected valid data (such as images, sound clips, identification results, and metadata containing species A) are transmitted via ad-hoc network links between devices through multiple hops. The data eventually converges to a gateway device near the reserve management station, which then uploads the data to the cloud platform via satellite link or wired network.

[0201] Data Analysis and Application: Researchers remotely accessed and analyzed monitoring data via a cloud platform. The data, incorporating visual and auditory information, provided valuable insights into species A's diurnal activity rhythms, home range, population structure, and reproductive behavior. Simultaneously, the system also detected signs of poaching or illegal intrusion, providing timely early warnings for the patrol and management of the protected area.

[0202] This case fully demonstrates the practicality, reliability, and efficiency of the invention in complex field environments, providing strong technical support for wildlife research and conservation.

[0203] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0204] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.

[0205] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.

[0206] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0207] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0208] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.

Claims

1. A wildlife intelligent monitoring device based on multimodal fusion and self-organizing networks, characterized in that, include: The outer casing (4) is used to house and protect the internal modules; A solar panel (1) is fixed to the top of the outer casing (4) to convert solar energy into electrical energy and provide auxiliary energy to extend the equipment's outdoor endurance. A binocular infrared trigger camera module (2) is mounted in parallel on the front panel of the housing (4) for collecting stereoscopic visual information and depth perception of wild animals; Microphone array opening (3), the array is opened on the front panel of the outer shell (4); a microphone array module is installed inside the microphone array opening (3), the microphone array module is used to collect ambient sound and acoustic signals emitted by animals, and to locate the sound source; The outer shell (4) is also equipped with an active processing unit, a communication and self-organizing network module, and a power management module. The main control and processing unit is connected to the binocular infrared trigger camera module and the microphone array module, respectively, for controlling data acquisition, synchronizing visual and auditory signals, and running target detection and recognition algorithms. The main control and processing unit is equipped with the WS-YOLO algorithm, which undergoes model quantization, pruning, and hardware acceleration optimization on the device side to adapt to the resource limitations of the embedded platform. The WS-YOLO algorithm is an improvement based on the YOLOv11n framework, introducing a wavelet convolution module, a spatial and channel collaborative attention module, and an adaptive sliding gradient loss function to improve the performance of false trigger filtering and wildlife target detection, specifically including: Wavelet convolution is introduced into the backbone and neck network of WS-YOLO to achieve joint extraction of spatial and frequency domain information through discrete wavelet decomposition. By integrating the spatial and channel collaborative attention module into the neck of the network, spatial and channel attention are modeled jointly in a cascaded manner, enabling the network to focus on discriminative target regions and suppress irrelevant background noise; The adaptive sliding gradient loss function is applied to both the target prediction and classification branches as a plug-in module to replace the standard BCE / Focal loss in the classification and confidence outputs. The communication and self-organizing network module is connected to the main control and processing unit to realize wireless networking and data transmission between devices, and transmits data through multi-hop relay in areas without public network coverage. The power management module, connected to the solar panel, is used to power the various functional modules.

2. The intelligent wildlife monitoring device based on multimodal fusion and self-organizing network according to claim 1, characterized in that, The binocular infrared trigger camera module and the microphone array module adopt a hardware synchronization mechanism to ensure that visual and auditory data are time-aligned during acquisition.

3. The intelligent wildlife monitoring device based on multimodal fusion and self-organizing network according to claim 1, characterized in that, The communication and self-organizing network module supports multiple wireless communication protocols and automatically selects or switches communication modes according to the network environment; the wireless communication protocols include LoRa, ZigBee, and Wi-Fi Mesh.

4. The intelligent wildlife monitoring device based on multimodal fusion and self-organizing network according to claim 1, characterized in that, The power management module includes a rechargeable lithium battery and a solar charging controller, and employs an intelligent power consumption management strategy to extend the device's battery life.

5. The intelligent wildlife monitoring device based on multimodal fusion and self-organizing network according to claim 1, characterized in that, The main control and processing unit is also used to extract depth information from the images acquired by the binocular infrared camera, and combine the sound source localization results of the microphone array to perform multimodal data fusion in order to obtain richer target information.

6. A wildlife intelligent monitoring system constructed based on the device described in any one of claims 1-5, characterized in that, The system connects the binocular infrared trigger camera module, microphone array module, main control and processing unit, communication and self-organizing network module, and power management module through a self-organizing network to collaboratively complete the task of monitoring wild animals in a large area without network coverage.

7. A method for intelligent monitoring of wild animals implemented by the device according to any one of claims 1-5, characterized in that, Includes the following: Data acquisition is performed by synchronously activating a binocular infrared camera and microphone array via infrared triggering. The collected visual and auditory data are timestamped and aligned. On the device side, the WS-YOLO algorithm is used to perform false trigger filtering and species identification on image data; Effective monitoring data can be stored locally or transmitted to the data center via an ad hoc network.

8. The intelligent wildlife monitoring method based on multimodal fusion and self-organizing networks according to claim 7, characterized in that, It also includes fusing visual analysis results with acoustic information to form a comprehensive monitoring event report.

Citation Information

Patent Citations

  • Wild animal feature recognition supervision system based on multiple sources and multiple modes

    CN120337141A

  • Artificial intelligence system utilizing microphone array and fisheye camera

    US20190236416A1