Multi-modal bird monitoring method and system based on end-side artificial intelligence
Through the multimodal bird monitoring method of edge-side artificial intelligence, voiceprint signals and image streams are collected and processed in real time, solving the problems of low efficiency of traditional monitoring and high latency of cloud solutions, and realizing high-precision, low-power bird monitoring suitable for complex environments.
Patent Information
- Application Number
- CN202510731846.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional bird monitoring relies on manual observation or a single sensor, which is inefficient and has a high false detection rate. Cloud-based AI solutions have high latency and a high risk of privacy leakage. A single modality is not robust enough in complex scenarios.
A multimodal bird monitoring method based on edge artificial intelligence is adopted. By deploying it on edge devices, bird voiceprint signals and image streams are collected in real time, local preprocessing is performed, and features are extracted using voiceprint recognition and image recognition modules. Feature fusion and classification are performed through a multimodal fusion module, and bird positioning information is obtained by combining machine learning linear regression, supporting low-power communication transmission.
It achieves low-latency, high-precision bird monitoring, suitable for environments without network coverage, increases species identification accuracy to 98.5%, and reduces power consumption to <1W, making it suitable for biodiversity surveys and illegal hunting monitoring.
Smart Images

Figure CN120632774A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ecological monitoring and artificial intelligence technology, and in particular to a multimodal bird monitoring method and system based on end-side artificial intelligence. Background Art
[0002] The existing technology has the following problems: Traditional bird monitoring relies on manual observation or single sensors (such as infrared cameras and recording equipment), which suffer from low efficiency, high false detection rates, and difficulty in providing all-weather coverage. Cloud AI solutions require transmitting large amounts of data to servers, which can lead to high latency, privacy risks, and network dependence. A single modality (image only or sound only) is not robust enough in complex scenarios, such as visual occlusion in dense forests and voiceprint interference from environmental noise. Summary of the Invention
[0003] In response to the above-mentioned shortcomings in the existing technology, the present invention provides a multimodal bird monitoring method and system based on end-side artificial intelligence, which solves the problems of high latency and high false detection rate of traditional monitoring solutions. It is suitable for wild environments without network coverage and has the advantages of low power consumption and high precision.
[0004] To achieve the above-mentioned purpose, the present invention adopts a technical solution: a multimodal bird monitoring method based on device-side artificial intelligence, comprising the following steps: S1: Use edge devices deployed at monitoring nodes to collect bird voiceprint signals and image streams in real time and perform local preprocessing; S2: Using the voiceprint recognition module to extract bird voiceprint features based on the preprocessed bird voiceprint signal; S3: Using the image recognition module to extract refined bird image features based on the preprocessed bird image stream; S4: Use the multimodal fusion module to fuse and classify bird soundprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; S5: The classification results and bird location information are stored locally or transmitted to the central platform via a low-power communication protocol to complete multimodal bird monitoring based on edge-side artificial intelligence.
[0005] Furthermore, the bird voiceprint signal in S1 is collected by a multi-channel microphone array, and the sound source positioning is achieved based on beamforming technology; the image stream is captured by a wide-angle camera.
[0006] Furthermore, the local preprocessing in S1 includes performing bilinear interpolation, Gaussian filtering, normalization and affine transformation on the bird voiceprint signal and image stream.
[0007] Furthermore, the voiceprint recognition module in S2 is a lightweight convolutional neural network.
[0008] Furthermore, the image recognition module in S3 includes four sub-models with the same structure, each of which includes an Embedding network and a PoolFormer network; The PoolFormer network includes a Channel MLP network, a feature normalization network, and a feature cascade network; The output of the PoolFormer network for:
[0009] in, is average pooling, is the first input feature map positions; The fully connected layer in the PoolFormer network is a Channel MLP network, and the formula is:
[0010] in, For Channel MLP network, is the input feature map, and is the 1*1 convolution weight, and is the bias term, is the GELU activation function; The output of the submodule for:
[0011] in, is the submodule of the current stage, For the Embedding network, Output results for the submodule of the previous stage, For the pooling operation, is the normalized network; The output of the image recognition module for: .
[0012] Furthermore, the multimodal fusion module in S4 includes a weighted feature pyramid network, a feature convolutional network, and a Bottleneck network, which are connected by a bidirectional feature fusion structure; The weighted feature pyramid network Layer Output for:
[0013] in, To learn the weights, is a convolutional network, is spatial downsampling, is the input data of the multimodal fusion module, is spatial upsampling; The output of the Bottleneck network for:
[0014] in, and Convolutional network for annotation; The bidirectional feature fusion structure Layer Output for:
[0015] in, It is a bidirectional feature fusion structure. Layer output The core operation of and They are weighted feature pyramid network and The output of the layer; The output of the multimodal fusion module for:
[0016] in, is the connection layer, 、 、 and It is the fusion feature output by the bidirectional feature fusion structure.
[0017] The present invention also adopts a technical solution: a system for a multimodal bird monitoring method based on end-side artificial intelligence, comprising: Data acquisition and processing module: used to collect bird voiceprint signals and image streams and perform local preprocessing; Voiceprint recognition module: used to extract bird voiceprint features based on preprocessed bird voiceprint signals; Image recognition module: used to extract refined bird image features based on the preprocessed bird image stream; Multimodal fusion module: used to fuse and classify bird voiceprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; Data transmission and storage module: used to store classification results and bird location information locally or transmit them to the central platform through low-power communication protocols.
[0018] The beneficial effects of the present invention are: This invention synchronizes voiceprint and image data based on timestamps to solve the problem of temporal and spatial inconsistency of heterogeneous data. At the same time, dual-modal fusion improves the accuracy of species recognition in complex scenarios to 98.5%. This invention uses sound source localization (beamforming technology) to narrow the image detection area and improve computing efficiency. It activates the camera based on environmental trigger conditions (such as detecting motion or a specific frequency band sound pattern), reducing standby power consumption. Device-side computing reduces the average daily power consumption of a single device to less than 1W, and supports solar power supply. This paper uses knowledge distillation and quantization technology to compress the bimodal model to less than 10MB, making it suitable for low-computing power devices; The present invention can be extended to fields such as biodiversity surveys and illegal hunting monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flow chart of a multimodal bird monitoring method based on end-side artificial intelligence in the present invention.
[0020] Figure 2 This is the structure diagram of the image recognition module.
[0021] Figure 3 This is the structure diagram of the multimodal fusion module. DETAILED DESCRIPTION
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0023] Example 1, as Figure 1 As shown in FIG, a multimodal bird monitoring method based on edge-side artificial intelligence includes the following steps: S1: Use edge devices deployed at monitoring nodes to collect bird voiceprint signals and image streams in real time and perform local preprocessing; S2: Using the voiceprint recognition module to extract bird voiceprint features based on the preprocessed bird voiceprint signal; S3: Using the image recognition module to extract refined bird image features based on the preprocessed bird image stream; S4: Use the multimodal fusion module to fuse and classify bird soundprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; S5: The classification results and bird location information are stored locally or transmitted to the central platform via a low-power communication protocol to complete multimodal bird monitoring based on edge-side artificial intelligence.
[0024] The bird voiceprint signal in S1 is collected through a multi-channel microphone array, and the sound source positioning is achieved based on beamforming technology; the image stream is captured by a wide-angle camera.
[0025] In this embodiment, raw audio is collected through a microphone array, converted into a spectrogram through FFT, and Mel-frequency cepstral coefficient (MFCC) features are extracted through a lightweight voiceprint model (such as a SqueezeNet variant) to output species probability and call type (courtship, alarm, etc.).
[0026] The video stream is captured by the camera, and the bird area is located using a lightweight target detection model. The super-resolution algorithm is combined to enhance the detailed features of small targets (such as distant birds).
[0027] The local preprocessing in S1 includes bilinear interpolation, Gaussian filtering, normalization and affine transformation processing on the bird voiceprint signal and image stream to achieve preliminary filtering and processing of the input data and provide structured information for the feature extraction model.
[0028] The voiceprint recognition module in S2 is a lightweight convolutional neural network that supports environmental noise filtering and species classification.
[0029] like Figure 2 As shown in the figure, the image recognition module is based on an improved YOLO model, which implements bird target detection and posture analysis. Because birds are relatively small and their features are more detailed, YOLO incorporates a model that extracts more detailed features. This model consists of four sub-models with the same structure, each of which provides more refined feature extraction results for the next module.
[0030] The image recognition module in S3 includes four sub-models with the same structure, each of which includes an Embedding network and a PoolFormer network; The PoolFormer network includes a Channel MLP network, a feature normalization network, and a feature cascade network; The output of the PoolFormer network for:
[0031] in, is average pooling, is the first input feature map positions; The fully connected layer in the PoolFormer network is a Channel MLP network, and the formula is:
[0032] in, For Channel MLP network, is the input feature map, and is the 1*1 convolution weight, and is the bias term, is the GELU activation function; The output of the submodule for:
[0033] in, is the submodule of the current stage, For the Embedding network, Output results for the submodule of the previous stage, For the pooling operation, is the normalized network; The output of the image recognition module for: .
[0034] like Figure 3 As shown, the multimodal fusion module improves classification confidence through decision-level fusion (such as weighted voting) or feature-level fusion (such as feature vector concatenation). This module primarily utilizes a feature pyramid network, a feature convolutional network, and a bottleneck network, combined through a bidirectional feature fusion architecture. This model classifies the high-dimensional information output from the previous process and assigns it to the corresponding bird species, enabling classification of individual birds. Furthermore, it uses machine learning linear regression methods to determine the relative positioning of each bird in the image.
[0035] The multimodal fusion module in S4 includes a weighted feature pyramid network, a feature convolutional network, and a Bottleneck network, which are connected by a bidirectional feature fusion structure; The weighted feature pyramid network Layer Output for:
[0036] in, To learn the weights, is a convolutional network, is spatial downsampling, is the input data of the multimodal fusion module, is spatial upsampling; The output of the Bottleneck network for:
[0037] in, and Convolutional network for annotation; The bidirectional feature fusion structure Layer output for:
[0038] in, It is a bidirectional feature fusion structure. Layer output The core operation of and They are weighted feature pyramid network and The output of the layer; The output of the multimodal fusion module for:
[0039] in, is the connection layer, 、 、 and It is the fusion feature output by the bidirectional feature fusion structure.
[0040] In this embodiment, the fusion strategy is set as follows: if the voiceprint confidence is greater than 90% and the image confidence is less than 70%, the voiceprint result is used as the main result; if the dual-modal results conflict, the local re-inference mechanism is activated to intercept key frames and audio clips for secondary analysis.
[0041] In one embodiment of the present invention, a national nature reserve in Shaanxi Province, China is used as an application scenario to perform species identification and monitoring of endangered birds, while verifying the performance stability of the method in cold and high humidity environments.
[0042] (1) System deployment architecture 1) Hardware layer Edge computing device: uses Rockchip RK3588 quad-core A76 processor (main frequency 2.4GHz) + 8GB LPDDR4X memory module, integrated NPU unit (computing power 3TOPS), packaged in an IP67-level waterproof chassis.
[0043] Voiceprint collection unit: Deploys an 8-channel digital microphone array (sensitivity -32dB@1kHz, frequency response 20Hz-20kHz), equipped with a waterproof and dustproof mesh cover.
[0044] Image acquisition unit: Equipped with a Sony IMX585 20MP CMOS camera (equivalent focal length 28mm, F1.8 aperture), supports an operating temperature range of -40°C to 70°C, and has infrared night vision function (850nm wavelength, maximum fill light distance 30m).
[0045] Energy module: 30W polycrystalline silicon solar panel + 20000mAh lithium iron phosphate battery pack, supporting MPPT charging control algorithm to ensure stable power supply under short daylight conditions in winter.
[0046] 2) Data Layer Real-time acquisition of voiceprint signals (48kHz sampling rate, 16-bit quantization) and image streams (1080P@15fps), hardware-level synchronous triggering through FPGA, and timestamp accuracy of ±2ms.
[0047] 3) Local preprocessing module Voiceprint preprocessing: Adaptive spectral subtraction (ASR) is used to suppress ambient noise, short-time energy (STE) is used to detect speech segments, continuous audio streams are segmented into 3-second segments, and Mel-frequency cepstral coefficients (MFCC, 20-dimensional features) and statistical features such as spectral centroid and zero-crossing rate are calculated.
[0048] Image preprocessing: Dehazing is implemented based on dark channel prior theory (applicable to the rime weather common in protected areas). Bilateral filtering (BF) is used to remove salt and pepper noise. Histogram equalization (HE) is used to enhance contrast. Finally, the image size is normalized to 640 × 640 pixels.
[0049] 4) Algorithm layer Voiceprint recognition module: The MobileNetV3-Small structure (input size 128×256, number of channels reduced to 0.5 times) is used to extract voiceprint features, and the attention mechanism (SENet) is used to enhance the channel weights of bird song features.
[0050] Knowledge distillation technology was introduced, and the pre-trained VGGish model (16-layer CNN) was used as the teacher network. The distillation temperature was set to 3, and the feature map similarity loss weight was 0.1. The final model parameters were compressed to 2.8MB, and the Top-1 classification accuracy reached 94.7% (based on the CLOTHO bird voiceprint dataset test).
[0051] It supports call type classification (six categories, including courtship, alarm, and territory declaration), and implements multi-label classification through support vector machines (SVM), with an average F1-score of 83.2%.
[0052] Image recognition module: Based on the improved YOLOv8 model, the EfficientPS structure is integrated to extract refined features: The front-end network uses EfficientNet-B0 to extract basic features (output feature map size 32×32×1280) The intermediate feature enhancement module contains 4 PoolFormer units (each unit contains Channel MLP+SpaceMLP structure).
[0053] A dynamic routing mechanism is used to optimize the detection head, and the number of feature extraction layers is automatically increased for small bird targets (area accounting for less than 5%).
[0054] The detection results output bird bounding boxes (coordinate accuracy ±2 pixels), species classification (covering 68 common bird species in the protected area), and posture estimation (5 postures including flying, standing, and foraging), with mAP@0.5 reaching 92.3% and a single-frame inference time of 120ms.
[0055] Multimodal fusion module: A feature pyramid network (FPN) is used to achieve cross-modal alignment of voiceprint features (256 dimensions) and image features (512 dimensions), and the feature weights are dynamically adjusted through the attention-guided feature-level fusion (AGF) strategy.
[0056] A lightweight Transformer structure (2-layer encoder, hidden layer dimension 256) is used for final classification. The model parameter size is controlled at 3.6MB. The species recognition accuracy after fusion reaches 98.5%, which is 3.8% and 6.2% higher than that of single modality (94.7% for voiceprint and 92.3% for image).
[0057] 5) Application layer Local storage module: uses an industrial-grade microSDXC card (128GB, read and write speed 90MB / s), supports cyclic overwrite storage, and can save 30 days of monitoring data (including voiceprint features, image thumbnails and structured metadata).
[0058] Data transmission module: Integrates the LoRaWAN communication protocol (SF10 spreading factor, 433MHz frequency band), with a transmission distance of up to 8.3km under line-of-sight conditions. It transmits key data (species occurrence frequency, activity heat map, etc.) on a daily basis, with a data packet size controlled at 256 bytes and transmission power consumption less than 0.5W.
[0059] (2) Experimental design and data collection Quadrat setting: 5 monitoring quadrats (area 2×2km²) were set up in the protected area, and 3 sets of monitoring equipment were deployed in each quadrangle to form a triangular array (spacing 1.2km).
[0060] Data annotation: Voiceprint data: Ornithology experts at the reserve combined on-site observations to label 1,248 valid voiceprint segments with species, achieving a consistency of 95.6%.
[0061] Image data: 8,362 surveillance images were annotated, including 763 individual birds, with pixel-level accuracy.
[0062] Multimodal association: A total of 2,143 voiceprint-image pairs were collected based on timestamp alignment, with temporal and spatial errors controlled within ±1.3 seconds and ±8.7 meters.
[0063] (3) Experimental results and analysis The species recognition performance is shown in Table 1: Table 1 Species recognition performance
[0064] The system performance indicators are shown in Table 2: Table 2 System performance indicators
[0065] The power consumption management effect is shown in Table 3: Table 3 Power consumption management effect
[0066] The robustness test is shown in Table 4: Table 4 Robustness test
[0067] This example demonstrates the effectiveness of the multimodal bird monitoring method based on on-device AI in real-world environments. High-precision identification: Through the complementary fusion of voiceprints and images, the species identification accuracy reaches 97.5% and the behavior classification accuracy reaches 93.9%, which is significantly better than a single modality solution.
[0068] Low-power operation: The average daily power consumption is only 0.68W, it supports solar power supply, and can still operate stably in low temperature environments in winter.
[0069] Real-time guarantee: The edge computing architecture controls the response time within 1.3 seconds, meeting real-time monitoring requirements.
[0070] Data privacy protection: The local processing mode avoids the risk of external transmission of sensitive ecological data and only transmits structured results to the management center.
[0071] Strong environmental adaptability: The system can maintain high recognition performance in adverse weather conditions such as rain, snow, and dense fog, and is suitable for complex field environments.
[0072] This method has successfully documented certain endangered species and provided protected area management with accurate bird distribution heat maps, effectively supporting biodiversity conservation efforts. Experimental data demonstrates that this approach outperforms traditional monitoring methods in both technical performance and practical application, demonstrating significant innovative value and potential for widespread adoption.
[0073] Example 2, a system for a multimodal bird monitoring method based on device-side artificial intelligence, comprising: Data acquisition and processing module: used to collect bird voiceprint signals and image streams and perform local preprocessing; Voiceprint recognition module: used to extract bird voiceprint features based on preprocessed bird voiceprint signals; Image recognition module: used to extract refined bird image features based on the preprocessed bird image stream; Multimodal fusion module: used to fuse and classify bird voiceprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; Data transmission and storage module: used to store classification results and bird location information locally or transmit them to the central platform through low-power communication protocols.
[0074] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the invention.
Claims
1. A multimodal bird monitoring method based on device-side artificial intelligence, characterized in that: The following steps are involved: S1: Use edge devices deployed at monitoring nodes to collect bird voiceprint signals and image streams in real time and perform local preprocessing; S2: Using the voiceprint recognition module to extract bird voiceprint features based on the preprocessed bird voiceprint signal; S3: Using the image recognition module to extract refined bird image features based on the preprocessed bird image stream; S4: Use the multimodal fusion module to fuse and classify bird soundprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; S5: The classification results and bird location information are stored locally or transmitted to the central platform via a low-power communication protocol to complete multimodal bird monitoring based on edge-side artificial intelligence.
2. The multimodal bird monitoring method based on device-side artificial intelligence according to claim 1 is characterized in that: The bird voiceprint signal in S1 is collected through a multi-channel microphone array, and the sound source positioning is achieved based on beamforming technology; the image stream is captured by a wide-angle camera.
3. The multimodal bird monitoring method based on device-side artificial intelligence according to claim 1 is characterized in that: The local preprocessing in S1 includes bilinear interpolation, Gaussian filtering, normalization and affine transformation processing on the bird voiceprint signal and image stream.
4. The multimodal bird monitoring method based on device-side artificial intelligence according to claim 1 is characterized in that: The voiceprint recognition module in S2 is a lightweight convolutional neural network.
5. The multimodal bird monitoring method based on device-side artificial intelligence according to claim 1 is characterized in that: The image recognition module in S3 includes four sub-models with the same structure, each of which includes an Embedding network and a PoolFormer network; The PoolFormer network includes a Channel MLP network, a feature normalization network, and a feature cascade network; The output of the PoolFormer network for: in, is average pooling, is the first input feature map positions; The fully connected layer in the PoolFormer network is a Channel MLP network, and the formula is: in, For Channel MLP network, is the input feature map, and is the 1*1 convolution weight, and is the bias term, is the GELU activation function; The output of the submodule for: in, is the submodule of the current stage, For the Embedding network, Output results for the submodule of the previous stage, For the pooling operation, is the normalized network; The output of the image recognition module for: 。 6. The multimodal bird monitoring method based on device-side artificial intelligence according to claim 1 is characterized in that: The multimodal fusion module in S4 includes a weighted feature pyramid network, a feature convolutional network, and a Bottleneck network, which are connected by a bidirectional feature fusion structure; The weighted feature pyramid network Layer Output for: in, To learn the weights, is a convolutional network, is spatial downsampling, is the input data of the multimodal fusion module, is spatial upsampling; The output of the Bottleneck network for: in, and Convolutional network for annotation; The bidirectional feature fusion structure Layer Output for: in, It is a bidirectional feature fusion structure. Layer Output The core operation of and They are weighted feature pyramid network and The output of the layer; The output of the multimodal fusion module for: in, is the connection layer, 、 、 and It is the fusion feature output by the bidirectional feature fusion structure.
7. A system for a multimodal bird monitoring method based on device-side artificial intelligence according to any one of claims 1 to 6, characterized in that: include: Data acquisition and processing module: used to collect bird voiceprint signals and image streams and perform local preprocessing; Voiceprint recognition module: used to extract bird voiceprint features based on preprocessed bird voiceprint signals; Image recognition module: used to extract refined bird image features based on the preprocessed bird image stream; Multimodal fusion module: used to fuse and classify bird voiceprint features and bird image features, and use machine learning linear regression methods to obtain bird location information in the image; Data transmission and storage module: used to store classification results and bird location information locally or transmit them to the central platform through low-power communication protocols.
Citation Information
Cited By
Low-cost long-distance mobile electroencephalogram monitoring data transmission implementation method
CN121037716A