Wild animal feature recognition supervision system based on multiple sources and multiple modes
Through a multi-source and multi-modal wildlife feature recognition supervision system, combining video, infrared, microphone and manual observation data, the deep learning framework is used to integrate cross-modal features, solving the problem of low accuracy of single-modal recognition, and achieving unified processing and efficient identification of different data.
Patent Information
- Application Number
- CN202510434742.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
When a single image or sound data is recognized and regulated in the prior art, the recognition accuracy is low and data of different scales and formats under different modes cannot be processed, and the applicability is poor.
A multi-source and multi-modal wild animal feature recognition supervision system is adopted, including a multi-source data acquisition module, a multi-modal processing module, a dual-channel feature extraction module, a cross-modal integration module and an analysis supervision module. Through video cameras, infrared monitoring cameras, microphone arrays and manual observation data acquisition, data processing and feature integration are combined with convolutional neural networks and Bayesian deep learning frameworks, data processing and feature integration are achieved to achieve cross-modal semantic association and feature recognition.
It improves the accuracy and robustness of wildlife recognition, can process data of different resolutions, sizes, distances and formats, and enhances the applicability and recognition effect of the system.
Smart Images

Figure CN120337141A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of wild animal feature recognition and supervision, and specifically relates to a wild animal feature recognition and supervision system based on multi-source and multi-modal data. Background Art
[0002] With the increasing deterioration of the global ecological environment and the continuous reduction of wild animal resources, the protection of wild animals has become a hot issue of global concern. In order to effectively protect wild animals, it is necessary to comprehensively and accurately monitor their population numbers, distribution ranges, activity habits, etc. Due to the complex and changeable activity environments of wild animals, and they are often in remote areas, it brings great challenges to the monitoring work.
[0003] With the development of technology, automated monitoring devices such as video cameras and infrared monitoring cameras have gradually been applied to the field of wild animal monitoring. These devices can continuously and automatically record the activities of wild animals, and combined with the wild animal feature recognition and supervision system, realize the identification and monitoring of wild animals.
[0004] However, most of the current wild animal feature recognition and supervision systems mainly identify and supervise based on single image or sound data. If in the case of long distance or low light conditions, the clarity of the image data may be reduced. At the same time, the single wild animal call data may be interfered by environmental noise, affecting the recognition effect, resulting in a low overall recognition accuracy, and being unable to process different scales and formats of data under different modalities, requiring manual adjustment of parameters, with poor applicability. Summary of the Invention
[0005] This application provides a wild animal feature recognition and supervision system based on multi-source and multi-modal data, aiming to solve the problems in the prior art that the recognition and supervision based on single image or sound data lead to a decline in recognition accuracy, and at the same time being unable to process different scales and formats of data under different modalities, with poor applicability.
[0006] A wild animal feature recognition and supervision system based on multi-source and multi-modal data includes a multi-source data acquisition module, a multi-modal processing module, a dual-channel feature extraction module, a cross-modal integration module, and an analysis and supervision module;
[0007] The multi-source data acquisition module includes a data source acquisition unit and a preprocessing unit. The multi-source data acquisition module is the basic part of the whole system, responsible for collecting relevant data of wild animals from different sources and different types of data sources;
[0008] The multi-modal processing module is used to uniformly process data with different resolutions, different sizes, distances, and different formats;
[0009] The dual-channel feature extraction module is used to process data of different modalities, facilitating subsequent feature integration;
[0010] The dual-channel feature extraction module includes an image branch module and a sound branch module. The image branch module is constructed by building an FPN network with a convolutional neural network;
[0011] The cross-modal integration module is used to achieve deep integration of dual-channel features and establish cross-modal semantic associations;
[0012] The analysis and supervision module is used to train and optimize the comparison model with a large amount of labeled data.
[0013] Further, the data acquisition unit specifically includes video camera monitoring data, infrared monitoring camera data, microphone array recording data, and manually observed shooting data.
[0014] Further, the preprocessing unit includes data cleaning and data denoising.
[0015] Further, the multi-modal processing module includes a format conversion unit and an aggregation unit. The format conversion unit is used to convert the preprocessed images into a unified frame sequence image format and uniformly adjust the input image size to the range of 32×32 pixels to 1024×1024 pixels to adapt to animal target detection with different resolutions.
[0016] Further, the format conversion unit can also convert sound data into a unified spectrogram format, with the sampling rate covering 0 - 100 kHz, and convert the data into a high-quality data set to provide a basis for subsequent feature extraction.
[0017] Further, the aggregation unit is the core part of the multi-modal processing module and is responsible for receiving the data set from the format conversion unit.
[0018] Further, the sound branch module is mainly responsible for processing the collected sound data to extract the acoustic features of wild animals, and uses audio analysis technology to extract the signal information that can characterize the key sound frames of wild animals from the preprocessed sound data.
[0019] Further, the cross-modal integration module includes a deep learning model unit and an integration unit. The deep learning model constructs a three-dimensional environmental map based on the Bayesian deep learning framework and SLAM, projects the key sound frame coordinates and image position coordinates at the same time axis in the aggregation unit into the same coordinate system, and calculates the integration correlation parameter M according to the key sound frame coordinates and image position coordinates ij ;
[0020] The integration unit can identify key features from complex multi-modal sound data and image data.
[0021] Furthermore, the analysis and supervision module includes a database unit, a comparison model unit, and a supervision unit. The database unit is the data center of the entire system, responsible for classifying and storing a large amount of wildlife-related data according to the species, time, and geographical location of wild animals.
[0022] The comparison model unit is constructed based on a deep learning model and has the ability to identify features.
[0023] The supervision unit can detect changes in the number of wild animals and changes in the activity range according to the comparison model unit. Compared with the prior art, the present application has at least the following beneficial effects:
[0024] Based on further analysis and research of the problems in the prior art, the present application obtains video camera monitoring data, infrared monitoring camera data, microphone array recording data, and manual observation and shooting data to achieve the purpose of multi-source data collection, making the data sources complement each other, providing comprehensive wildlife activity information, and facilitating the improvement of subsequent recognition effects. At the same time, combined with the multi-modal processing module, it can uniformly process data with different resolutions, sizes, distances, and formats, improve applicability, facilitate subsequent feature extraction and recognition, and improve the accuracy and robustness of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic flow chart of a wild animal feature recognition and supervision system based on multi-source and multi-modal provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0027] A wild animal feature recognition and supervision system based on multi-source and multi-modal provided by the present application includes a multi-source data collection module, a multi-modal processing module, a dual-channel feature extraction module, a cross-modal integration module, and an analysis and supervision module.
[0028] The multi-source data collection module includes a data source acquisition unit and a preprocessing unit. The multi-source data collection module is the basic part of the entire system, responsible for collecting wildlife-related data from different sources and different types of data sources. The data sources include videos, photos, and sounds.
[0029] Among them, the data acquisition unit specifically includes video camera monitoring data, infrared monitoring camera data, microphone array recording data, and manual observation and shooting data, and specifically includes the following contents:
[0030] a) Video camera monitoring data
[0031] Deploy high-definition visible light video cameras in wildlife activity areas to continuously record the activities of wild animals. The video cameras support day-night mode switching, with color imaging during the day and supplementary light imaging at night. They are used to capture the visual characteristics of the body surface texture, markings, and limb morphology of wild animals.
[0032] b) Infrared monitoring camera data
[0033] The infrared monitoring camera is equipped with a thermal imaging device with a non-cooled infrared detector, which can capture the animal body temperature radiation characteristics in a completely dark, thick fog, or jungle occlusion environment, and identify hidden animals through the heat source distribution, playing a role in complementing the target detection blind area of the visible light video camera, facilitating the continuous recording of the nocturnal activity patterns of wild animals.
[0034] c) Microphone array recording data
[0035] It can simultaneously receive the sound waves of wild animal calls from different directions. By analyzing the characteristics such as frequency, duration, pitch, and rhythm of the calls, it can reflect information such as the species, gender, age, and behavior status of wild animals, providing strong data support for the research and protection of wild animals.
[0036] d) Manual observation and shooting data
[0037] Professional personnel conduct field observations and record the characteristics and activities of wild animals on-site. They mainly record occasional events in areas not covered by the equipment, which can provide a more comprehensive understanding of the biological characteristics of wild animals (such as claw print size, fecal morphology). Manual observation data can be used as labeled data for model training, facilitating the supplementation of the deficiencies of automated monitoring equipment and providing more abundant wild animal information.
[0038] The preprocessing unit specifically includes the following:
[0039] a) Data cleaning
[0040] Process blurred images: The image data captured by video cameras and infrared monitoring cameras may become blurred due to poor shooting conditions. Blurred images will reduce the accuracy of image recognition. Therefore, it is necessary to improve the image quality through image enhancement and sharpening techniques, and directly mark the blurred images as invalid data and eliminate them.
[0041] b) Data denoising
[0042] Filtering process: For sound data, low-pass filtering and high-pass filtering methods are used to remove noise within a specific frequency range to avoid the noise affecting the sound recognition effect.
[0043] Noise reduction algorithm: For image data, the image is smoothed by means of mean filtering to reduce noise points. Mean filtering is a spatial domain filtering method mainly used to remove noise points in the image while trying to retain the detailed information of the image. Its basic principle is to replace the value of the current pixel with the average value of the neighboring pixels. In this way, the image is smoothed, and the difference between the noise points in the image and the surrounding pixels is reduced after processing, so as to achieve the purpose of noise reduction.
[0044] The multi-modal processing module includes a format conversion unit and an aggregation unit, which are used to uniformly process data with different resolutions, sizes, distances, and formats, facilitating subsequent feature extraction and recognition. The specific contents are as follows:
[0045] a) Format conversion unit
[0046] Convert the preprocessed image into a unified frame sequence image format, and uniformly adjust the input image size to the range of 32×32 pixels to 1024×1024 pixels to adapt to animal target detection with different resolutions. Convert the sound data into a unified spectrogram format, and the sampling rate should cover 0 - 100kHz for subsequent processing and analysis.
[0047] After being processed by the preprocessing module, the data is converted into a high-quality data set, providing a basis for subsequent feature extraction.
[0048] The aggregation unit is the core part of the multi-modal processing module and is responsible for receiving the data set from the format conversion unit. It supports the access of multiple protocols such as RTSP (Real-Time Streaming Protocol), HTTP (Photo Upload), and MQTT (Sensor Metadata). The data source type is automatically identified through a protocol parsing engine. And deploy PTP (Precision Time Protocol) to achieve microsecond-level time synchronization between devices, map the dual-channel data to a unified time axis, and eliminate the clock drift of each sensor. Add accurate timestamps to the image data frames and sound data respectively to synchronize the spatio-temporal alignment and data binding time.
[0049] At the same time, the aggregation unit uses a hierarchical buffer to cope with the data rate difference. The video data is allocated to use a high-priority buffer (capacity ≥ 5 minutes of the original stream) to prevent frame loss due to network fluctuations. The photo and audio data are allocated to use a low-latency memory pool to ensure the instant processing of single-shot capture data (such as the trigger of an infrared camera to take a photo), and automatically clean the expired cache based on the LRU (Least Recently Used) algorithm to avoid storage overflow.
[0050] The dual-channel feature extraction module includes an image branch module and a sound branch module. By respectively processing data of different modalities through the image branch module and the sound branch module, it is convenient for subsequent feature integration, can make full use of multi-source data information, and improve the accuracy and robustness of target detection and recognition. The specific contents are as follows:
[0051] a) Image Branch Module
[0052] The image branch module is constructed by building an FPN network with a convolutional neural network. It can achieve multi-scale feature extraction and integration. The FPN network extracts features from feature maps at different levels to achieve multi-scale object detection. FPN (Feature Pyramid Network), namely the feature pyramid network, is an important network structure in deep learning.
[0053] In a convolutional neural network, as the network depth increases, the size of the image frame gradually decreases, the semantic information becomes richer, but the spatial information is gradually lost; while the low-level image frame has a higher resolution and contains more spatial detail information, but the semantic information is relatively weak. Through the FPN module, the strong semantic information of the high level can be combined with the detail information of the low level, and the hierarchical features in the image frame can be automatically learned, so as to effectively identify the types of wild animals and obtain key image frames.
[0054] b) The sound branch module is mainly responsible for processing the collected sound data to extract the acoustic features of wild animals. These features include but are not limited to frequency, pitch, rhythm, etc., and are another important basis for identifying the types and behaviors of wild animals.
[0055] Use audio analysis technology to extract the signal information that can represent the key sound frames of wild animals from the preprocessed sound data.
[0056] Time-frequency transformation: Use the short-time Fourier transform (STFT) to convert the sound signal into a time-frequency domain representation, and extract the key time-frequency domain in the time-frequency domain sound signal as the key sound frame to improve the pertinence and accuracy of feature extraction.
[0057] The cross-modal integration module includes a deep learning model unit and an integration unit, which are used to achieve the deep integration of dual-channel features. By establishing the cross-modal semantic association, it can effectively solve the misjudgment problem caused by the ambiguity of single-modal data.
[0058] The deep learning model constructs a three-dimensional environment map based on the Bayesian deep learning framework and SLAM, projects the key sound frame coordinates and image position coordinates under the unified time axis in the convergence unit into the same coordinate system, and calculates the integration association parameter M according to the key sound frame coordinates and image position coordinates ij , where the calculation formula is:
[0059]
[0060] Among them, is the spatial coordinate of the key sound frame, is the spatial coordinate of the key image frame, and are the feature vectors of the sound and image positions in this spatial coordinate respectively, and σ is the environmental perception parameter.
[0061] The integration unit can judge the key features from complex multi-modal sound data and image data. The integration unit takes the integration correlation parameter M ij as an independent node and outputs multiple integration correlation parameters M at different time nodes during this period in real time ij . The preset node range value judges multiple integration correlation parameters M ij and the numerical value of the node range value . If the numerical value of a certain integration correlation parameter M ij is infinitely close to 's numerical value, then the key sound frame and key image frame in this modality are judged as key features.
[0062] The analysis and supervision module includes a database unit, a comparison model unit, and a supervision unit. It is used to train and optimize the comparison model through a large amount of labeled data, improve the recognition accuracy and robustness of the model, and conduct real-time monitoring and analysis of the recognition results, and count the species, quantity, and activity range of wild animals.
[0063] The database unit is the data center of the entire system and is responsible for classifying and storing a large amount of wild animal-related data according to the species, time, and geographical location of wild animals. These data come from a wide range of sources, including the original images obtained from field monitoring devices and the image data after accurate recognition. It is convenient for subsequent retrieval and invocation during recognition.
[0064] With the continuous operation of the monitoring devices and the continuous influx of new data, the database unit needs to update the data in a timely manner. When new wild animal images or data are transmitted, the database unit will add them to the corresponding category storage location and update the data index for quick query.
[0065] The comparison model unit is constructed based on a deep learning model and has the ability of feature recognition. The comparison model unit needs to be trained with a large amount of labeled wild animal data in the database unit before use. By continuously adjusting the parameters and structure of the model, the model can accurately identify different species of wild animals and their features. At the same time, during the training process, technical means such as regularization, data augmentation, and attention mechanism are adopted to improve the generalization ability and robustness of the model, reduce the overfitting phenomenon, and facilitate improving the feature comparison effect.
[0066] The key features identified by the integration unit are input into the comparison model unit. The comparison model unit automatically extracts the key sound frames and key image frames from the key features, including the morphology, texture, and voiceprint features of wild animals in this modality. Then, the key features are compared with the database to identify information such as the species, quantity, and activity range of the animals, providing a scientific basis for wild animal protection.
[0067] The supervision unit can detect changes in the quantity of wild animals, changes in the activity range, etc. according to the comparison model unit. Once an abnormal situation is detected, the supervision unit will promptly send an alarm notification to relevant personnel so that timely measures can be taken for intervention and protection.
[0068] In the above wild animal feature recognition and supervision system based on multi-source and multi-modal, the purpose of multi-source data collection is achieved by obtaining video camera monitoring data, infrared monitoring camera data, microphone array recording data, and manual observation and shooting data, enabling the data sources to complement each other, providing comprehensive wild animal activity information, and facilitating the improvement of subsequent recognition effects. At the same time, combined with the multi-modal processing module, it can uniformly process data with different resolutions, different sizes, distances, and different formats, facilitating subsequent feature extraction and recognition, and improving the accuracy and robustness of recognition.
[0069] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
Claims
1. A wild animal feature recognition and supervision system based on multi-source and multi-modal, characterized in that, It includes a multi-source data acquisition module, a multi-modal processing module, a dual-channel feature extraction module, a cross-modal integration module, and an analysis and supervision module; The multi-source data acquisition module includes a data source acquisition unit and a preprocessing unit. The multi-source data acquisition module is the basic part of the whole system and is responsible for collecting relevant data of wild animals from different sources and different types of data sources; The multi-modal processing module is used to uniformly process data with different resolutions, different sizes, distances, and different formats; The dual-channel feature extraction module is used to process data of different modalities to facilitate subsequent feature integration; The dual-channel feature extraction module includes an image branch module and a sound branch module. The image branch module is constructed by building an FPN network with a convolutional neural network; The cross-modal integration module is used to achieve deep integration of dual-channel features and establish cross-modal semantic associations; The analysis and supervision module is used to train and optimize the comparison model through a large number of labeled data.
2. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 1, wherein The data acquisition unit specifically includes video camera monitoring data, infrared monitoring camera data, microphone array recording data, and manual observation and shooting data.
3. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 1, wherein, The preprocessing unit includes data cleaning and data denoising.
4. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 1, characterized in that, The multi-modal processing module includes a format conversion unit and an aggregation unit. The format conversion unit is used to convert the preprocessed images into a unified frame sequence image format and uniformly adjust the input image size to the range of 32×32 pixels to 1024×1024 pixels to adapt to animal target detection with different resolutions.
5. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 4, characterized in that, The format conversion unit can also convert sound data into a unified spectrogram format, with the sampling rate covering 0-100kHz, and convert the data into a high-quality data set to provide a basis for subsequent feature extraction.
6. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 4, characterized in that, The aggregation unit is the core part of the multi-modal processing module and is responsible for receiving the data set from the format conversion unit.
7. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 1, characterized in that, The sound branch module is mainly responsible for processing the collected sound data to extract the acoustic features of wild animals, and uses audio analysis technology to extract the signal information that can represent the key sound frames of wild animals from the preprocessed sound data.
8. A multi-source and multi-modal based wild animal feature recognition and supervision system according to claim 1, characterized in that, The cross-modal integration module includes a deep learning model unit and an integration unit. The deep learning model constructs a three-dimensional environmental map based on the Bayesian deep learning framework and SLAM, projects the key sound frame coordinates and image position coordinates under the same time axis in the convergence unit into the same coordinate system, and calculates the integration correlation parameter M according to the key sound frame coordinates and image position coordinates ij ; The integration unit can judge the key features from complex multi-modal sound data and image data.
9. The wildlife feature recognition and supervision system based on multi-source and multi-modal according to claim 1, characterized in that, The analysis and supervision module includes a database unit, a comparison model unit, and a supervision unit. The database unit is the data center of the whole system and is responsible for classifying and storing a large amount of wild animal-related data according to the species, time, and geographical location of wild animals; The comparison model unit is constructed based on a deep learning model and has the ability of feature recognition; The supervision unit can detect changes in the number of wild animals and changes in the activity range according to the comparison model unit.
Citation Information
Cited By
Wild animal intelligent monitoring device and method based on multi-modal fusion and ad hoc network
CN120876978A
Intelligent wildlife monitoring equipment and methods based on multimodal fusion and self-organizing networks
CN120876978B