A distributed collaborative perception multi-modal bird identification method and system

By using a distributed collaborative sensing system and employing heterogeneous cameras and microphone arrays for multimodal data acquisition and processing, the problem of low accuracy in bird identification under complex environments has been solved, achieving high-precision and environmentally adaptable bird identification.

CN122113002APending Publication Date: 2026-05-29ZHEJIANG UNIHOME TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIHOME TECHNOLOGY CO LTD
Filing Date
2026-04-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing bird identification technologies have low accuracy in complex natural environments, making it difficult to meet the needs of ecological monitoring and biodiversity conservation. Multimodal fusion technologies lack spatiotemporal consistency processing and environmental adaptability.

Method used

A distributed collaborative sensing system is constructed, which uses heterogeneous camera linkage and microphone array for multimodal data acquisition. Combined with a spatiotemporal synchronization and adaptive data processing framework, the system dynamically adjusts the processing strategy to adapt to environmental changes through hierarchical feature fusion and intelligent recognition decision system.

Benefits of technology

It achieves high-precision bird identification in complex environments, improves identification accuracy and environmental adaptability, has autonomous optimization capabilities, and adapts to changes in the ecological environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113002A_ABST
    Figure CN122113002A_ABST
Patent Text Reader

Abstract

The application discloses a kind of distributed collaborative perception multi-modal bird identification method and system, it is related to bird identification technical field;Distributed sensor network is constructed, and the data collected is time-aligned by space-time synchronization protocol, and multi-dimensional perception data is formed, wherein multi-dimensional perception data includes image, audio and environment, and its technical key points are as follows: the multi-modal data system of vision, acoustics and environmental parameters is constructed: through heterogeneous camera linkage architecture, wide-angle monitoring unit continuously scans monitoring area, and bird activity is detected in real time based on interframe difference analysis, triggers long-focus tracking unit to predict target trajectory according to historical position data, dynamically adjusts focal length and angle, to ensure that clear image is obtained;Microphone array is combined with sound source positioning algorithm, accurately separates bird song in complex sound field, provides multi-dimensional data support for identification, and adopts hierarchical feature fusion strategy, and image and audio features are spliced after dimension reduction, to effectively combine multi-dimensional data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bird identification technology, specifically to a distributed collaborative sensing multimodal bird identification method and system. Background Technology

[0002] In the fields of ecological environment monitoring and biodiversity conservation, birds, as important indicator species of ecosystem health, play a crucial role in maintaining ecological balance and assessing environmental changes through accurate monitoring of their population dynamics, distribution range, and behavioral habits. For example, in wetland ecosystems, the migration routes and habitat choices of migratory birds directly reflect the quality of the wetland environment; in forest ecosystems, changes in bird species diversity can intuitively reflect the stability of the forest ecosystem. Therefore, achieving efficient and accurate bird identification is the foundation for carrying out ecological research and conservation work.

[0003] Currently, bird identification technology mainly relies on single-modal data processing and image-based visual recognition technology. This technology classifies birds by extracting visual features such as feather texture and body outline, and is often used in scenarios such as fixed-point monitoring in nature reserves and birdwatching enthusiasts' photography and recording. However, this technology has obvious limitations in complex natural environments. For example, in areas with high forest canopy density, insufficient light can lead to blurry images, making it difficult to extract clear bird features. During migratory bird seasons, when large flocks of birds fly together, mutual occlusion frequently occurs, causing ambiguity in target identification in images. Taking a national nature reserve as an example, the accuracy of image-based identification systems drops significantly on cloudy or rainy days or during low light periods in the early morning, failing to meet the needs of long-term, stable monitoring.

[0004] Audio recognition technology based on bird calls identifies species by analyzing the unique acoustic characteristics of bird calls, such as frequency and rhythm. It is suitable for scenarios such as nighttime monitoring and remote data collection in sparsely populated areas. However, in practical applications, this technology is easily affected by environmental noise. For example, at monitoring points near highways or industrial areas, noise from vehicles and machinery can severely mask bird calls. During the rainy season, continuous rain can also significantly reduce the signal-to-noise ratio of audio signals. Data from a coastal wetland monitoring project shows that strong winds generate a lot of noise that interferes with sound recognition, which greatly affects the false judgment rate of audio-based recognition systems, making it difficult to effectively distinguish bird species with similar calls.

[0005] Although some studies have attempted to improve recognition performance using multimodal fusion techniques, existing solutions mostly remain at the level of simple data splicing, lacking in-depth processing of the spatiotemporal consistency between images, audio, and environmental parameters. For example, some systems simply input synchronously acquired image and audio data directly into the classification model without considering the differences in time delay and spatial positioning of different modal data, resulting in poor feature fusion performance. In addition, existing methods mostly use fixed weights for modal fusion, which cannot dynamically adjust the contribution of each modality according to real-time environmental changes. In complex and ever-changing natural environments, it is difficult to fully leverage the complementary advantages of multimodal data.

[0006] In summary, existing bird identification technologies have significant shortcomings in terms of accuracy and environmental adaptability. There is an urgent need for a bird identification method that can deeply integrate multimodal data and adaptively handle environmental changes to meet the growing technological demands in the fields of ecological monitoring and biodiversity conservation. Summary of the Invention

[0007] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a distributed collaborative sensing multimodal bird identification method and system, which solves the problems mentioned in the background technology.

[0008] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: A distributed collaborative sensing multimodal bird identification method, the specific steps of which are as follows: S1. Distributed collaborative perception: Construct a distributed sensor network that includes visual perception units, acoustic perception units, and environmental perception units, and achieve data alignment through a spatiotemporal synchronization protocol to form multidimensional perception data.

[0009] S101, Heterogeneous Camera Linkage Architecture: Employing a collaborative working mode between a wide-angle monitoring unit and a telephoto tracking unit, the wide-angle monitoring unit performs a panoramic scan of the monitoring area at a preset frequency. By analyzing changes in consecutive frames, it detects bird activity. When activity is detected, the telephoto tracking unit is triggered to take targeted shots. The telephoto tracking unit predicts the bird's future position based on its historical movement trajectory, adjusts parameters such as focal length and angle in advance, and continuously tracks the target during shooting, achieving comprehensive monitoring of bird activity.

[0010] S102, Acoustic Positioning System: Deploys a microphone array to determine the location of the sound source by comparing the time difference of sound arriving at different microphones, and achieves three-dimensional positioning of the sound source by enhancing the target sound signal and suppressing background noise, providing a spatial reference for multimodal fusion.

[0011] S103. Environmental Parameter Acquisition: Integrates multiple types of sensors to collect environmental parameters in real time, constructs environmental feature vectors, and provides background information for subsequent data processing.

[0012] S104. Spatiotemporal synchronization mechanism: Add timestamps to the data of each sensor, unify the time base through time alignment algorithm, ensure the spatiotemporal consistency of multimodal data, and lay the foundation for subsequent fusion processing.

[0013] S2, Adaptive Data Processing Framework: Preprocesses data from each modality, including noise suppression, feature enhancement, and semantic segmentation, and employs a parameter dynamic adjustment mechanism based on model performance feedback.

[0014] S201. Image preprocessing process: First, the image is classified at the pixel level, dividing different regions in the image into bird regions and background regions; then, according to the image noise characteristics, an appropriate filtering method is selected to remove noise; finally, by adjusting parameters such as image contrast and brightness, bird features are highlighted to improve recognition accuracy.

[0015] S202. Audio enhancement processing: First, the audio signal is analyzed to identify and remove the dominant environmental noise component. Then, residual noise is further suppressed, and acoustic parameters that reflect the characteristics of bird species are extracted.

[0016] S203. Dynamic parameter adjustment: Based on the real-time performance feedback mechanism, the preprocessing parameters are dynamically adjusted according to the recognition accuracy to achieve adaptive optimization.

[0017] S3. Hierarchical Feature Fusion Architecture: Design feature-level and decision-level fusion strategies to achieve efficient integration of multi-source information through modal complementarity evaluation and dynamic weight allocation.

[0018] S301, Feature-level fusion: Dimensionality reduction is performed on image and audio features to extract the main feature information. Then, the processed features are concatenated to form a joint feature vector, preserving the complementarity of multimodal information.

[0019] S302, Decision-level Fusion: Construct an integrated system that includes multiple classification methods. Each classification method is trained for different types of features. The contribution of each classification method is adjusted according to data quality through a dynamic weight allocation strategy.

[0020] S303, Modal Complementarity Assessment: By analyzing the relationship between image and audio features, the complementarity of the two modal information is quantified, a modal-environment association model is established, and adaptive weight adjustment is achieved.

[0021] S4. Intelligent Recognition and Decision-Making System: It adopts a learning strategy that can quickly adapt to new species, and combines ecological knowledge graphs and historical data verification mechanisms to construct a triple-verification recognition and decision-making process.

[0022] S401, Dual-track learning strategy: By learning how to quickly adapt to new species, the distribution of the feature space is optimized, making the features of samples of the same type closer in space, and the features of samples of different types further apart.

[0023] S402, Triple Verification Mechanism: By calibrating the confidence level of the recognition results, using ecological knowledge of birds for logical verification, and comparing and analyzing with historical data, the reliability and rationality of the recognition results are ensured.

[0024] S403, Active Learning Optimization: Select the most valuable samples for manual annotation and continuously optimize the recognition performance through incremental learning.

[0025] S5. Dynamic Ecological Model: Establish an ecological niche model based on existing knowledge and real-time data, and dynamically update species distribution predictions based on real-time observation data and environmental parameters.

[0026] S501. Initial Model Construction: An initial niche model is constructed based on a global bird distribution database to provide prior knowledge for new regions.

[0027] S502, Online Learning and Updates: As new observation data accumulates, the distribution prediction results are continuously optimized to achieve dynamic updates.

[0028] S503, Transfer Learning Optimization: To address the problem of insufficient data in new regions, the distribution model of similar ecological regions is used as a reference. By adjusting the model parameters, the model is adapted to the characteristics of the new region.

[0029] S6. System self-optimization mechanism: Design a feedback mechanism and multi-objective optimization method that can sense changes in the environment to achieve adaptive adjustment of system parameters and continuous performance optimization.

[0030] S601, Environmental Perception Feedback: Real-time monitoring of environmental changes, automatic adjustment of system parameters and processing strategies to ensure stable performance under different environmental conditions.

[0031] S602, Multi-objective optimization: Simultaneously optimize multiple objectives such as recognition accuracy, processing speed, and energy consumption, and achieve a balanced improvement in overall system performance by searching for the optimal combination of system parameters.

[0032] S603, Continuous Performance Evaluation: Regularly conduct self-evaluation, trigger model fine-tuning through error rate analysis, and achieve continuous improvement of the system.

[0033] Furthermore, the heterogeneous camera linkage architecture includes: adopting a collaborative working mode of a wide-angle monitoring unit and a telephoto tracking unit. The wide-angle monitoring unit performs a panoramic scan of the monitoring area at a preset frequency and detects bird activity by analyzing changes in continuous frame images. This triggers the telephoto tracking unit to predict the bird's position based on historical movement trajectories and adjust parameters in advance for targeted shooting. At the same time, resource utilization is optimized through multi-target task allocation, dynamic exposure control ensures image clarity, high dynamic range imaging improves image quality, and continuous target tracking maintains stable monitoring of the target.

[0034] Furthermore, the adaptive data processing framework includes: performing semantic segmentation on the image to divide it into bird regions and background regions; employing progressive training optimization, initially training on a general dataset and then finely training on a specific region dataset; accelerating the transfer of complex model knowledge to a lightweight model through knowledge transfer; utilizing context enhancement processing to solve the recognition problem when birds are partially occluded or the background is complex; performing adaptive denoising processing based on the local noise characteristics of the image; and optimizing image enhancement by adjusting parameters such as image contrast and brightness.

[0035] Furthermore, the hierarchical feature fusion architecture includes: performing dimensionality reduction processing on image and audio features and then concatenating them to form a joint feature vector; constructing an integrated system containing multiple classification methods and training it for different types of features; evaluating the quality of each modality data by analyzing feature uncertainty; dynamically adjusting the weights of each classification method based on the evaluation results and environmental conditions; and selecting the most discriminative feature subset from the original features for optimization.

[0036] Furthermore, the intelligent recognition decision-making system includes: optimizing the feature space distribution through a dual-track learning strategy to make the features of similar samples closer and the features of different samples further apart; calibrating and optimizing the confidence level of the recognition result probability; integrating knowledge such as bird ecological habits to construct an ecological knowledge graph; comparing the recognition result with the knowledge graph for logical verification; analyzing historical data of the same region and season for comparative analysis; and selecting the sample with the most uncertain prediction result for active learning sample screening.

[0037] Furthermore, the dynamic ecological model includes: constructing an initial niche model based on a global bird distribution database and ecological environment parameters; performing cluster analysis on different ecological regions; continuously updating the distribution prediction model as new observation data accumulates; migrating the distribution model from similar ecological regions and adaptively adjusting it using local data when data for new regions is insufficient; adjusting model parameters by analyzing the differences in characteristics between the source and target domains; evaluating the reliability of the model based on factors such as data volume, environmental similarity, and prediction consistency; and triggering an update process when the data falls below a threshold.

[0038] Furthermore, the system's self-optimization mechanism includes: real-time acquisition of environmental parameters and analysis of their changing trends through an environmental perception unit; automatic adjustment of image and audio data preprocessing parameters based on changes in environmental parameters; simultaneous optimization of multiple objectives such as recognition accuracy, processing speed, and energy consumption, and searching for the optimal parameter combination; retention of the optimal solution for each generation during the optimization process; dynamic adjustment of parameter search strategies according to the optimization progress; periodic comparison of manually labeled data with recognition results to calculate the error rate index; and optimization of recognition performance through incremental learning when the error rate exceeds a threshold.

[0039] Furthermore, it also includes a data storage step, which includes: using a hierarchical storage strategy to store high-frequency access data on high-speed devices and low-frequency data on high-capacity devices, and automatically migrating cold data according to access frequency and time interval; reducing storage overhead while ensuring reliability through data redundancy storage methods; providing a unified data access interface to support multiple query methods; and caching frequently queried data to reduce access time.

[0040] Furthermore, when the method is applied to a mobile terminal device, it also includes: designing an edge computing architecture to perform data processing locally on the device to reduce transmission latency, reducing computational complexity and the demand for device resources through model lightweighting optimization, optimizing the recognition algorithm to support offline recognition capabilities, and designing an intuitive and easy-to-use user interface to visualize the recognition results and receive user feedback.

[0041] (III) Beneficial Effects This invention provides a distributed collaborative sensing multimodal bird identification method and system, which has the following beneficial effects: This invention achieves significant technological improvements in bird identification accuracy, environmental adaptability, system evolution capability, and data management efficiency through deep multimodal data fusion, dynamic adaptive mechanisms, and systematic optimization strategies. Specific beneficial effects are as follows: 1. A multimodal data system integrating visual, acoustic, and environmental parameters is constructed. Through a heterogeneous camera linkage architecture, a wide-angle monitoring unit continuously scans the monitoring area and detects bird activity in real time based on inter-frame difference analysis. This triggers a telephoto tracking unit to predict the target trajectory based on historical location data and dynamically adjust the focal length and angle to ensure clear image acquisition. A microphone array combined with a sound source localization algorithm accurately separates bird calls in complex sound fields, providing multidimensional data support for identification. In the data processing stage, a hierarchical feature fusion strategy is adopted to stitch together image and audio features after dimensionality reduction. Dynamic weight allocation is used to optimize the contribution of classification algorithms such as support vector machines and random forests. A triple verification mechanism is introduced for identification decisions: probability calibration technology improves the confidence of the results, an ecological knowledge base verifies the rationality of species distribution, and historical data comparison evaluates seasonal adaptability. With the synergistic effect of multiple technologies, the identification ambiguity problem in low-light and high-noise scenarios is effectively solved, achieving high-precision classification of bird species.

[0042] 2. Robustness is improved through environmental perception and parameter self-optimization. In the image preprocessing stage, filtering algorithms are automatically matched according to noise characteristics, and contrast is enhanced through adaptive histogram equalization. Audio processing uses a multi-level noise reduction strategy to suppress environmental noise while preserving key acoustic features. The system establishes a mapping relationship between environmental conditions and modal weights to achieve intelligent decision-making. Audio modal weights are automatically increased in low-light environments, and image feature analysis is enhanced in noisy scenes. At the same time, environmental parameters such as temperature and humidity are monitored in real time. When parameter fluctuations exceed the threshold, the preprocessing algorithm parameters are automatically adjusted to ensure that the data processing strategy is adapted to the current environment, enabling the system to operate stably under different climate and geographical conditions.

[0043] 3. A dynamic ecological model is constructed, establishing an initial prediction model based on global bird distribution data and continuously updating it iteratively with new observational data. In new areas with scarce data, transfer learning techniques are employed, referencing historical models from similar ecological areas and fine-tuning them with local data to shorten the model deployment cycle. Intelligent optimization algorithms are used to simultaneously balance recognition accuracy, processing speed, and energy consumption. Manually labeled data is periodically compared, and incremental learning is triggered when the error rate exceeds a threshold to correct model biases. This mechanism enables the system to proactively adapt to changes in the ecological environment and maintain long-term high-performance operation. Attached Figure Description

[0044] Figure 1 This is a flowchart of the identification method in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the data processing in an embodiment of the present invention. Figure 3 This is a flowchart illustrating the overall process architecture of the identification method in this embodiment of the invention. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] As key indicator species of ecosystems, birds are of great significance for ecological and environmental research and biodiversity conservation through accurate monitoring of their population dynamics, distribution patterns and behavioral habits. Traditional single-modal bird identification technology has significant limitations, and existing multimodal identification technology cannot meet the requirements of high precision and high reliability in practical applications. In the field of ecological environment monitoring, image-based bird identification technology has poor robustness in complex environments, while audio-based identification technology is susceptible to noise interference. Existing multimodal identification technologies mostly involve simple data splicing and lack environmental adaptability. Therefore, there is an urgent need for a bird identification method that deeply integrates multimodal data and has environmental adaptability.

[0047] Research Background: (1) Development of image-based bird recognition technology: Early image-based bird recognition relied on manually designed feature extraction methods, such as Scale Invariant Feature Transform (SIFT) and Histogram of Oriented Gradients (HOG). SIFT generates a 128-dimensional feature vector by detecting extreme points in the image and calculating the scale and orientation information of key points. Under ideal conditions, it can achieve a stable recognition accuracy, but it is poorly adaptable to scenes with changes in lighting, perspective, and complex backgrounds. HOG describes the contour features of the target by statistically analyzing the gradient orientation histogram of local regions of the image. It has some effect in simple backgrounds, but its feature discrimination ability is insufficient when faced with birds with varied postures and similar feather textures.

[0048] The development of deep learning has driven innovation in bird image recognition technology. Methods based on convolutional neural networks (CNNs), such as ResNet and VGG, automatically extract image features through multi-layer convolution and pooling operations. On publicly available bird image datasets, they can improve the recognition accuracy to about 85%. However, these models still have limitations in practical applications: for example, image noise increases and texture details become blurred in low-light environments, resulting in incomplete feature extraction; in highly occluded scenes, parts of the bird's body are obscured by branches and leaves, making it difficult for the model to obtain complete morphological features, thus reducing the recognition accuracy.

[0049] (2) Development of audio-based bird identification technology: Traditional audio-based bird identification primarily involves extracting Mel-frequency cepstral coefficients (MFCCs) and classifying them using support vector machines (SVMs) or hidden Markov models (HMMs). MFCCs, by mimicking the characteristics of the human auditory system, convert audio signals into cepstral coefficients, effectively characterizing bird calls in quiet environments. When combined with an SVM classifier, they can achieve an 80% recognition accuracy. However, in natural environments, interference signals such as wind, rain, and human activity noise overlap with bird calls, leading to biases in MFCC feature extraction. For example, the low-frequency components of wind may mask key frequencies in bird calls, causing misclassification and resulting in lower recognition accuracy in noisy environments.

[0050] (3) Current status of distributed collaborative sensing multimodal bird recognition technology: Existing distributed collaborative sensing multimodal bird identification methods attempt to fuse image and audio data to improve performance, but they generally suffer from the following problems: First, the data fusion method is simple, often using feature-level splicing or decision-level voting, without considering the temporal and spatial consistency of different modal data; second, there is a lack of environmental adaptation mechanisms, making it impossible to dynamically adjust the weights of each modality according to environmental changes such as illumination and noise; third, bird ecological knowledge is not fully utilized, and the identification results lack ecological rationality verification. These problems prevent the average identification accuracy of existing multimodal methods from being effectively improved, making it difficult to meet the high-precision requirements of actual ecological monitoring.

[0051] Example: Please refer to Figures 1-3 This embodiment provides a distributed collaborative sensing multimodal bird identification method, which mainly consists of multimodal data acquisition, fusion, optimization, and verification, as detailed below: 1. Distributed collaborative sensing: 1.1 Heterogeneous Camera Linkage Architecture: The system employs a wide-angle monitoring unit and a telephoto tracking unit working in tandem. The wide-angle camera uses a low-resolution, wide-field-of-view device to perform panoramic scanning of the monitoring area, i.e., macroscopic observation, to capture the movement position and operational status of birds. The telephoto camera is equipped with a high-resolution, variable-focal-length lens to capture bird details, i.e., microscopic observation, to capture detailed images of bird feathers, postures, etc. The two cameras are connected to an edge computing device via Ethernet to achieve real-time data transmission and processing.

[0052] The inter-frame difference method for detecting bird activity works by comparing the pixel differences between two adjacent frames. In actual monitoring, the camera continuously acquires image sequences. For adjacent frames t and t-1, the difference value D(x,y,t) = |I(x,y,t) - I(x,y,t-1)| is calculated pixel by pixel. In the formula, I(x,y,t) and I(x,y,t-1) are the pixel values ​​at coordinates (x,y) of the frames t and t-1, respectively. These pixel values ​​come from the image data acquired by the camera in real time. The color and brightness information of each pixel constitute the pixel value. During calculation, the absolute value of the difference between corresponding pixel values ​​in two frames is directly taken to obtain the difference of each pixel. Then, the image is divided into multiple small regions, such as multiple 16×16 pixel sub-regions, and the sum of the difference values ​​of all pixels in each region, Sum(D), is calculated. This sum reflects the degree of image change in the region. When the degree of image change in a region exceeds 12% of the total number of pixels in that region (Total(Pixels)), it means that there is obvious dynamic change in that region, thus determining the presence of bird activity. This judgment method is based on a large number of experiments and actual monitoring experience to determine the threshold. In practical applications, it can effectively filter out minor interferences, accurately detect bird activity, and thus precisely trigger the telephoto tracking unit.

[0053] The Kalman filter algorithm is used to predict bird trajectories. The data source primarily consists of the bird's previous position, velocity, and other state information, which is obtained from acquired images using image recognition algorithms. In the prediction equation x... k|k-1 =F k x k-1|k-1 +B k u k In the formula, x k|k-1 For the optimal state estimate at time k, x k-1|k-1 F represents the optimal state estimate obtained based on all information at time k-1; k B is the state transition matrix, which is pre-defined based on the physical laws and time intervals of bird movement, and is used to describe the natural change trend of the state from time k-1 to time k; k To control the input matrix; in bird tracking scenarios, since there are usually no human control factors, u k Generally, it is set to 0; then the state at the current moment can be predicted using the optimal estimate from the previous moment. Update equation x k|k =x k|k-1 +K k (z k -H k x k|k-1 In the formula, x k|k-1 The optimal state estimate at time k is obtained by combining the observation value z at the current time. k Correct the predicted values; z k This refers to real-time observation information such as the current location of birds, obtained through image recognition and other methods; H k The observation matrix is ​​used to transform the predicted state into the observation space and compare it with the actual observed values; K k The Kalman gain dynamically adjusts the correction magnitude based on prediction and observation errors; by continuously repeating the prediction and update steps, it achieves real-time and accurate prediction of bird movement trajectories.

[0054] The PID control algorithm calibrates the camera position based on the deviation e(t) between the bird's actual position and the camera's field of view center; the deviation e(t) is obtained by comparing the bird's position in the image, determined by the image recognition algorithm, with a pre-set image center position; the proportional element K... p e(t) is the control quantity generated based on the current deviation; the larger the deviation, the larger the control quantity. (Integral term) Accumulate past deviations to eliminate the system's steady-state error; differential element Adjusting the control input in advance based on the trend of deviation increases system stability. Through the synergistic effect of these three components, the output control input u(t) is determined, i.e., the proportional component K. p e(t), integration stage and differential elements The accumulation of values ​​drives the camera to adjust its position, ensuring that the bird is always centered in the frame and the image is clear. Here, t is a time variable, and K... p K i K d These are the proportional, integral, and derivative coefficients, respectively, and the optimal parameter values ​​can be determined through experimental debugging.

[0055] Taking forest monitoring as an example, when the wide-angle camera detects abnormal pixel changes in the canopy area, and the abnormality exceeds 12% of the total number of pixels in that area, the telephoto camera is triggered. Based on the bird's previous movement trajectory, the Kalman filter algorithm predicts its flight direction and speed, and adjusts the focal length of the telephoto lens. Then, through PID control, the lens angle is rotated towards the abnormal area to keep it in the center of the image, thereby capturing a clear picture of the abnormal area.

[0056] 1.2 Acoustic Positioning System: It consists of a circular array of 8 omnidirectional microphones with an element spacing of 0.5 meters. Audio signals are acquired through an audio acquisition card at a sampling rate of 44.1 kHz and a quantization precision of 16 bits. The microphone array is connected to an edge computing device to achieve real-time signal processing.

[0057] TDOA positioning is based on the characteristics of sound propagation, simulating the human ear's ability to locate sound sources. Sound travels at a fixed speed c, approximately 340 m / s at room temperature. When birds call, due to the varying distances between the microphones and the sound source, there is a time difference τ between the sound reaching each microphone. ij This time information is recorded by a high-precision clock built into the microphone. The time difference is obtained by comparing the arrival times of sounds recorded by different microphones. Based on the speed of sound c and the time difference τ... ij Using formula d ij =cτ ij =||SM i |-|SM j|| Calculate the distance difference from the sound source to the two microphones; where S represents the location of the sound source, M i and M j These are the positions of the i-th and j-th microphones, respectively, and the microphone position information has been accurately measured and determined when the microphone array is installed. By measuring the distance difference between multiple microphone pairs and combining it with the geometric layout of the microphone array, the specific coordinates of the sound source in three-dimensional space can be calculated using triangulation or hyperbolic positioning algorithms.

[0058] STFT converts time-domain audio signals into frequency-domain signals. Its calculation process involves segmenting the audio signal into frames, multiplying each frame by a window function, such as a Hanning window, to reduce spectral leakage, and then performing a Fourier transform on each frame to obtain the amplitude and phase information of different frequency components. By analyzing the resulting spectrum, the dominant frequency components of environmental noise can be identified. These components can be determined through statistical analysis of a large amount of actual audio data. For example, in wetland environments, wind noise is mostly concentrated in the 100–200 Hz frequency band. Band-stop filters are designed for these dominant noise frequencies, and by setting the passband and stopband frequency ranges of the filters, noise in the corresponding frequency bands can be suppressed. For residual noise, a DNN is used for further suppression. The DNN is trained with a large amount of labeled noisy frequency and clean audio data to learn the mapping relationship from noisy frequency to clean audio, and the network parameters are adjusted with the goal of minimizing the mean square error. The processed audio signal is used to extract acoustic features such as MFCC. The MFCC extraction process includes pre-emphasis to enhance high-frequency signals, framing, windowing, STFT, Mel filtering to convert linear frequencies into Mel frequencies that conform to human hearing characteristics, discrete cosine transform, and other steps, and finally obtains feature vectors that can effectively characterize the features of bird calls.

[0059] The formula for converting a time-domain audio signal to the frequency domain is: ; In the formula, x(n) is the time-domain audio signal, n is the time frame index, m is the intra-frame sampling point index, x(n+m) is the audio signal under the time frame index + intra-frame sampling point index, w(m) is the Hanning window function, M is the window length, set to 512, and N is the number of FFT points, set to 1024. The complex exponential modulation factor is the core component for converting from the time domain to the frequency domain. Based on the idea of ​​Fourier transform, j is the imaginary unit, satisfying j 2=-1, making this factor a complex number that can simultaneously represent the amplitude and phase information of the frequency; N is the number of FFT points, which is the computation length when performing the Fast Fourier Transform (FFT), usually greater than or equal to the frame length M. If it is insufficient, zeros are padded to the frame. It determines the resolution of the frequency domain. The larger N is, the finer the frequency details that can be resolved in the frequency domain; k is the frequency index, corresponding to different frequency points in the frequency domain. k ranges from 0 to N-1, and each k corresponds to a specific frequency. The actual frequency value is kf. s / N,f s It is the audio sampling rate. By using k, the time-domain signal can be decomposed into different frequency components.

[0060] Noise suppression involves analyzing the spectrogram to identify the dominant frequencies of environmental noise and designing band-stop filters to remove them. For example, in wetland environments, a fourth-order Butterworth band-stop filter is designed for wind noise in the 100–200 Hz range. For residual noise, a deep neural network (DNN) is used for further suppression. The network structure contains three fully connected layers with 256, 128, and 64 neurons in each layer, respectively, and is trained using the minimum mean square error (MSE) as the loss function.

[0061] The 13-dimensional Mel frequency cepstral coefficients (MFCC) are extracted from the processed audio. The specific steps include pre-emphasis, framing, windowing, STFT, Mel filtering, and discrete cosine transform (DCT).

[0062] 1.3 Environmental Parameter Acquisition and Spatiotemporal Synchronization: Integrated with digital temperature, humidity, and light intensity sensors, it collects environmental data every 10 seconds; each sensor connects via I... 2 The C bus connects to the main control chip, adding millisecond-accurate timestamps to the data.

[0063] The network time protocol (NTP) is used to calibrate the clock with the server to ensure that the time error of each sensor is less than 8ms; the edge computing device acts as an NTP client and periodically requests time synchronization from the server.

[0064] The Dynamic Time Warping (DTW) algorithm is used to align the temporal dimensions of image, audio, and environmental data. DTW finds the optimal matching path by calculating the similarity between two time series. Combined with Geographic Information System (GIS) or sensor physical coordinates, spatial location information is assigned to each data point, achieving spatiotemporal consistency of multimodal data and laying the foundation for subsequent fusion processing.

[0065] 2. Adaptive Data Processing Framework: 2.1 Image preprocessing workflow: The improved U-Net network model employs an encoder-decoder structure. The encoder consists of four downsampling modules, each containing two 3×3 convolutional layers and one 2×2 max-pooling layer. The decoder consists of four upsampling modules, each containing two 3×3 convolutional layers and one 2×2 deconvolutional layer. Network training is divided into two stages: One method involves pre-training on the ImageNet and COCO datasets for 15 epochs, freezing the parameters of the first 3 layers of the network, using a stochastic gradient descent (SGD) optimizer, a learning rate of 0.01, and a batch size of 32.

[0066] Secondly, fine-tuning was performed, which involved training on a bird image dataset for 25 epochs, unfreezing all parameters, using the Adam optimizer with a learning rate of 0.001, an L2 regularization coefficient of 0.0001, and a batch size of 16 to ensure the accuracy of semantic segmentation.

[0067] After training is completed, image denoising and enhancement are performed. When calculating the local region variance of the image to evaluate the noise level, the image is first divided into multiple small regions. For the pixel value in each region, its average value is calculated. Then, the squares of the differences between the average value and each pixel value are summed and divided by the total number of pixels in the region to obtain the local region variance of the image. The filtering method is selected based on the comparison result of the variance and the set threshold. During filtering, Gaussian filtering effectively removes random noise from low-noise images by weighting pixel values ​​with weights determined by a Gaussian function; nonlocal mean filtering compares the similarity of different regions in the image and weights the pixel values ​​of similar regions, better preserving image details and is suitable for high-noise images; CLAHE enhances local contrast by performing block histogram equalization on the image; and gamma correction improves the overall brightness and contrast of the image by adjusting the power of pixel gray values, highlighting key features of birds.

[0068] 2.2 Audio Enhancement Processing: Based on the STFT analysis results, a band-stop filter is designed to remove the dominant noise frequency band. For example, for traffic noise in the urban environment of 200-500Hz, a fifth-order elliptic band-stop filter is designed. The residual noise is further processed by a DNN. The network input is 256-dimensional spectral features, and the output is 128-dimensional denoised features. The training is conducted for 100 epochs.

[0069] In the MFCC calculation process, the number of Mel filters is increased from 26 to 40, and the DCT transform retains the first 13 coefficients, improving the feature representation capability.

[0070] 2.3 Dynamic parameter adjustment: In the Q-learning algorithm, the state space S contains parameters such as recognition accuracy, image noise level, and audio signal-to-noise ratio, which are obtained through real-time monitoring and calculation; the action space A is a set of adjustable preprocessing parameters; the reward function R is set according to the change in recognition accuracy; the system tries different actions in each state, obtains reward feedback based on the change in recognition accuracy after the action is executed, updates the Q-table using the reward feedback, and records the expected reward value of different actions in each state; in actual operation, the action with the largest expected reward value is selected according to the Q-table, that is, the optimal preprocessing parameter adjustment strategy is determined.

[0071] The reward function R is used to evaluate the impact of different preprocessing parameter adjustments on recognition accuracy, and uses this feedback to guide the direction of parameter optimization; the specific judgment rules are as follows: If, after performing a parameter adjustment, the bird recognition accuracy improves by more than 5% compared to before the adjustment, a reward of +10 is given. This indicates that the parameter adjustment strategy has significantly improved the recognition effect and has a positive effect on improving system performance. Such operations should be repeated in subsequent decisions. When parameter adjustments cause a drop in recognition accuracy exceeding 3%, a penalty of -5 is applied. This indicates that the adjustment has a negative impact on system performance, and similar operations should be avoided as much as possible in subsequent decisions to prevent further deterioration of recognition performance. If the change in recognition accuracy is between -3% and +5% after parameter adjustment, or if it has no significant impact on recognition accuracy, the reward value is set to 0. This means that the adjustment has little impact on system performance and does not have a significant bias in subsequent decision-making. Through this reward mechanism, the system gradually learns the optimal parameter combination that can improve recognition accuracy as it tries different parameter adjustment actions.

[0072] 3. Layered feature fusion architecture: 3.1 Feature-level fusion: When reducing the dimensionality of image features, PCA is used to reduce the dimensionality of image features. First, the covariance matrix of the image feature vectors is calculated. The covariance matrix reflects the correlation between feature vectors. Eigenvalue decomposition is performed on the covariance matrix to obtain feature vectors and eigenvalues. The magnitude of the eigenvalues ​​represents the amount of information contained in the corresponding feature vector. The top K principal components with larger eigenvalues ​​are selected, and the original high-dimensional image feature vectors are projected into a low-dimensional space composed of these principal components, reducing the dimensionality while retaining 90% to 95% of the image information.

[0073] When performing dimensionality reduction on audio features, LDA is used to reduce the dimensionality of the audio features, and the within-class scatter matrix S is calculated. w and the inter-class scatter matrix S b S wS describes the dispersion of features among samples of the same category. b It describes the degree of difference between features of samples of different categories; by solving the generalized eigenvalue problem, a dimension reduction matrix is ​​obtained, and the original audio features are projected into a low-dimensional space, making intra-class features more compact and inter-class features more dispersed, thereby improving inter-class discriminative power. The dimensionality-reduced image and audio features are concatenated to form a joint feature vector, integrating the effective information from both modalities to provide a more comprehensive feature representation for subsequent classification.

[0074] The dimensionality-reduced image and audio features are concatenated into a 70-dimensional joint feature vector: Fusion_Feature=[Image_Feature reduced; Audio_Feature reduced ]; Image_Feature after PCA dimensionality reduction reduced (50-dimensional) and the audio feature vector Audio_Feature after dimensionality reduction by LDA reduced (20-dimensional) Vertical concatenation is performed, that is, the audio feature vector is connected below the image feature vector to form a 70-dimensional joint feature vector Fusion_Feature; this joint feature vector integrates the effective information of both image and audio modalities, providing a more comprehensive and richer feature representation for subsequent classifiers, and is used for the classification and identification of bird species.

[0075] 3.2 Decision-level integration: When ensemble classifiers are used, the ensemble classification system dynamically allocates the weights of each classifier based on data quality and environmental conditions, namely feature variance (Var), entropy (Entropy), light intensity (Light), and noise level (Noise). Feature variance reflects the stability of features; a small variance indicates that the feature values ​​are more concentrated. Entropy measures the uncertainty of features; a large entropy indicates that the features contain more complex information. Environmental conditions directly affect the quality and reliability of data from each modality. Using the weighted summation formula The classification decision is made, where PredictedClass represents the identified bird category, n is the number of classifiers, and w i The weights for the i-th classifier can be dynamically determined based on data and environmental conditions. (Probability) ijIt is the predicted probability of the i-th classifier that the sample belongs to class j. Each classifier first predicts the input sample and outputs the probability of belonging to each class. Then, these probabilities are weighted and summed according to the assigned weights. Finally, the sample is classified into the class with the largest weighted sum. At the same time, the recursive feature elimination (RFE) method is used to continuously remove features that contribute less to the classification and optimize the feature combination, thereby reducing the number of features while maintaining a high classification accuracy.

[0076] ; In the formula, α i β i The weighting coefficients for data quality and environmental conditions were determined experimentally. i Environment i Let be the data quality and environmental adaptability scores corresponding to the i-th classifier, respectively; for each classifier i, calculate its corresponding data quality score, Quality. i Environmental adaptability score i Data quality score i The evaluation is based on the calculation of feature variance and entropy. A smaller feature variance indicates higher data stability, while a lower entropy indicates lower data uncertainty. A higher combined score indicates better data quality. The environmental adaptability score is also included. i Then, based on current environmental conditions such as light intensity and noise level, and combined with the historical performance of each classifier in different environments, scores are assigned. For example, in low-light environments, classifiers that rely less on image features are given higher environmental adaptability scores; α i and β i These are weighting coefficients determined through optimization using a large amount of experimental data, used to adjust the relative importance of data quality and environmental adaptability in weight calculation.

[0077] 3.3 Modal complementarity assessment: The mutual information (MI) between image and audio features is calculated to assess modal complementarity. Mutual information measures the degree of dependence between two random variables. A larger MI value indicates a stronger correlation between the two modal features, but a potentially weaker complementarity. Based on the MI value and environmental parameters, the fusion weights of image and audio features are dynamically adjusted using a weighted formula. In nighttime scenes, image quality degrades, while audio information becomes relatively more important. When the calculated MI value is 0.35, the system automatically adjusts the weights, increasing the audio weight from 0.4 to 0.7, making recognition more reliant on audio features and thus improving recognition accuracy.

[0078] The degree of complementarity between image and audio features is quantified by calculating mutual information (MI). The formula for calculating mutual information is as follows: ; In the formula, X and Y represent image features and audio features, respectively, p(x,y) is the joint probability distribution of X and Y, and p(x) and p(y) are the marginal probability distributions of X and Y, respectively. The joint probability distribution and marginal probability distribution are obtained by statistical calculation of a large amount of sample data, and then the mutual information value is obtained. The larger the MI value, the stronger the correlation between the two modal features, and the weaker the complementarity may be.

[0079] 4. Intelligent recognition and decision-making system: 4.1 Dual-track learning strategy: The center loss function Lcenter is used to optimize the feature space distribution, making features of similar samples more clustered and features of dissimilar samples more separated. It is weighted and summed with the classification loss function Lclassification to form the total loss function Ltotal. By adjusting the weight coefficients λ1 and λ2, feature optimization and classification accuracy are balanced. When training the new species White-crowned Swallowtail, by optimizing the total loss function, the intra-class distance of the White-crowned Swallowtail feature vector is reduced by 0.2 and the inter-class distance is increased by 0.3, thereby improving the model's ability to recognize the new species.

[0080] The formula upon which the center loss function is based is as follows: ; In the formula, n is the sample size, f i It is the feature vector of the i-th sample. The category y to which sample i belongs i The feature center.

[0081] The modality weight adjustment formula calculates the mutual information (MI) between image features and recognition results. Image Mutual information (MI) between audio features and recognition results Audio Then calculate MI. Image MI Image With MI Audio The ratio of the sum (reflecting the relative importance of image features) is multiplied by an adjustment factor. α Plus adjustment coefficient β The final weight of the image features in the fusion process is obtained as Final_Weight. Image Final weight of audio features Audio Then it is determined by 1 - Final_Weight Image calculate.

[0082] 4.2 Triple verification mechanism: The temperature scaling method calibrates the predicted probabilities of the model output by adjusting the temperature parameter T, minimizing the negative log-likelihood loss L. nllTo make the probability more consistent with the actual confidence level, the identification results are verified from the perspectives of ecological rationality and data consistency by combining bird ecological knowledge graphs and historical data, in which bird ecological knowledge graphs contain information such as bird distribution and habits.

[0083] Minimize the negative log-likelihood loss L nll The calculation process will be introduced as follows: For each sample i, first predict its probability vector P. i Perform an exponentiation operation on each element of P, that is, exponentiation on each element of P. i Each element in the array is raised to the power of T to obtain P. i T The purpose of this step is to adjust the temperature parameter T to change the distribution of the predicted probability, making the probability value smoother or sharper, thereby calibrating the confidence level.

[0084] Next, for P i T Taking the natural logarithm (log) of each element yields a new set of values. Since the logarithmic function is monotonically increasing within its domain, taking the logarithm does not change the relative probability. At the same time, it converts multiplication into addition, which facilitates subsequent summation calculations.

[0085] Summing the results of the above operations for each sample yields the total log-likelihood of all samples; finally, taking the negative average of the sums gives the negative log-likelihood loss L. nll During model training or parameter tuning, by continuously trying different temperature parameters T, the corresponding L is calculated. nll Choose to make L nll The minimum T value is used as the optimal temperature parameter, so that the predicted probability output by the model more accurately reflects the true probability of the sample belonging to each category.

[0086] 4.3 Proactive Learning and Optimization: Calculate the uncertainty value U of the sample prediction result, such as the entropy value. The larger the entropy value, the higher the uncertainty of the prediction result. Set an uncertainty threshold U. threshold When the sample U>U threshold When these samples are used as active learning samples, they are manually labeled and added to the training set. The incremental learning algorithm is used to update the model, so that the model can improve its ability to identify difficult samples under the training of a small number of key samples. Wherein, the uncertainty value U is obtained by predicting the probability p for each category. j With its logarithm logp j Multiplying them together, we get p. j logp jThe logarithm is used here to amplify the influence of low-probability categories, resulting in higher entropy values ​​for situations where the prediction results are more uncertain, such as when the probabilities of multiple categories are relatively even. Then, p is calculated for all categories. j logp j We sum the results. Since we want a larger entropy value to indicate higher uncertainty, and the above summation result is usually negative, we add a negative sign in front of it. Finally, we get the entropy value of the sample prediction result, which is the uncertainty value U.

[0087] 5. Dynamic Ecological Model: Based on a large amount of global bird distribution data and ecological environment parameters, an initial niche model is constructed using the maximum entropy model. Under known constraints, the maximum entropy model maximizes the entropy of the model, i.e., maximizes the uncertainty of the model, thus obtaining the most reasonable distribution prediction.

[0088] As new observational data accumulates, the random forest model is updated online. The random forest makes predictions by constructing multiple decision trees and combining their results. It has good generalization ability. In regional transfer learning, the differences in ecological characteristics between the source and target regions are compared, and the model parameters are adjusted by a weighted formula to adapt the model to the new regional environment. In the monitoring of the new park, the reliability of the model has been significantly improved after 3 months of updates. The specific process of adjusting the model parameters using the weighted formula is as follows: The difference between Source_Parameter and Target_Parameter is used as the bias of the biometric feature. This bias is then combined with the weights of the corresponding biometric features to obtain the influence coefficients (w) of different biometric features on the model. i (Weight of the i-th biometric feature) × (Source_Parameter) i -Target_Parameter i ), Source_Parameter i and Target_Parameter i The i-th ecological feature parameter is the source region and the target region respectively. The result is the influence coefficient of the i-th biological parameter on the model parameter. Using this method, the influence coefficients corresponding to all biological features are summed to obtain the updated model parameter New_Parameter.

[0089] 6. System self-optimization mechanism: It monitors changes in environmental parameters in real time, such as temperature, humidity, light intensity, and wind speed. When the parameter changes exceed preset thresholds, such as temperature changes exceeding 5°C or light intensity changes exceeding 50%, it automatically adjusts the image and audio data preprocessing algorithm parameters. For example, when the temperature rises and causes a decrease in camera image quality, it automatically enhances the image noise reduction intensity; when the wind speed increases and causes an increase in environmental noise, it raises the audio noise reduction threshold.

[0090] The Particle Swarm Optimization (PSO) algorithm is employed to simultaneously optimize multiple objectives, including recognition accuracy, processing speed, and energy consumption. In PSO, each particle represents a set of system parameter solutions, such as image preprocessing parameters and classifier weights. The particles navigate through the solution space, continuously updating their positions to find the optimal solution, ensuring long-term stable system operation and continuous performance improvement. The particle update process is as follows: Determine the current iteration number t, the particle number i, and the dimension d. For each particle i in dimension d, obtain its current velocity v. id t and multiplied by inertia weight ω Inertial weight ω Used to balance the global and local search capabilities of particles, with a larger [capacity / sufficiency]. ω It is beneficial for searching particles over a wider range, and smaller ones ω This helps the particle to perform a precise search in a local area, and this step yields the first part of the velocity. ω v id t .

[0091] Next, the cognitive part c1r is calculated. 1d t (p id -x id t ), where c1 is the cognitive learning factor, used to adjust the step size of the particle learning towards its historical best position; r 1d t It is a random number within the interval [0,1], introducing randomness into the calculation process and preventing the algorithm from getting trapped in local optima; p id It is the historical best position of particle i in dimension d, x id t It is the current position of particle i in dimension d, (p id -x id t This represents the difference between the particle's current position and its historical best position, multiplied by c1 and r. 1d t This yields the velocity increment that the particle learns from its historical best position.

[0092] Then, calculate the social component C2R.2d t (g d -x id t Here, c2 is the social learning factor, used to adjust the step size of the particle learning towards the global optimal position; r 2d t Similarly, it is a random number within [0,1]; g d It is the globally optimal position found by the entire particle swarm in dimension d, g d -x id t This represents the difference between the particle's current position and the global optimal position, multiplied by c2 and r. 2d t Then, the velocity increment of the particle learning towards the global optimal position is obtained.

[0093] Finally, the three parts of speed ω v id t c1r 1d t (p id -x id t ) and C2R 2d t (g d -x id t Adding these together, we obtain the updated velocity v of particle i in dimension d. id t+1 .

[0094] After obtaining the updated speed v id t+1 Then, compare it with the current position x of particle i in dimension d. id t Add them together to obtain the updated position x of particle i in dimension d. id t+1 .

[0095] The system will periodically compare manually labeled data with the model's recognition results to calculate the model's error rate. The specific calculation method is as follows: count the number of incorrectly identified samples, divide it by the total number of samples participating in the evaluation, and the resulting ratio is the model's error rate. For example, in an evaluation with 100 samples, if the model incorrectly identifies 12 samples, then the error rate of this evaluation is 12 / 100=12%.

[0096] When the calculated error rate exceeds a preset threshold, the system triggers an incremental learning mechanism. At this time, the system automatically collects recently identified incorrect samples. These samples include bird images and audio data that the model has difficulty classifying accurately, along with their corresponding correct category labels. Then, using a gradient-based incremental learning algorithm, the system calculates gradients based on these newly added incorrect samples, building upon the original model parameters. Specifically, for each new sample, the model first performs forward propagation to calculate the prediction result, and then uses backpropagation to calculate the gradient of the error between the prediction result and the true label with respect to each parameter. Finally, the model parameters are updated according to a certain learning rate based on the calculated gradients, enabling the model to better learn the features of these incorrect samples, thereby reducing the error rate, continuously optimizing model performance, and ensuring that the system maintains high recognition accuracy and stability during long-term operation.

[0097] The distributed collaborative sensing multimodal bird identification system and method proposed in this application achieves end-to-end optimization of bird identification technology from data collection to intelligent decision-making by constructing innovative modules such as distributed collaborative sensing, adaptive data processing, and hierarchical feature fusion. Unlike traditional single-modal or simple fusion identification schemes, this invention fully utilizes the complementary characteristics of multimodal data such as images and audio, and combines algorithms such as reinforcement learning and swarm intelligence optimization to dynamically adapt to complex environmental changes. Simultaneously, by introducing ecological models and active learning mechanisms, it not only significantly improves the accuracy and stability of bird identification but also possesses self-optimization and ecological prediction capabilities. This scheme provides a highly innovative and practical technical means for the field of bird ecological monitoring, effectively addressing the shortcomings of existing technologies in terms of environmental adaptability, identification accuracy, and system optimization, and is of great significance for promoting the development of biodiversity conservation and ecological environment monitoring technologies.

[0098] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A distributed collaborative sensing multimodal bird identification method, characterized in that, The identification method includes the following steps: A distributed sensor network is constructed, and the collected data is time-aligned through a spatiotemporal synchronization protocol to form multidimensional sensing data, which includes image data, audio data, and environmental data. The multidimensional sensing data is preprocessed, and a dynamic mapping relationship between recognition accuracy and preprocessing parameters is established. The preprocessing process is adjusted based on real-time performance feedback. The preprocessing of multidimensional sensing data includes image data preprocessing and audio enhancement processing. The preprocessed multidimensional sensing data is dimensionality reduced to retain key feature information. The processed key features are then concatenated to form a joint feature vector. At the same time, an integrated library of classification methods is built. By evaluating the feedback of each classification method under different data quality, weights are dynamically allocated to form an adaptive fusion strategy. By analyzing the correlation between image and audio features, the complementarity of the two modal information is quantified, a modality-environment correlation model is established, and adaptive weight adjustment is performed based on real-time acquired environmental data to output optimized feature data. The confidence level of the identification results is corrected by probabilistic calibration technology to filter out low-confidence predictions; the ecological rationality of the identification results is verified based on the bird ecological knowledge base to exclude conclusions that conflict with biological laws; historical bird distribution data in the same region and season are analyzed, and the accuracy of the current results is evaluated by statistical matching degree. The final identification result is output by comprehensively verifying the results from multiple dimensions.

2. The distributed collaborative sensing multimodal bird identification method according to claim 1, characterized in that, Image data is acquired using a heterogeneous camera linkage architecture, which includes a wide-angle monitoring camera and a telephoto tracking camera. Among them, the wide-angle monitoring camera captures panoramic images of the monitoring area at preset time intervals, obtains continuous image frames, and extracts the areas where the pixels change by comparing the color and brightness data of each pixel in adjacent image frames; when the area of ​​the changed area exceeds the preset ratio threshold of the captured image area, an upload command is issued to determine that bird activity has been detected and to trigger the telephoto tracking camera to track and capture images. The telephoto tracking camera extracts historical location data of birds and establishes a predictive relationship between location and time based on the historical location data to estimate the location data of birds at the next moment. Then, the parameters of the telephoto tracking camera are adjusted according to the estimated location data. During filming, the real-time location information of the birds is continuously acquired and compared with the current camera's perspective to calibrate the camera angle.

3. The distributed collaborative sensing multimodal bird identification method according to claim 2, characterized in that, Audio data is acquired using a microphone array. The microphone array receives sound signals and obtains the time data of sound arrival at each microphone. By analyzing the time differences in sound arrival at different microphones and considering the spatial relationships of the microphones, a geometric calculation model is established to determine the sound source location. After acquiring the sound signal, audio enhancement processing is performed. The specific process of audio enhancement processing is as follows: First, analyze the characteristics of environmental noise, eliminate the dominant components of environmental noise, suppress residual noise, and highlight bird sound signals; The processed bird sound signals are used to extract sound feature parameters that reflect the differences between bird species as valid audio. The sound feature parameters include the intensity of different frequency bands and the duration of sound.

4. The distributed collaborative sensing multimodal bird identification method according to claim 3, characterized in that: Image data preprocessing utilizes an improved deep learning semantic segmentation model to perform semantic segmentation on the images, dividing the acquired image data into bird regions and background regions; The model was pre-trained on a general image dataset, with the first 3-8 convolutional layers frozen during the pre-training phase, and trained using a stochastic gradient descent optimizer. Then, fine-tuning was performed on bird image data from a preset region, unfreezing all layers and using the Adam optimizer to reduce the learning rate and add L2 regularization. A distillation learning method was employed, using EfficientNet-B7 as the teacher model and the lightweight MobileNetV3 as the student model. The knowledge distillation loss was calculated by adjusting the temperature parameter, transferring knowledge from the teacher model to the student model. Based on the local variance of the image, the noise level is estimated, and an appropriate filtering method is selected for adaptive denoising. The image parameters are then enhanced and optimized by adjusting adaptive histogram equalization and gamma correction.

5. The distributed collaborative sensing multimodal bird identification method according to claim 4, characterized in that: Dimensionality reduction processing of multidimensional sensory data includes image dimensionality reduction and audio dimensionality reduction, followed by fusion of the reduced image and audio features; the specific details are as follows: Principal component analysis is used to reduce the dimensionality of image features. By calculating the covariance matrix of the image feature vector, eigenvalue decomposition is performed to retain the principal components that interpret the image information, thus mapping the original high-dimensional image features to a low-dimensional space. Linear discriminant analysis is used to reduce the dimensionality of audio features. The intra-class and inter-class scatter matrices of the audio features are calculated, the generalized eigenvalue problem is solved, and the most discriminative eigenvectors are found to reduce the dimensionality of the audio features. The dimensionality-reduced image and audio features are concatenated to construct an integrated system. Each classifier is trained separately for different types of features. The quality of modal data is evaluated by analyzing the statistical values ​​of the features. Based on the evaluation results and environmental conditions, weights are dynamically assigned to each classifier. At the same time, a recursive feature elimination method is used to select the most discriminative feature subset from the original features for optimization.

6. The distributed collaborative sensing multimodal bird identification method according to claim 5, characterized in that: When verifying the recognition results, a dual-track learning process is adopted to optimize the feature space distribution through the center loss function. During the training process, the center loss weight is adjusted to make the features of samples of the same class similar and the features of samples of different classes far apart. The confidence level of the identification results was calibrated using a temperature scaling method. Temperature parameters were determined on the validation set by minimizing the negative log-likelihood loss. The ecological habits, distribution range, and seasonal activities of birds were integrated. An ecological knowledge graph was constructed using the Neo4j graph database. The identification results were compared with the knowledge in the graph to verify the ecological rationality of the identification results. Analyze historical bird distribution data for the same region and season, count the bird species that have appeared in history and their frequencies, and compare and evaluate them with the current identification results; at the same time, calculate the entropy value of the prediction result for each sample, select samples with high entropy values ​​as active learning sample selection objects, and perform incremental learning on the model every time the number of newly labeled samples reaches a preset threshold.

7. The distributed collaborative sensing multimodal bird identification method according to claim 6, characterized in that: After outputting the final identification results, an initial niche model is constructed, the details of which are as follows: Hierarchical clustering algorithm is used to cluster different ecological regions based on environmental characteristics and bird distribution patterns; as new observation data accumulates, the distribution prediction model is updated online using a random forest model, and the model parameters are updated using an incremental learning algorithm. When there is insufficient data in the new region, the model is migrated from a similar ecological region. Using limited local data, the model parameters are adjusted to adapt to the new region by analyzing the differences in mean and variance of the characteristics of the source and target regions. The reliability coefficient of the model is comprehensively evaluated based on the amount of data, the degree of environmental similarity, and the consistency of prediction results. When the reliability coefficient is lower than the set threshold, the training process update instruction is triggered.

8. The distributed collaborative sensing multimodal bird identification method according to claim 7, characterized in that: Environmental data includes temperature, humidity, and light intensity. The trend of change is analyzed by calculating the moving average and standard deviation of the environmental data. When the parameter change exceeds the set threshold, the parameters of image preprocessing and audio enhancement processing are adjusted based on the change. The particle swarm optimization algorithm is used to simultaneously optimize recognition accuracy, processing speed and energy consumption. During the optimization process, the best solution of each generation is retained to form an elite solution set to guide subsequent searches. Regularly compare manually labeled data with recognition results, calculate the error rate metric, and adjust the recognition performance through incremental learning when the error rate exceeds the preset error threshold.

9. The distributed collaborative sensing multimodal bird identification method according to claim 8, characterized in that: The data storage adopts a hierarchical storage strategy, storing frequently accessed data in high-speed storage devices and storing historical data with access frequency below the access threshold in backup devices; Based on access time and frequency, migrate data that is not accessed within a unit of time from high-speed storage devices to low-speed storage devices: use data redundancy storage methods; It provides a unified data access interface, caches frequently queried data, and deletes used data within the time period of 0.1h to 0.5h when the cache space is insufficient, thereby reducing the data access time.

10. A distributed collaborative sensing multimodal bird identification system, using the method described in any one of claims 1 to 9, characterized in that, The system includes: The data acquisition module constructs a distributed sensor network to collect image data, audio data, and environmental data, and performs spatiotemporal synchronization to form multidimensional sensing data. The data preprocessing module preprocesses the multidimensional sensing data and adjusts the preprocessing process based on the dynamic mapping relationship between recognition accuracy and preprocessing parameters. The feature processing module reduces the dimensionality of the preprocessed multidimensional perceptual data and concatenates them to form a joint feature vector, and builds an integrated library of classification methods to achieve adaptive fusion. The modality-environment association module analyzes the correlation between image and audio features, establishes a modality-environment association model, and adjusts the weights based on environmental data. The results processing module outputs the final identification results through probability calibration, ecological knowledge base verification, and historical data matching analysis. The niche model building module builds and updates the niche model and performs model migration when there is insufficient data in new regions. The environmental data processing module analyzes the changing trends of environmental data, adjusts preprocessing parameters, and optimizes system performance. The data storage module adopts a hierarchical storage strategy and a data redundancy storage method, provides a unified data access interface, and manages data caching.