A method and system for bird flock recognition based on ultra-high-definition video

By synchronously obtaining video streams and audio streams, dynamically adjusting video acquisition parameters, and combining space-time and time-frequency characteristics, the accuracy and robustness of bird flock recognition are improved, solving the problem of low recognition accuracy of traditional methods in complex environments.

CN120012031BActive Publication Date: 2025-07-18SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510502649.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional bird flock recognition methods rely on single video data or audio data, and the recognition accuracy rate is not high in complex environments, making it difficult to meet the needs of ecological protection and environmental monitoring.

Method used

By synchronously obtaining video streams and audio streams, using spectrum features to dynamically adjust the acquisition parameters of the video stream, combining spatiotemporal feature extraction and time-frequency feature extraction, spatial consistency matching and confidence weighting calculations are performed, and the identification results of bird targets are output.

Benefits of technology

It improves the accuracy and robustness of bird flock identification, reduces the computational complexity, and meets practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012031B_ABST
    Figure CN120012031B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for bird flock recognition based on ultra-high-definition video, which relates to the technical field of bird recognition. The method includes synchronously acquiring a video stream and an audio stream of a target area, and dynamically adjusting the acquisition parameters of the video stream based on the spectral features extracted from the audio stream; extracting spatio-temporal features from the dynamically adjusted video stream to output the spatial coordinates and visual confidence of bird targets; extracting time-frequency features and performing sound source localization on the audio stream to output the sound source azimuth angle and acoustic confidence; performing spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determining it as a valid candidate area when the spatial distance between the two is less than a preset threshold; performing weighted calculation on the visual confidence and the acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds the determination threshold, outputting the recognition result that there are birds in the target area. The present invention improves the accuracy and robustness of bird flock recognition through the effective fusion of multi-modal information and dynamic adjustment strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bird recognition, and particularly to a method and system for bird flock recognition based on ultra-high-definition video. Background Art

[0002] With the increasing demands for ecological protection and environmental monitoring, the accurate recognition and monitoring of bird activities in the natural environment have become an important research direction. Especially against the background of the rapid development of ultra-high-definition video technology, how to use ultra-high-definition video data to achieve efficient and accurate bird flock recognition is of great significance for fields such as ecological protection, wildlife management, and environmental monitoring. The present invention relates to a method and system for bird flock recognition based on ultra-high-definition video, belonging to the cross field of computer vision and audio processing technology, aiming to improve the accuracy and robustness of bird flock recognition by combining multi-modal information of video and audio.

[0003] Traditional bird flock recognition methods mainly rely on single video data or audio data. Although video-based recognition methods can intuitively capture the visual characteristics of birds, their recognition accuracy will be significantly affected in complex environments (such as light changes, occlusion, etc.). And audio-based recognition methods can capture the call characteristics of birds, but their recognition effect will also be greatly reduced in the presence of background noise or multi-source interference.

[0004] Therefore, it is necessary to provide a method and system for bird flock recognition based on ultra-high-definition video to solve the above technical problems. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method and system for bird flock recognition based on ultra-high-definition video, which improves the accuracy and robustness of bird flock recognition through the effective fusion of multi-modal information and dynamic adjustment strategies.

[0006] The present invention provides a method for bird flock recognition based on ultra-high-definition video, and the method includes the following steps:

[0007] Synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream;

[0008] Extract spatio-temporal features from the dynamically adjusted video stream, and output the spatial coordinates and visual confidence of the bird target;

[0009] Extract time-frequency features and perform sound source localization on the audio stream, and output the sound source azimuth angle and acoustic confidence;

[0010] Perform spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determine it as an effective candidate area when the spatial distance between the two is less than a preset threshold;

[0011] Perform weighted calculation on the visual confidence and acoustic confidence of the effective candidate region, and when the fusion confidence obtained from the weighted calculation exceeds the determination threshold, output the recognition result that there are birds in the target region.

[0012] Preferably, synchronously acquire the video stream and audio stream of the target region, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream, including:

[0013] Extract the energy distribution ratios of the preset high-frequency band and preset low-frequency band in the audio stream;

[0014] When the energy distribution ratio of the high-frequency band exceeds the first threshold, increase the resolution and frame rate of the video stream according to a preset proportional coefficient;

[0015] When the energy distribution ratio of the low-frequency band exceeds the second threshold, decrease the resolution and frame rate of the video stream according to a preset proportional coefficient;

[0016] If the energy distribution ratios of both the high-frequency band and the low-frequency band do not exceed the corresponding thresholds, keep the current acquisition parameters unchanged.

[0017] Preferably, perform spatio-temporal feature extraction on the dynamically adjusted video stream and output the spatial coordinates and visual confidence of the bird target, including:

[0018] Perform spatio-temporal joint modeling on at least three adjacent frames of the video stream, and extract the fusion features including temporal motion features and spatial texture features;

[0019] Based on the fusion features, generate candidate regions including spatial coordinates and initial confidence through a pre-trained target detection network;

[0020] Perform non-maximum suppression processing on the candidate regions, and output the spatial coordinates and corresponding visual confidence of the final bird target.

[0021] Preferably, perform time-frequency feature extraction and sound source localization on the audio stream, and output the sound source azimuth angle and acoustic confidence, including:

[0022] Perform short-time Fourier transform on the audio stream to obtain the time-frequency spectrogram;

[0023] Detect bird call features on the time-frequency spectrogram and mark potential sound sources;

[0024] Determine the sound source azimuth angle of the potential sound source by analyzing the time difference of the sound signals received by at least two microphones;

[0025] Calculate and output the acoustic confidence based on the similarity between the time-frequency spectrogram and a preset bird sound template, and in combination with the sound source azimuth angle.

[0026] Preferably, the spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determining it as a valid candidate region when the spatial distance between the two is less than a preset threshold, includes:

[0027] Map the spatial coordinates to the two-dimensional plane coordinate system of the video frame, and convert the sound source azimuth angle into the projection coordinates in the two-dimensional plane coordinate system;

[0028] Calculate the Euclidean distance between the spatial coordinates and the projection coordinates, and determine it as spatial consistency matching when the Euclidean distance is less than or equal to the preset threshold;

[0029] Perform confidence correction on the valid candidate region with successful matching, where:

[0030] When the Euclidean distance is less than 50% of the preset threshold, enhance the visual confidence and acoustic confidence according to a preset confidence enhancement ratio;

[0031] When the Euclidean distance is between 50% and 100% of the preset threshold, dynamically attenuate the confidence according to the proportional relationship between the Euclidean distance and the preset threshold.

[0032] Preferably, the dynamic attenuation includes:

[0033] Calculate the normalized proportional value of the Euclidean distance to the preset threshold, denoted as the attenuation coefficient;

[0034] Perform attenuation calculations on the visual confidence and acoustic confidence of the valid candidate region respectively;

[0035] When both the attenuated visual confidence and acoustic confidence are lower than the preset confidence lower limit, eliminate the valid candidate region.

[0036] Preferably, perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate region, and when the fused confidence obtained by the weighted calculation exceeds the determination threshold, output the recognition result that there are birds in the target region, including:

[0037] Calculate the dynamic weight distribution coefficient based on the ratio of the Euclidean distance to the preset threshold, where the weight coefficient of the visual confidence is linearly related to the reciprocal of the ratio;

[0038] Perform weighted fusion on the dynamic weight distribution coefficient, the visually confident degree after confidence correction, and the acoustic confidence, specifically expressed as:

[0039]

[0040] Among them, is the fusion confidence, is the Euclidean distance, is the preset threshold, and are the corrected visual confidence and acoustic confidence respectively;

[0041] When the fusion confidence is greater than or equal to the determination threshold, it is determined that there is a bird target in the effective candidate area, and the recognition result is output. Otherwise, it is determined that there is no bird target in the effective candidate area.

[0042] Preferably, the determination threshold is dynamically adjusted according to the acquisition parameters of the video stream, specifically including:

[0043] The determination threshold is determined by looking up a table in a pre-established mapping relationship table between the video stream acquisition parameters and the determination threshold, where the mapping relationship table contains the optimal determination thresholds corresponding to different combinations of resolution and frame rate.

[0044] The present invention provides a bird flock recognition system based on ultra-high-definition video for implementing a bird flock recognition method based on ultra-high-definition video. The system includes:

[0045] A parameter dynamic adjustment module, configured to synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream;

[0046] A video stream processing module, configured to extract spatio-temporal features from the dynamically adjusted video stream, and output the spatial coordinates and visual confidence of the bird target;

[0047] An audio stream processing module, configured to extract time-frequency features and perform sound source localization on the audio stream, and output the sound source azimuth angle and acoustic confidence;

[0048] A region determination module, configured to perform spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determine it as an effective candidate region when the spatial distance between the two is less than the preset threshold;

[0049] A result output module, configured to perform weighted calculation on the visual confidence and acoustic confidence of the effective candidate region, and output the recognition result that there are birds in the target area when the fusion confidence obtained by the weighted calculation exceeds the determination threshold.

[0050] Compared with the related technology, the bird flock recognition method and system based on ultra-high-definition video provided by the present invention have the following beneficial effects:

[0051] The present invention synchronously acquires the video stream and audio stream of the target area, and dynamically adjusts the acquisition parameters of the video stream based on the spectral features extracted from the audio stream to meet the recognition requirements under different environmental conditions. Meanwhile, by combining the spatio-temporal feature extraction of the video stream, the time-frequency feature extraction of the audio stream, and sound source localization, spatial consistency matching and confidence-weighted calculation are realized, and finally the recognition result of the presence of birds in the target area is output. The purpose of the present invention is to improve the accuracy and robustness of bird flock recognition through the effective fusion of multi-modal information and dynamic adjustment strategies, while reducing the computational complexity to meet the actual application requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of a method for bird flock recognition based on ultra-high definition video provided by the present invention;

[0053] Figure 2 is a module structure diagram of a bird flock recognition system based on ultra-high definition video provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only parts related to the present invention rather than all structures are shown in the drawings for the convenience of description. In addition, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.

[0055] In addition, it should be noted that only parts related to the present invention rather than all content are shown in the drawings for the convenience of description. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0056] Embodiment 1

[0057] The present invention provides a method for bird flock recognition based on ultra-high definition video, as shown in Figure 1 The method includes the following steps:

[0058] S1: Synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream.

[0059] Specifically, step S1 includes the following steps:

[0060] S11: Extract the energy distribution ratios of the preset high-frequency band and the preset low-frequency band in the audio stream.

[0061] In this embodiment, in the audio stream processing, first, the continuously input audio signal is frame-divided, and a Hanning window function is used for windowing operation. The length of each frame is 50 ms, and the frame shift is 25 ms. The preset high-frequency band (8 kHz to 16 kHz) and low-frequency band (100 Hz to 500 Hz) are separated by a band-pass filter bank, and the energy values of these two frequency bands in the time dimension are calculated respectively. The high-frequency band energy ratio refers to the ratio of the high-frequency band energy value to the total energy of the entire audio signal, and the calculation method of the low-frequency band energy ratio is the same. All energy values are calibrated in decibel units, and finally the dynamic ratio distributions of these two frequency bands are output.

[0062] S12: When the energy distribution ratio of the high-frequency band exceeds the first threshold, increase the resolution and frame rate of the video stream according to a preset proportional coefficient.

[0063] In this embodiment, when the high-frequency band energy ratio exceeds the preset first threshold (e.g., 30%), the adjustment mechanism of the video acquisition parameters is automatically triggered. The specific adjustment scheme includes: increasing the video resolution according to a preset magnification factor (such as 1.5 times), for example, from 1920×1080 pixels to 3840×2160 pixels (i.e., 4K resolution); at the same time, increasing the video frame rate from 30 frames per second to 60 frames per second. This adjustment strategy is based on the strong correlation between the high-frequency band energy and specific bird activities (such as wing flapping, chirping), and enhances the ability to capture details of bird activities by increasing the video acquisition parameters. During the parameter adjustment process, a smooth transition algorithm is used to avoid sudden changes in the picture quality. Among them, the resolution switching is achieved through bilinear interpolation technology, and the frame rate adjustment uses the frame skipping compensation technology to maintain the smooth playback of the video.

[0064] S13: When the energy distribution ratio of the low-frequency band exceeds the second threshold, decrease the resolution and frame rate of the video stream according to a preset proportional coefficient.

[0065] In this embodiment, when the energy ratio of the low-frequency band exceeds a preset second threshold (e.g., 40%), it is determined that there is low-frequency interference in the environment (such as wind noise, mechanical noise, etc.). At this time, the video acquisition parameters are reduced according to a preset reduction factor (e.g., 0.7 times). The specific operations include reducing the resolution from 4K to 1080p and reducing the frame rate from 60 frames per second to 30 frames per second. This parameter reduction strategy reduces the occupancy of system computing resources by reducing the amount of video stream data, and at the same time uses motion adaptive filtering technology to suppress the impact of low-frequency noise on video quality. In the specific implementation process, the resolution parameter is adjusted first, followed by the frame rate parameter, and the time interval between the two parameter adjustments is not less than 2 seconds to prevent frequent fluctuations in parameter settings.

[0066] S14: If the energy distribution ratios of both the high-frequency band and the low-frequency band do not exceed the corresponding thresholds, the current acquisition parameters are maintained unchanged.

[0067] In this embodiment, when the energy ratio of the high-frequency band does not exceed 30% and the energy ratio of the low-frequency band does not exceed 40%, the current video acquisition parameters are maintained unchanged. In this state, the fluctuations of the energy ratios of the two bands are continuously monitored: if the fluctuation amplitude is less than 5% within 5 consecutive seconds, it is determined that the environment is in a stable state and enters the low-power operation mode; if the fluctuation amplitude exceeds 5%, the threshold judgment process is restarted. In order to prevent the system from frequently switching parameters near the threshold, a hysteresis interval is specially set (e.g., the high-frequency threshold fluctuates up and down by 3%, and the low-frequency threshold fluctuates up and down by 5%). Only when the energy ratio exceeds the range of these hysteresis intervals will the parameter adjustment operation be performed.

[0068] In addition, the involved first threshold and second threshold are both obtained through experimental calibration. These thresholds have undergone a large number of actual scenario tests and performance optimizations to ensure the best recognition effect under different environmental conditions.

[0069] S2: Extract spatio-temporal features from the dynamically adjusted video stream, and output the spatial coordinates and visual confidence of the bird target.

[0070] Specifically, step S2 includes the following steps:

[0071] S21: Perform spatio-temporal joint modeling on at least three adjacent frames of the video stream, and extract the fusion features including time motion features and spatial texture features.

[0072] In this embodiment, during the spatio-temporal feature extraction of the video stream, first, the dynamically adjusted video sequence is preprocessed, and three consecutive frames of images are selected to form a spatio-temporal analysis unit. A three-dimensional convolutional neural network (3DCNN) is used to jointly model the video sequence. The size of the first-layer convolutional kernel is set to 3×3×3 (time × height × width), the stride is 1×1×1, and a total of 64 filters are used to extract spatio-temporal features. In the time dimension, the motion features of the birds are captured by calculating the optical flow field between adjacent frames, and the Farneback dense optical flow algorithm is used to calculate the pixel-level displacement vector; in the spatial dimension, an improved ResNet50 network is used to extract multi-scale texture features, including static features such as feather texture and beak shape. The time motion features and the spatial texture features are cascaded and fused at the feature level to form a 1280-dimensional fused feature vector.

[0073] S22: Based on the fused features, generate candidate regions containing spatial coordinates and initial confidence levels through a pre-trained object detection network.

[0074] In this embodiment, based on the extracted fused feature vector, it is input into a pre-trained Faster R-CNN object detection network for bird object detection. This network consists of two parts: a Region Proposal Network (RPN) and a detection network. The RPN network generates approximately 2000 candidate regions, and each candidate region outputs spatial coordinates (center point coordinates and width and height) and an initial confidence score. The detection network performs secondary classification and regression on the candidate regions, and uses the softmax function to calculate the probability value belonging to the bird category as the initial confidence level. The confidence threshold is set to 0.7 to filter out low-quality proposal boxes. When training the network, the cross-entropy loss function and the smooth L1 loss function are jointly optimized, and it is trained to convergence on a dataset containing 50 common bird species.

[0075] S23: Perform non-maximum suppression processing on the candidate regions, and output the spatial coordinates and corresponding visual confidence levels of the final bird objects.

[0076] In this embodiment, non-maximum suppression (NMS) processing is performed on the candidate regions output by the detection network, and the overlap threshold (IOU) is set to 0.5. The specific processing process includes: first, all candidate regions are sorted in descending order of confidence, and the candidate box with the highest score is selected as the reference; calculate the intersection-over-union ratio of other candidate boxes with the reference box, and delete the low-score candidate boxes with an intersection-over-union ratio greater than 0.5; iterate the above process until all candidate boxes are processed. The finally output bird objects include accurate spatial coordinates and normalized visual confidence levels, where the confidence level is calibrated through the sigmoid function to ensure the comparability of objects at different scales. For the case where multiple birds appear simultaneously, the top 10 detection results with the highest confidence levels are retained to meet the real-time requirements.

[0077] S3: Extract the time-frequency features and localize the sound source of the audio stream, and output the sound source azimuth angle and acoustic confidence.

[0078] Specifically, step S3 includes the following steps:

[0079] S31: Perform a short-time Fourier transform on the audio stream to obtain a time-frequency spectrogram.

[0080] In this embodiment, during the audio stream processing, first, the collected audio signal is preprocessed and encoded in 16-bit PCM format with a sampling rate of 48 kHz. A short-time Fourier transform (STFT) is performed on the continuous audio stream, and a Hamming window function is used for frame processing. The window length is 1024 sampling points (about 21.3 ms), and the frame shift is 512 sampling points. The spectral components of each frame are calculated through FFT to generate a time-frequency spectrogram with a frequency resolution of 46.9 Hz and a time resolution of 10.7 ms. The spectrogram is converted to Mel scale, and 40 Mel filter banks are used to non-linearly compress the spectral energy, and finally, a time-frequency-energy three-dimensional feature time-frequency spectrogram is output.

[0081] S32: Detect the bird call features on the time-frequency spectrogram and mark the potential sound sources.

[0082] In this embodiment, for bird call feature detection on the time-frequency spectrogram, first, the time-frequency regions with bird call characteristics are located through spectral centroid analysis and spectral flatness calculation. A GMM-HMM-based acoustic model is used to mark the potential sound sources. When training the model, a bird call database with more than 100 hours, containing 5000 call samples of 50 common bird species, is used. For each detected sound source event, the following feature parameters are extracted: fundamental frequency contour (50 - 8000 Hz), harmonic structure (at least 3 harmonics), time modulation feature (amplitude modulation of 10 - 300 Hz), and its start and end times and frequency range are recorded. By calculating the matching degree of these features with the typical bird call features, the potential sound sources with a confidence greater than 0.6 are preliminarily screened out.

[0083] S33: Determine the sound source azimuth angle of the potential sound source by analyzing the time difference of the sound signals received by at least two microphones.

[0084] In this embodiment, an array composed of at least two microphones arranged in space is used for sound source localization. The time difference of arrival (TDOA) of the same sound signal at different microphones is calculated, and the generalized cross-correlation function (GCC-PHAT) method is adopted for time delay estimation with a time resolution of 0.1 ms. According to the geometric configuration of the microphone array (minimum spacing of 0.5 m), the azimuth angle of the sound source is calculated by the spherical intersection algorithm, with a horizontal direction resolution of 2° and a vertical direction resolution of 5°. Within a distance range of 3 m, the positioning accuracy can reach ±0.1 m. For stable sound sources with more than 10 consecutive frames, Kalman filtering is used for trajectory smoothing processing.

[0085] S34: Calculate and output the acoustic confidence based on the similarity between the spectrogram and a preset bird sound template, and in combination with the azimuth angle of the sound source.

[0086] In this embodiment, the detected sound source features are matched with a preset bird sound template library, which contains the MFCC features (39 dimensions), prosody features, and spectral envelope features of each bird species. The dynamic time warping algorithm is used to calculate the similarity between the test sample and the template, and the acoustic confidence is calculated in combination with the stability of the azimuth angle of the sound source (the angle change is less than 5° within 5 consecutive frames). The confidence calculation formula is: Acoustic confidence = 0.7×Spectral similarity + 0.3×Azimuth stability, where the spectral similarity is normalized to the range of 0 - 1 through softmax. Finally, the azimuth angle (0 - 360°) and acoustic confidence (0 - 1) of each sound source are output, and the update frequency is 10 Hz.

[0087] S4: Perform spatial consistency matching based on the spatial coordinates and the azimuth angle of the sound source, and determine it as a valid candidate area when the spatial distance between the two is less than a preset threshold.

[0088] Specifically, step S4 includes the following steps:

[0089] S41: Map the spatial coordinates to the two-dimensional plane coordinate system of the video frame, and convert the azimuth angle of the sound source into the projection coordinates in the two-dimensional plane coordinate system.

[0090] In the process of spatial consistency matching, first, the spatial coordinates obtained from video detection are converted from the pixel coordinate system to the world coordinate system. A three-dimensional coordinate system with the optical center of the camera as the origin is established. Through the camera calibration parameters (including focal length, principal point coordinates, and distortion coefficients) and the known installation height (such as 3 m), the two-dimensional image coordinates are converted into three-dimensional ground coordinates, where the Z-axis coordinate is fixed at 0 (assuming that birds are active near the ground). The coordinate conversion is realized by using a perspective transformation matrix, and the conversion error is controlled within the range of ±0.1 m.

[0091] Convert the sound source azimuth information into projection coordinates in the same world coordinate system. According to the installation position of the microphone array (maintaining a fixed distance of 1 meter from the camera) and the azimuth data (horizontal angle , pitch angle ), calculate the three-dimensional coordinates of the sound source through the spherical coordinate conversion formula: , where is the preset estimated value of the sound source distance (default 3 meters). Considering the sound source localization error, perform Gaussian smoothing on the coordinates ( set to 0.2 meters), and finally obtain the sound source projection coordinates.

[0092] S42: Calculate the Euclidean distance between the spatial coordinates and the projection coordinates, and determine spatial consistency matching when the Euclidean distance is less than or equal to the preset threshold.

[0093] In this embodiment, calculate the Euclidean distance between the video detection coordinates and the sound source projection coordinates. Since the Z coordinate is fixed at 0, set the preset threshold to 30% of the diagonal length of the video detection frame (typical value is 0.5 - 1.5 meters), and determine spatial consistency matching when D is less than or equal to the preset threshold. To improve the calculation efficiency, use the KD-tree data structure to perform fast neighborhood search on the detection targets, and the processing speed can reach 1000 matches per second.

[0094] S43: Perform confidence correction on the valid candidate regions that match successfully, where:

[0095] When the Euclidean distance is less than 50% of the preset threshold, enhance the visual confidence and acoustic confidence according to the preset confidence improvement ratio;

[0096] When the Euclidean distance is between 50% and 100% of the preset threshold, perform dynamic attenuation on the confidence according to the proportional relationship between the Euclidean distance and the preset threshold.

[0097] Among them, the dynamic attenuation includes:

[0098] First, calculate the normalized proportional value of the Euclidean distance to the preset threshold, denoted as the attenuation coefficient.

[0099] Second, perform attenuation calculations on the visual confidence and acoustic confidence of the valid candidate regions respectively.

[0100] Finally, when both the attenuated visual confidence and acoustic confidence are lower than the preset confidence lower limit, eliminate the valid candidate regions.

[0101] S5: Perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate regions. When the fusion confidence obtained from the weighted calculation exceeds the determination threshold, output the recognition result that there are birds in the target region.

[0102] Specifically, step S5 includes the following steps:

[0103] S51: Calculate the dynamic weight distribution coefficient based on the ratio of the Euclidean distance to the preset threshold, where the weight coefficient of the visual confidence is linearly related to the reciprocal of the ratio.

[0104] In the fusion confidence calculation stage, first dynamically allocate weights according to the spatial consistency matching result. For each valid candidate region, determine the weight ratio of the visual confidence and the acoustic confidence based on the ratio of its Euclidean distance to the preset threshold (denoted as d / D). The weight coefficient of the visual confidence is set to (1 d / D), and the weight coefficient of the acoustic confidence is d / D. When the distance is closer to the upper limit of the threshold, the weight proportion of the acoustic confidence is higher, and vice versa, the visual confidence dominates. The weight allocation process uses a linear interpolation algorithm to ensure a smooth transition of the weight coefficient within the range of the distance from 0 to D.

[0105] S52: Perform weighted fusion on the dynamic weight distribution coefficient, the visually corrected confidence, and the acoustic confidence, which is specifically expressed as:

[0106]

[0107] where, is the fusion confidence, is the Euclidean distance, is the preset threshold, and are the corrected visual confidence and acoustic confidence respectively.

[0108] In this embodiment, perform weighted fusion on the dynamic weight, the corrected visual confidence, and the acoustic confidence. Specifically, calculate according to the above formula for each candidate region. For the case of multi-modal data conflict (such as high V and extremely low A), the system sets a conflict detection mechanism, and triggers an artificial review flag when the difference between the two types of confidence exceeds 0.5.

[0109] S53: When the fusion confidence is greater than or equal to the determination threshold, determine that there is a bird target in the valid candidate region and output the recognition result; otherwise, determine that there is no bird target in the valid candidate region.

[0110] In this embodiment, the dynamic adjustment of the determination threshold is achieved through a pre-established parameter mapping table. This mapping table uses video resolution and frame rate as indexes and stores the experimentally calibrated optimal thresholds. For example, the threshold is 0.65 in the 4K@60fps mode and 0.75 in the 1080p@30fps mode. The current video stream acquisition parameters are obtained in real time, and the corresponding threshold is quickly retrieved through a hash table. For parameter combinations that are not covered, the nearest neighbor interpolation method is used to calculate the threshold. For example, the threshold for 3840×1600@45fps is taken as the weighted average of the thresholds for 4K@60fps and 1080p@30fps.

[0111] When the fusion confidence reaches or exceeds the found determination threshold, it is determined that there is a bird target in the area. The recognition result including the spatial coordinates, confidence value, and timestamp is output, and the original data is recorded for subsequent model optimization. For valid candidate areas with a fusion confidence lower than the determination threshold, two-level filtering is performed: First, the invalid areas with a fusion confidence less than 0.3 are discarded, and the remaining areas are stored in the buffer for 5 seconds. If the confidence rises above the threshold during this period, they are reactivated. The final output result is encapsulated in JSON format and includes structured data such as the target ID, coordinate set, and confidence curve.

[0112] Embodiment 2

[0113] The present invention provides a bird flock recognition system based on ultra-high-definition video for implementing a bird flock recognition method based on ultra-high-definition video. Refer to Figure 2 as shown, the system includes:

[0114] A parameter dynamic adjustment module 100 for synchronously acquiring the video stream and audio stream of the target area and dynamically adjusting the acquisition parameters of the video stream based on the spectral features extracted from the audio stream.

[0115] A video stream processing module 200 for extracting spatio-temporal features from the dynamically adjusted video stream and outputting the spatial coordinates and visual confidence of the bird target.

[0116] An audio stream processing module 300 for extracting time-frequency features and performing sound source localization on the audio stream, and outputting the sound source azimuth angle and acoustic confidence.

[0117] A region determination module 400 for performing spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determining it as a valid candidate area when the spatial distance between the two is less than a preset threshold.

[0118] A result output module 500 for performing weighted calculation on the visual confidence and acoustic confidence of the valid candidate area, and outputting the recognition result that there are birds in the target area when the fusion confidence obtained by the weighted calculation exceeds the determination threshold.

[0119] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0120] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0121] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity, or device including the element.

Claims

1. A bird flock recognition method based on ultra-high definition video, characterized in that, The method includes the following steps: Synchronously obtain the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream; Extract spatio-temporal features from the dynamically adjusted video stream, and output the spatial coordinates and visual confidence of the bird target; Extract time-frequency features and perform sound source localization on the audio stream, and output the sound source azimuth angle and acoustic confidence; Perform spatial consistency matching based on the spatial coordinates and the sound source azimuth angle. When the spatial distance between the two is less than a preset threshold, it is determined as a valid candidate area, specifically including: Map the spatial coordinates to the two-dimensional plane coordinate system of the video screen, and convert the sound source azimuth angle into the projection coordinates in the two-dimensional plane coordinate system; Calculate the Euclidean distance between the spatial coordinates and the projection coordinates. When the Euclidean distance is less than or equal to the preset threshold, it is determined as spatial consistency matching; Perform confidence correction on the valid candidate area with successful matching, where: When the Euclidean distance is less than 50% of the preset threshold, enhance the visual confidence and acoustic confidence according to a preset confidence enhancement ratio; When the Euclidean distance is between 50% and 100% of the preset threshold, dynamically attenuate the confidence according to the proportional relationship between the Euclidean distance and the preset threshold; Perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate area. When the fusion confidence obtained by the weighted calculation exceeds the decision threshold, output the recognition result that there are birds in the target area, specifically including: The weighted calculation of the visual confidence and acoustic confidence of the valid candidate area. When the fusion confidence obtained by the weighted calculation exceeds the decision threshold, output the recognition result that there are birds in the target area, including: Calculate the dynamic weight distribution coefficient based on the ratio of the Euclidean distance to the preset threshold, where the weight coefficient of the visual confidence is linearly related to the reciprocal of the ratio; Perform weighted fusion of the dynamic weight distribution coefficient with the visually corrected confidence and acoustic confidence, specifically expressed as: Among them, is the fusion confidence, is the Euclidean distance, is the preset threshold, and are the corrected visual confidence and acoustic confidence respectively; When the fusion confidence is greater than or equal to the decision threshold, determine that there is a bird target in the valid candidate area and output the recognition result. Otherwise, determine that there is no bird target in the valid candidate area.

2. The method for identifying a flock of birds based on an ultra-high-definition video according to claim 1, wherein The synchronously obtaining the video stream and audio stream of the target area, and dynamically adjusting the acquisition parameters of the video stream based on the spectral features extracted from the audio stream includes: Extract the energy distribution ratios of the preset high-frequency band and the preset low-frequency band in the audio stream; When the energy distribution ratio of the high-frequency band exceeds the first threshold, increase the resolution and frame rate of the video stream according to a preset proportional coefficient; When the energy distribution ratio of the low-frequency band exceeds the second threshold, decrease the resolution and frame rate of the video stream according to a preset proportional coefficient; If the energy distribution ratios of both the high-frequency band and the low-frequency band do not exceed the corresponding thresholds, keep the current acquisition parameters unchanged.

3. The method for identifying bird flocks based on ultra-high-definition video according to claim 2, wherein, The extracting spatio-temporal features from the dynamically adjusted video stream, and outputting the spatial coordinates and visual confidence of the bird target includes: Perform spatio-temporal joint modeling on at least three adjacent frames of the video stream, and extract fused features containing temporal motion features and spatial texture features; Based on the fused features, generate candidate regions containing spatial coordinates and initial confidence levels through a pre-trained object detection network; Perform non-maximum suppression processing on the candidate regions, and output the spatial coordinates and corresponding visual confidence levels of the final bird targets.

4. The method for identifying a bird flock based on an ultra-high definition video according to claim 2, wherein, The performing time-frequency feature extraction and sound source localization on the audio stream, and outputting the sound source azimuth angle and acoustic confidence level includes: Perform short-time Fourier transform on the audio stream to obtain a time-frequency spectrogram; Detect bird call features on the time-frequency spectrogram and mark potential sound sources; Determine the sound source azimuth angle of the potential sound source by analyzing the time difference of sound signals received by at least two microphones; Based on the similarity between the time-frequency spectrogram and a preset bird sound template, and in combination with the sound source azimuth angle, calculate and output the acoustic confidence level.

5. The method for identifying a flock of birds based on an ultra-high-definition video according to claim 4, wherein, The dynamic attenuation includes: Calculate the normalized ratio value of the Euclidean distance to a preset threshold, denoted as the attenuation coefficient; Perform attenuation calculations on the visual confidence level and acoustic confidence level of the effective candidate regions respectively; When both the attenuated visual confidence level and acoustic confidence level are lower than the preset confidence level lower limit, eliminate the effective candidate regions.

6. The method for identifying a flock of birds based on an ultra-high-definition video according to claim 5, wherein The decision threshold is dynamically adjusted according to the acquisition parameters of the video stream, specifically including: Determine the decision threshold by looking up a table from a pre-established mapping relationship table between the video stream acquisition parameters and the decision threshold, where the mapping relationship table contains the optimal decision thresholds corresponding to different combinations of resolutions and frame rates.

7. A bird flock recognition system based on ultra-high-definition video, which is used to execute a bird flock recognition method based on ultra-high-definition video according to any one of claims 1 to 6, characterized in that, The system includes: A parameter dynamic adjustment module, configured to synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectral features extracted from the audio stream; A video stream processing module, configured to perform spatio-temporal feature extraction on the dynamically adjusted video stream, and output the spatial coordinates and visual confidence levels of bird targets; An audio stream processing module, configured to perform time-frequency feature extraction and sound source localization on the audio stream, and output the sound source azimuth angle and acoustic confidence level; A region determination module, configured to perform spatial consistency matching based on the spatial coordinates and the sound source azimuth angle, and determine it as an effective candidate region when the spatial distance between the two is less than a preset threshold; A result output module, configured to perform weighted calculation on the visual confidence level and acoustic confidence level of the effective candidate regions, and output the recognition result that there are birds in the target area when the fused confidence level obtained by the weighted calculation exceeds the decision threshold.

Citation Information

Patent Citations

  • Video object detection and segmentation method based on space-time double-branch network

    CN110097568A

  • Monitoring method and device, electronic equipment and storage medium

    CN116684548A

  • Bird type identification method and identification device, and electronic equipment

    CN118861987A