Bird flock identification method and system based on ultra-high-definition video

By synchronously obtaining video streams and audio streams, dynamically adjusting video stream acquisition parameters, and combining spatiotemporal feature extraction and sound source positioning, the high accuracy and robustness of bird flock recognition are achieved, solving the recognition accuracy problem of traditional methods in complex environments.

CN120012031AActive Publication Date: 2025-05-16SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510502649.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional bird flock recognition methods rely on single video data or audio data, resulting in the recognition accuracy being affected in complex environments, especially in the case of light changes, occlusion, background noise or interference from multiple sound sources.

Method used

By synchronously obtaining the video stream and audio stream in the target area, dynamically adjusting the acquisition parameters of the video stream based on the spectrum characteristics extracted in the audio stream, and combining the spatiotemporal feature extraction of the video stream and the time-frequency feature extraction of the audio stream and the sound source positioning, spatial consistency matching and confidence weighting calculation are achieved, and the identification results of birds present in the target area are finally output.

Benefits of technology

It improves the accuracy and robustness of bird flock recognition, adapts to the identification needs under different environmental conditions, reduces the computational complexity, and meets the practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012031A_ABST
    Figure CN120012031A_ABST
Patent Text Reader

Abstract

The invention provides a bird flock recognition method and system based on an ultra-high-definition video, and relates to the technical field of bird recognition, and the method comprises the steps: synchronously obtaining a video stream and an audio stream of a target region, and dynamically adjusting the collection parameters of the video stream based on the spectrum features extracted from the audio stream; performing spatial-temporal feature extraction on the dynamically adjusted video stream, and outputting spatial coordinates and visual confidence of the bird target; performing time-frequency feature extraction and sound source localization on the audio stream, and outputting a sound source azimuth angle and an acoustic confidence coefficient; performing space consistency matching based on the space coordinates and the sound source azimuth angle, and determining the region as an effective candidate region when the space distance between the two is smaller than a preset threshold value; the visual confidence coefficient and the acoustic confidence coefficient of the effective candidate area are subjected to weighted calculation, when the fusion confidence coefficient obtained through weighted calculation exceeds a judgment threshold value, the recognition result that birds exist in the target area is output, and through effective fusion and dynamic adjustment strategies of multi-modal information, the accuracy and robustness of bird flock recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bird identification, and in particular to a method and system for bird flock identification based on ultra-high-definition video. Background Art

[0002] With the growing demand for ecological protection and environmental monitoring, accurate identification and monitoring of bird activities in natural environments has become an important research direction. Especially in the context of the rapid development of ultra-high-definition video technology, how to use ultra-high-definition video data to achieve efficient and accurate bird flock identification is of great significance to the fields of ecological protection, wildlife management, and environmental monitoring. The present invention relates to a bird flock identification method and system based on ultra-high-definition video, which belongs to the intersection of computer vision and audio processing technology, and aims to improve the accuracy and robustness of bird flock identification by combining multimodal information of video and audio.

[0003] Traditional bird flock recognition methods mainly rely on single video data or audio data. Although video-based recognition methods can intuitively capture the visual characteristics of birds, their recognition accuracy will be significantly affected in complex environments (such as lighting changes, occlusion, etc.). Audio-based recognition methods can capture the characteristics of bird calls, but their recognition effect will be greatly reduced in the presence of background noise or interference from multiple sound sources.

[0004] Therefore, it is necessary to provide a bird flock recognition method and system based on ultra-high-definition video to solve the above technical problems. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a bird flock recognition method and system based on ultra-high-definition video, which improves the accuracy and robustness of bird flock recognition through the effective fusion of multimodal information and dynamic adjustment strategy. The present invention provides a method for identifying a flock of birds based on ultra-high-definition video, the method comprising the following steps: Synchronously acquiring a video stream and an audio stream of a target area, and dynamically adjusting acquisition parameters of the video stream based on spectral features extracted from the audio stream; Extracting spatiotemporal features from the dynamically adjusted video stream, and outputting spatial coordinates and visual confidence of the bird target; Extracting time-frequency features and locating sound sources from the audio stream, and outputting sound source azimuth and acoustic confidence; Performing spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determining the region as a valid candidate region when the spatial distance between the two is less than a preset threshold; A weighted calculation is performed on the visual confidence and the acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds a determination threshold, an identification result indicating that there is a bird in the target area is output.

[0006] Preferably, the synchronously acquiring the video stream and the audio stream of the target area, and dynamically adjusting the acquisition parameters of the video stream based on the spectrum features extracted from the audio stream, includes: Extracting the energy distribution ratio of a preset high frequency band and a preset low frequency band in the audio stream; When the energy distribution ratio of the high frequency band exceeds a first threshold, increasing the resolution and frame rate of the video stream according to a preset proportionality coefficient; When the energy distribution ratio of the low frequency band exceeds a second threshold, reducing the resolution and frame rate of the video stream according to a preset proportionality coefficient; If the energy distribution ratios of the high frequency band and the low frequency band do not exceed the corresponding thresholds, the current acquisition parameters are maintained unchanged.

[0007] Preferably, the step of extracting spatiotemporal features from the dynamically adjusted video stream and outputting the spatial coordinates and visual confidence of the bird target comprises: Performing spatiotemporal joint modeling on at least three adjacent frames of the video stream to extract fusion features including temporal motion features and spatial texture features; Based on the fused features, a candidate region including spatial coordinates and initial confidence is generated through a pre-trained target detection network; Non-maximum suppression processing is performed on the candidate area, and the spatial coordinates of the final bird target and the corresponding visual confidence are output.

[0008] Preferably, the extracting time-frequency features and locating the sound source of the audio stream, and outputting the sound source azimuth and acoustic confidence, comprises: Performing short-time Fourier transform on the audio stream to obtain a time-frequency spectrum diagram; detecting bird call features on the time-frequency spectrum diagram and marking potential sound sources; Determine the sound source azimuth of the potential sound source by analyzing the time difference of the sound signals received by at least two microphones; Based on the similarity between the time-frequency spectrum diagram and a preset bird sound template, and in combination with the sound source azimuth, the acoustic confidence is calculated and output.

[0009] Preferably, the performing spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determining as a valid candidate area when the spatial distance between the two is less than a preset threshold, includes: Mapping the spatial coordinates to a two-dimensional plane coordinate system of a video screen, and converting the sound source azimuth angle into a projection coordinate in the two-dimensional plane coordinate system; Calculating the Euclidean distance between the spatial coordinates and the projection coordinates, and determining that the spatial consistency match is achieved when the Euclidean distance is less than or equal to a preset threshold; The confidence of the successfully matched valid candidate area is corrected, where: When the Euclidean distance is less than 50% of the preset threshold, enhancing the visual confidence and the acoustic confidence according to a preset confidence enhancement ratio; When the Euclidean distance is between 50% and 100% of the preset threshold, the confidence is dynamically attenuated according to a proportional relationship between the Euclidean distance and the preset threshold.

[0010] Preferably, the dynamic attenuation includes: Calculate the normalized ratio of the Euclidean distance to a preset threshold value, and record it as an attenuation coefficient; Performing attenuation calculation on the visual confidence and the acoustic confidence of the valid candidate area respectively; When the attenuated visual confidence and acoustic confidence are both lower than the preset confidence lower limit, the valid candidate area is eliminated.

[0011] Preferably, the visual confidence and acoustic confidence of the valid candidate area are weightedly calculated, and when the fusion confidence obtained by the weighted calculation exceeds the judgment threshold, the recognition result that there are birds in the target area is output, including: Calculating a dynamic weight allocation coefficient based on the ratio of the Euclidean distance to the preset threshold, wherein the weight coefficient of the visual confidence is linearly related to the inverse of the ratio; The dynamic weight distribution coefficient is weighted and fused with the visual confidence and acoustic confidence after confidence correction, which is specifically expressed as: in, is the fusion confidence, is the Euclidean distance, is the preset threshold, and They are the corrected visual confidence and acoustic confidence respectively; When the fusion confidence is greater than or equal to the determination threshold, it is determined that there is a bird target in the valid candidate area and the recognition result is output; otherwise, it is determined that there is no bird target in the valid candidate area.

[0012] Preferably, the determination threshold is dynamically adjusted according to the acquisition parameters of the video stream, specifically including: The determination threshold is determined by table lookup from a pre-established mapping relationship table between video stream acquisition parameters and determination thresholds, wherein the mapping relationship table contains optimal determination thresholds corresponding to different resolution and frame rate combinations.

[0013] The present invention provides a bird flock recognition system based on ultra-high-definition video, which is used to perform a bird flock recognition method based on ultra-high-definition video. The system comprises: A parameter dynamic adjustment module, used to synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectrum features extracted from the audio stream; A video stream processing module is used to extract spatiotemporal features of the dynamically adjusted video stream and output the spatial coordinates and visual confidence of the bird target; An audio stream processing module, used for extracting time-frequency features and locating sound sources on the audio stream, and outputting sound source azimuth and acoustic confidence; A region determination module, configured to perform spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determine the region as a valid candidate region when the spatial distance between the two is less than a preset threshold; The result output module is used to perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds the judgment threshold, output the recognition result that there are birds in the target area.

[0014] Compared with the related art, the bird flock identification method and system based on ultra-high-definition video provided by the present invention has the following beneficial effects: The present invention synchronously acquires the video stream and audio stream of the target area, and dynamically adjusts the acquisition parameters of the video stream based on the spectral features extracted from the audio stream to adapt to the recognition requirements under different environmental conditions. At the same time, the spatiotemporal feature extraction of the video stream and the time-frequency feature extraction and sound source localization of the audio stream are combined to achieve spatial consistency matching and confidence weighted calculation, and finally output the recognition result of the presence of birds in the target area. The present invention aims to improve the accuracy and robustness of bird flock recognition through the effective fusion and dynamic adjustment strategy of multimodal information, while reducing the computational complexity to meet practical application needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A flow chart of a bird flock recognition method based on ultra-high-definition video provided by the present invention; Figure 2 A module structure diagram of a bird flock identification system based on ultra-high-definition video provided by the present invention. DETAILED DESCRIPTION

[0016] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention, rather than all structures, are shown in the accompanying drawings. In addition, the embodiments of the present invention and the features in the embodiments may be combined with each other without conflict.

[0017] It should also be noted that, for ease of description, only the parts related to the present invention, but not all of the contents, are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the operations (or steps) as sequential processes, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to methods, functions, procedures, subroutines, subprograms, etc.

[0018] Embodiment 1 The present invention provides a method for identifying bird flocks based on ultra-high-definition video. Figure 1 As shown, the method comprises the following steps: S1: synchronously acquiring a video stream and an audio stream of a target area, and dynamically adjusting acquisition parameters of the video stream based on frequency spectrum features extracted from the audio stream.

[0019] Specifically, step S1 includes the following steps: S11: Extracting the energy distribution ratio of a preset high frequency band and a preset low frequency band in the audio stream.

[0020] In this embodiment, in the audio stream processing, the continuously input audio signal is firstly framed, and the Hanning window function is used for windowing operation, with each frame length of 50ms and a frame shift of 25ms. The preset high frequency band (8kHz to 16kHz) and low frequency band (100Hz to 500Hz) are separated by a bandpass filter group, and the energy values ​​of these two frequency bands in the time dimension are calculated respectively, where the high frequency band energy ratio refers to the ratio of the high frequency band energy value to the total energy of the entire audio signal, and the calculation method of the low frequency band energy ratio is the same. All energy values ​​are calibrated in decibel units, and finally the dynamic proportional distribution of the two frequency bands is output.

[0021] S12: When the energy distribution ratio of the high frequency band exceeds a first threshold, the resolution and frame rate of the video stream are increased according to a preset proportionality coefficient.

[0022] In this embodiment, when the proportion of high-frequency band energy exceeds a preset first threshold (e.g., 30%), the adjustment mechanism of the video acquisition parameters is automatically triggered. The specific adjustment scheme includes: increasing the video resolution according to a preset magnification factor (e.g., 1.5 times), for example, from 1920×1080 pixels to 3840×2160 pixels (i.e., 4K resolution); and increasing the video frame rate from 30 frames per second to 60 frames per second. This adjustment strategy is based on the strong correlation between high-frequency band energy and specific bird activities (such as wing flapping and chirping), and enhances the ability to capture details of bird activities by improving video acquisition parameters. During the parameter adjustment process, a smooth transition algorithm is used to avoid sudden changes in picture quality, in which resolution switching is achieved through bilinear interpolation technology, and frame rate adjustment uses frame skipping compensation technology to maintain smooth video playback.

[0023] S13: When the energy distribution ratio of the low frequency band exceeds a second threshold, reducing the resolution and frame rate of the video stream according to a preset proportionality coefficient.

[0024] In this embodiment, when the proportion of low-frequency energy exceeds a preset second threshold (for example, 40%), it is determined that there is low-frequency interference (such as wind noise, mechanical noise, etc.) in the environment. At this time, the video acquisition parameters are reduced according to a preset reduction factor (such as 0.7 times). Specific operations include reducing the resolution from 4K to 1080p and the frame rate from 60 frames per second to 30 frames per second. This parameter reduction strategy reduces the occupancy of system computing resources by reducing the amount of video stream data, and uses motion adaptive filtering technology to suppress the impact of low-frequency noise on video quality. In the specific implementation process, the resolution parameter is adjusted first, followed by the frame rate parameter, and the time interval between the two parameter adjustments is not less than 2 seconds to prevent frequent fluctuations in parameter settings.

[0025] S14: If the energy distribution ratios of the high frequency band and the low frequency band do not exceed the corresponding thresholds, the current acquisition parameters are maintained unchanged.

[0026] In this embodiment, when the high-frequency band energy proportion does not exceed 30% and the low-frequency band energy proportion does not exceed 40%, the current video acquisition parameters are maintained unchanged. In this state, the fluctuation of the energy proportion of the two frequency bands is continuously monitored: if the fluctuation amplitude is less than 5% within 5 consecutive seconds, the environment is determined to be in a stable state and enters the low-power operation mode; if the fluctuation amplitude exceeds 5%, the threshold judgment process is restarted. In order to prevent the system from frequently switching parameters near the threshold, a hysteresis interval is specially set (for example, the high-frequency threshold fluctuates by 3% and the low-frequency threshold fluctuates by 5%). The parameter adjustment operation will only be performed when the energy proportion exceeds these hysteresis intervals.

[0027] In addition, the first and second thresholds involved are obtained through experimental calibration. These thresholds have been tested in a large number of actual scenarios and performance optimized to ensure that the best recognition effect can be achieved under different environmental conditions. S2: Extracting spatiotemporal features from the dynamically adjusted video stream, and outputting the spatial coordinates and visual confidence of the bird target.

[0028] Specifically, step S2 includes the following steps: S21: Performing spatiotemporal joint modeling on at least three adjacent frames of the video stream, and extracting fusion features including temporal motion features and spatial texture features.

[0029] In this embodiment, in the process of extracting spatiotemporal features of video streams, the dynamically adjusted video sequence is first preprocessed, and three consecutive frames of images are selected to form a spatiotemporal analysis unit. A three-dimensional convolutional neural network (3DCNN) is used to jointly model the video sequence, where the size of the first layer of convolution kernel is set to 3×3×3 (time×height×width), the step size is 1×1×1, and a total of 64 filters are used to extract spatiotemporal features. In the time dimension, the motion features of birds are captured by calculating the optical flow field between adjacent frames, and the pixel-level displacement vector is calculated using the Farneback dense optical flow algorithm; in the spatial dimension, the improved ResNet50 network is used to extract multi-scale texture features, including static features such as feather texture and beak shape. The temporal motion features and spatial texture features are cascaded and fused at the feature layer to form a 1280-dimensional fused feature vector.

[0030] S22: Based on the fused features, a candidate region including spatial coordinates and initial confidence is generated through a pre-trained target detection network.

[0031] In this embodiment, based on the extracted fusion feature vector, it is input into the pre-trained FasterR-CNN target detection network for bird target detection. The network consists of two parts: a region proposal network (RPN) and a detection network. The RPN network generates about 2,000 candidate regions, and each candidate region outputs spatial coordinates (center point coordinates and width and height) and an initial confidence score. The detection network performs secondary classification and regression on the candidate regions, and uses the softmax function to calculate the probability value belonging to the bird category as the initial confidence. The confidence threshold is set to 0.7 to filter out low-quality suggestion boxes. The cross entropy loss function and the smoothL1 loss function are jointly optimized during network training, and trained to convergence on a data set containing 50 common bird species.

[0032] S23: Perform non-maximum suppression processing on the candidate area, and output the spatial coordinates of the final bird target and the corresponding visual confidence.

[0033] In this embodiment, non-maximum suppression (NMS) processing is performed on the candidate regions output by the detection network, and the overlap threshold (IOU) is set to 0.5. The specific processing process includes: first, all candidate regions are sorted in descending order by confidence, and the highest-scoring candidate frame is selected as the benchmark; the intersection-over-union ratio of other candidate frames and the benchmark frame is calculated, and the low-scoring candidate frames with an intersection-over-union ratio greater than 0.5 are deleted; the above process is iterated until all candidate frames are processed. The final output bird target contains precise spatial coordinates and normalized visual confidence, where the confidence is calibrated by the sigmoid function to ensure the comparability of targets of different scales. In the case where multiple birds appear at the same time, the detection results with the top 10 confidence rankings are retained to meet real-time requirements.

[0034] S3: extracting time-frequency features and locating sound sources from the audio stream, and outputting sound source azimuth and acoustic confidence.

[0035] Specifically, step S3 includes the following steps: S31: Perform short-time Fourier transform on the audio stream to obtain a time-frequency spectrum diagram.

[0036] In this embodiment, during the audio stream processing, the collected audio signal is first preprocessed using a 16-bit PCM encoding format with a sampling rate of 48kHz. A short-time Fourier transform (STFT) is performed on the continuous audio stream, and a Hamming window function is used for frame processing, with a window length of 1024 sampling points (about 21.3ms) and a frame shift of 512 sampling points. The spectral components of each frame are calculated by FFT to generate a time-frequency spectrum with a frequency resolution of 46.9Hz and a time resolution of 10.7ms. The spectrum is converted to a Mel scale, and 40 Mel filter groups are used to perform nonlinear compression on the spectrum energy, and finally a time-frequency spectrum with three-dimensional characteristics of time-frequency-energy is output.

[0037] S32: Detect bird call features on the time-frequency spectrum diagram and mark potential sound sources.

[0038] In this embodiment, bird call feature detection is performed on the time-frequency spectrum diagram, and the time-frequency region with bird call characteristics is first located by spectral centroid analysis and spectral flatness calculation. The potential sound source is marked using an acoustic model based on GMM-HMM. A bird call database of more than 100 hours is used for model training, including 5,000 call samples of 50 common birds. For each sound source event detected, the following characteristic parameters are extracted: fundamental frequency contour (50-8000Hz), harmonic structure (at least 3 harmonics), time modulation characteristics (amplitude modulation of 10-300Hz), and its start and end time and frequency range are recorded. By calculating the matching degree of these features with the typical bird call features, potential sound sources with a confidence level greater than 0.6 are preliminarily screened out.

[0039] S33: Determine the sound source azimuth of the potential sound source by analyzing the time difference of the sound signals received by at least two microphones.

[0040] In this embodiment, an array of at least two microphones arranged in space is used to locate the sound source. The time difference (TDOA) of the same sound signal arriving at different microphones is calculated, and the generalized cross-correlation function (GCC-PHAT) method is used to estimate the delay, with a time resolution of 0.1ms. According to the geometric configuration of the microphone array (minimum spacing 0.5m), the azimuth of the sound source is calculated by the spherical intersection algorithm, with a horizontal resolution of 2° and a vertical resolution of 5°. Within a distance of 3m, the positioning accuracy can reach ±0.1m. For stable sound sources with more than 10 consecutive frames, Kalman filtering is used for trajectory smoothing.

[0041] S34: Based on the similarity between the time-frequency spectrum and the preset bird sound template, and in combination with the sound source azimuth, the acoustic confidence is calculated and output.

[0042] In this embodiment, the detected sound source features are matched with a preset bird sound template library, which contains MFCC features (39 dimensions), rhythmic features, and spectral envelope features for each type of bird. The dynamic time warping algorithm is used to calculate the similarity between the test sample and the template, and the acoustic confidence is calculated in combination with the stability of the sound source azimuth (the angle change is less than 5° within 5 consecutive frames). The confidence calculation formula is: acoustic confidence = 0.7 × spectral similarity + 0.3 × azimuth stability, where the spectral similarity is normalized to the range of 0-1 by softmax. Finally, the azimuth (0-360°) and acoustic confidence (0-1) of each sound source are output, and the update frequency is 10Hz.

[0043] S4: Perform spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determine the region as a valid candidate region when the spatial distance between the two is less than a preset threshold.

[0044] Specifically, step S4 includes the following steps: S41: Mapping the spatial coordinates to a two-dimensional plane coordinate system of the video screen, and converting the sound source azimuth angle into projection coordinates in the two-dimensional plane coordinate system.

[0045] In the spatial consistency matching process, the spatial coordinates obtained by video detection are first converted from the pixel coordinate system to the world coordinate system. A three-dimensional coordinate system with the optical center of the camera as the origin is established. The two-dimensional image coordinates are converted to three-dimensional ground coordinates through the camera calibration parameters (including focal length, principal point coordinates and distortion coefficient) and the known installation height (such as 3 meters), where the Z-axis coordinate is fixed to 0 (assuming that the birds are active near the ground). The coordinate conversion is achieved using the perspective transformation matrix, and the conversion error is controlled within the range of ±0.1 meters.

[0046] The azimuth angle information of the sound source is converted into projection coordinates in the same world coordinate system. According to the installation position of the microphone array (maintaining a fixed distance of 1 meter from the camera) and the azimuth angle data (horizontal angle , pitch angle ), the three-dimensional coordinates of the sound source are calculated by the spherical coordinate conversion formula: ,in is the preset sound source distance estimate (the default is 3 meters). Taking into account the sound source positioning error, the coordinates are Gaussian smoothed ( is set to 0.2 meters), and finally the projection coordinates of the sound source are obtained.

[0047] S42: Calculate the Euclidean distance between the spatial coordinates and the projection coordinates, and determine that the spatial consistency match is when the Euclidean distance is less than or equal to a preset threshold.

[0048] In this embodiment, the Euclidean distance between the video detection coordinates and the sound source projection coordinates is calculated. Since the Z coordinate is fixed to 0, the preset threshold is set to 30% of the diagonal length of the video detection frame (typical value is 0.5-1.5 meters). When D is less than or equal to the preset threshold, it is determined to be a spatial consistency match. In order to improve the calculation efficiency, the KD tree data structure is used to perform a fast neighborhood search for the detection target, and the processing speed can reach 1000 matches / second.

[0049] S43: Confidence correction is performed on the successfully matched valid candidate regions, where: When the Euclidean distance is less than 50% of the preset threshold, enhancing the visual confidence and the acoustic confidence according to a preset confidence enhancement ratio; When the Euclidean distance is between 50% and 100% of the preset threshold, the confidence is dynamically attenuated according to a proportional relationship between the Euclidean distance and the preset threshold.

[0050] Wherein, the dynamic attenuation includes: First, a normalized ratio value of the Euclidean distance and a preset threshold is calculated and recorded as an attenuation coefficient.

[0051] Secondly, the visual confidence and acoustic confidence of the valid candidate area are respectively attenuated and calculated.

[0052] Finally, when the attenuated visual confidence and acoustic confidence are both lower than the preset confidence lower limit, the valid candidate area is eliminated.

[0053] S5: performing weighted calculation on the visual confidence and the acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds a determination threshold, outputting a recognition result indicating that there are birds in the target area.

[0054] Specifically, step S5 includes the following steps: S51: Calculating a dynamic weight allocation coefficient based on the ratio of the Euclidean distance to the preset threshold, wherein the weight coefficient of the visual confidence is linearly related to the inverse of the ratio.

[0055] In the fusion confidence calculation stage, weights are first dynamically assigned based on the spatial consistency matching results. For each valid candidate region, the weight ratio of visual confidence and acoustic confidence is determined based on the ratio of its Euclidean distance to the preset threshold (denoted as d / D). The weight coefficient of visual confidence is set to (1 d / D), the weight coefficient of acoustic confidence is d / D. When the distance is closer to the upper threshold, the weight of acoustic confidence is higher, and vice versa, visual confidence is dominant. The weight allocation process uses a linear interpolation algorithm to ensure a smooth transition of the weight coefficient within the range of distance from 0 to D.

[0056] S52: weighted fusion of the dynamic weight distribution coefficient, the visual confidence and the acoustic confidence after confidence correction, specifically expressed as: in, is the fusion confidence, is the Euclidean distance, is the preset threshold, and They are the corrected visual confidence and acoustic confidence, respectively.

[0057] In this embodiment, the dynamic weight is weighted and fused with the corrected visual confidence and acoustic confidence. In specific implementation, each candidate region is calculated according to the above formula. For multimodal data conflicts (such as V is high and A is very low), the system sets a conflict detection mechanism, and triggers a manual review flag when the difference between the two types of confidence exceeds 0.5.

[0058] S53: When the fusion confidence is greater than or equal to the determination threshold, it is determined that there is a bird target in the valid candidate area and the recognition result is output; otherwise, it is determined that there is no bird target in the valid candidate area.

[0059] In this embodiment, the dynamic adjustment of the judgment threshold is achieved through a pre-established parameter mapping table. The mapping table is indexed by video resolution and frame rate, and stores the optimal threshold calibrated by experiments. For example: the threshold is 0.65 in 4K@60fps mode, and the threshold is 0.75 in 1080p@30fps mode. The current video stream acquisition parameters are obtained in real time, and the corresponding threshold is quickly retrieved through the hash table. For uncovered parameter combinations, the nearest neighbor interpolation method is used to calculate the threshold. For example, the threshold of 3840×1600@45fps is the weighted average of the 4K@60fps and 1080p@30fps thresholds.

[0060] When the fusion confidence reaches or exceeds the search threshold, it is determined that there is a bird target in the area. The recognition result containing spatial coordinates, confidence value and timestamp is output, and the original data is recorded for subsequent model optimization. For valid candidate areas with fusion confidence lower than the judgment threshold, two-level filtering is performed: first, invalid areas with fusion confidence less than 0.3 are discarded, and the remaining areas are stored in the buffer for 5 seconds. If the confidence rises above the threshold during this period, it is reactivated. The final output result is encapsulated in JSON format, including structured data such as target ID, coordinate set and confidence curve.

[0061] Embodiment 2 The present invention provides a bird flock recognition system based on ultra-high-definition video, which is used to perform a bird flock recognition method based on ultra-high-definition video. Figure 2 As shown, the system comprises: The parameter dynamic adjustment module 100 is used to synchronously acquire the video stream and the audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectrum features extracted from the audio stream.

[0062] The video stream processing module 200 is used to extract spatiotemporal features of the dynamically adjusted video stream and output the spatial coordinates and visual confidence of the bird target.

[0063] The audio stream processing module 300 is used to extract time-frequency features and locate the sound source of the audio stream, and output the sound source azimuth and acoustic confidence.

[0064] The region determination module 400 is used to perform spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determine the region as a valid candidate region when the spatial distance between the two is less than a preset threshold.

[0065] The result output module 500 is used to perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds the judgment threshold, output the recognition result that there are birds in the target area.

[0066] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0067] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, the storage medium including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically-erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0068] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

Claims

1. A bird flock recognition method based on ultra-high-definition video, characterized in that: The method comprises the following steps: Synchronously acquiring a video stream and an audio stream of a target area, and dynamically adjusting acquisition parameters of the video stream based on spectral features extracted from the audio stream; Extracting spatiotemporal features from the dynamically adjusted video stream, and outputting spatial coordinates and visual confidence of the bird target; Extracting time-frequency features and locating sound sources from the audio stream, and outputting sound source azimuth and acoustic confidence; Performing spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determining the region as a valid candidate region when the spatial distance between the two is less than a preset threshold; A weighted calculation is performed on the visual confidence and the acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds a determination threshold, an identification result indicating that there is a bird in the target area is output.

2. The method for bird flock recognition based on ultra-high-definition video according to claim 1, characterized in that: The synchronously acquiring the video stream and the audio stream of the target area, and dynamically adjusting the acquisition parameters of the video stream based on the spectrum features extracted from the audio stream, includes: Extracting the energy distribution ratio of a preset high frequency band and a preset low frequency band in the audio stream; When the energy distribution ratio of the high frequency band exceeds a first threshold, increasing the resolution and frame rate of the video stream according to a preset proportionality coefficient; When the energy distribution ratio of the low frequency band exceeds a second threshold, reducing the resolution and frame rate of the video stream according to a preset proportionality coefficient; If the energy distribution ratios of the high frequency band and the low frequency band do not exceed the corresponding thresholds, the current acquisition parameters are maintained unchanged.

3. The method for bird flock recognition based on ultra-high-definition video according to claim 2, characterized in that: The step of extracting spatiotemporal features from the dynamically adjusted video stream and outputting spatial coordinates and visual confidence of the bird target includes: Performing spatiotemporal joint modeling on at least three adjacent frames of the video stream to extract fusion features including temporal motion features and spatial texture features; Based on the fused features, a candidate region including spatial coordinates and initial confidence is generated through a pre-trained target detection network; Non-maximum suppression processing is performed on the candidate area, and the spatial coordinates of the final bird target and the corresponding visual confidence are output.

4. The method for bird flock recognition based on ultra-high-definition video according to claim 2, characterized in that: The extracting time-frequency features and locating the sound source of the audio stream, and outputting the sound source azimuth and acoustic confidence, comprises: Performing short-time Fourier transform on the audio stream to obtain a time-frequency spectrum diagram; detecting bird call features on the time-frequency spectrum diagram and marking potential sound sources; Determine the sound source azimuth of the potential sound source by analyzing the time difference of the sound signals received by at least two microphones; Based on the similarity between the time-frequency spectrum diagram and a preset bird sound template, and in combination with the sound source azimuth, the acoustic confidence is calculated and output.

5. The method for bird flock recognition based on ultra-high definition video according to claim 4, characterized in that: The performing spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determining as a valid candidate area when the spatial distance between the two is less than a preset threshold, includes: Mapping the spatial coordinates to a two-dimensional plane coordinate system of a video screen, and converting the sound source azimuth angle into a projection coordinate in the two-dimensional plane coordinate system; Calculating the Euclidean distance between the spatial coordinates and the projection coordinates, and determining that the spatial consistency match is achieved when the Euclidean distance is less than or equal to a preset threshold; The confidence of the successfully matched valid candidate area is corrected, where: When the Euclidean distance is less than 50% of the preset threshold, enhancing the visual confidence and the acoustic confidence according to a preset confidence enhancement ratio; When the Euclidean distance is between 50% and 100% of the preset threshold, the confidence is dynamically attenuated according to a proportional relationship between the Euclidean distance and the preset threshold.

6. The method for bird flock recognition based on ultra-high definition video according to claim 5, characterized in that: The dynamic attenuation includes: Calculate the normalized ratio of the Euclidean distance to a preset threshold value, and record it as an attenuation coefficient; Performing attenuation calculation on the visual confidence and the acoustic confidence of the valid candidate area respectively; When the attenuated visual confidence and acoustic confidence are both lower than the preset confidence lower limit, the valid candidate area is eliminated.

7. The method for bird flock recognition based on ultra-high definition video according to claim 6, characterized in that: The visual confidence and acoustic confidence of the valid candidate area are weightedly calculated, and when the fusion confidence obtained by the weighted calculation exceeds the determination threshold, the recognition result that there is a bird in the target area is output, including: Calculating a dynamic weight allocation coefficient based on the ratio of the Euclidean distance to the preset threshold, wherein the weight coefficient of the visual confidence is linearly related to the inverse of the ratio; The dynamic weight distribution coefficient is weighted and fused with the visual confidence and acoustic confidence after confidence correction, which is specifically expressed as: in, is the fusion confidence, is the Euclidean distance, is the preset threshold, and They are the corrected visual confidence and acoustic confidence respectively; When the fusion confidence is greater than or equal to the determination threshold, it is determined that there is a bird target in the valid candidate area and the recognition result is output; otherwise, it is determined that there is no bird target in the valid candidate area.

8. The method for bird flock recognition based on ultra-high definition video according to claim 7, characterized in that: The determination threshold is dynamically adjusted according to the acquisition parameters of the video stream, specifically including: The determination threshold is determined by table lookup from a pre-established mapping relationship table between video stream acquisition parameters and determination thresholds, wherein the mapping relationship table contains optimal determination thresholds corresponding to different resolution and frame rate combinations.

9. A bird flock identification system based on ultra-high-definition video, used to execute the bird flock identification method based on ultra-high-definition video according to any one of claims 1 to 8, characterized in that: The system comprises: A parameter dynamic adjustment module, used to synchronously acquire the video stream and audio stream of the target area, and dynamically adjust the acquisition parameters of the video stream based on the spectrum features extracted from the audio stream; A video stream processing module is used to extract spatiotemporal features of the dynamically adjusted video stream and output the spatial coordinates and visual confidence of the bird target; An audio stream processing module, used for extracting time-frequency features and locating sound sources on the audio stream, and outputting sound source azimuth and acoustic confidence; A region determination module, configured to perform spatial consistency matching based on the spatial coordinates and the sound source azimuth, and determine the region as a valid candidate region when the spatial distance between the two is less than a preset threshold; The result output module is used to perform weighted calculation on the visual confidence and acoustic confidence of the valid candidate area, and when the fusion confidence obtained by the weighted calculation exceeds the judgment threshold, output the recognition result that there are birds in the target area.

Citation Information

Patent Citations

  • Audio events triggering video analytics

    CN110033787A

  • Video object detection and segmentation method based on space-time double-branch network

    CN110097568A

  • Monitoring method and device, electronic equipment and storage medium

    CN116684548A

  • Bird type identification method and identification device, and electronic equipment

    CN118861987A

  • Automobile environment sound enhancement method and device, electronic equipment and storage medium

    CN119724232A

Cited By

  • Bird active identification method, device and system and storage medium

    CN120976975A

  • Automatic wonderful event marking method based on sound source localization and visual fusion

    CN121633994A

  • Bird voiceprint and vision fusion real-time identification method for oil exploitation operation area

    CN121861392A