A video bird quantity recognition method based on segmented images

CN122842147APending Publication Date: 2026-09-29CHONGQING YINGKA ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610913257.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

虽然这提高了识别效率,但由于缺乏对时间信息的利用,无法避免同一只鸟在多帧中被重复计数的问题

Benefits of technology

[0091]本发明的有益效果是:本发明通过在视频处理流程中引入高效目标检测、分割提示生成、跨帧关联跟踪和数量统计分析等多环节优化,有效提升了在复杂环境和大规模场景下的鸟类数量识别准确性和鲁棒性。相较于传统人工统计,本发明有效降低了人力和时间成本;相较于传统单帧检测方法,本发明避免了重复计数问题,实现了对鸟类个体跨帧轨迹的自动化跟踪与高精度识别。最后,本发明通过结合空间分割和时间一致性分析,显著提高了数量统计的精度,满足生态监测和鸟类行为研究对长周期、大范围、自动化观测的需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842147A_ABST
    Figure CN122842147A_ABST
Patent Text Reader

Abstract

A video-based bird count recognition method based on segmented images includes: a bird video acquisition module acquiring continuous video data of a monitoring area and sampling it to obtain a set of frame images; a bird target detection module automatically detecting bird targets in the frame image set, outputting a list of detection boxes, and calculating the center point coordinates of the detection boxes; an image segmentation module using the center point coordinates as cue points to perform instance-level segmentation of the frame images, while simultaneously performing cross-frame consistency tracking to obtain segmentation masks and ID tracking results; a target association and trajectory generation module, based on the segmentation mask and ID tracking results, performing bird individual matching and trajectory generation between multiple frame images, establishing associations for the same bird individual across different frame images, and outputting cross-frame trajectory information to a bird count statistics module; and a bird count statistics module statistically outputting the number of all bird individuals and their movement trajectories for each time period. The effect is to improve the efficiency and accuracy of bird count statistics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bird count recognition technology, and in particular to a video bird count recognition method based on segmented images. Background Technology

[0002] Existing methods for bird population identification are mainly based on static image recognition, manual observation and statistics, or simple video detection techniques. Traditional manual observation and annotation methods rely on experts to identify and count birds frame by frame in images or videos. This approach is time-consuming and labor-intensive, easily affected by subjective human factors, and extremely inefficient when facing large-scale monitoring tasks, making it unsuitable for long-term automated monitoring needs.

[0003] Another common approach is bird counting based on single-frame target detection. These methods typically utilize convolutional neural networks (CNNs) or lightweight detectors to independently identify and count birds in each frame. While this improves efficiency, the lack of temporal information means the same bird can be counted repeatedly across multiple frames. Other methods perform simple inter-frame deduplication or average counting on the video frame sequence, but this coarse processing lacks true cross-frame individual correlation modeling, making it difficult to accurately distinguish the trajectories of different birds during long-term tracking. Furthermore, existing segmentation-based image analysis often focuses on single-frame segmentation, lacking modeling of segmentation consistency over time, resulting in inconsistent segmentation results within the video and further impacting the accuracy of the counting.

[0004] The limitations of these traditional methods include: inefficiency due to reliance on manual statistics; high risk of repeated counting and difficulty in handling complex scenarios based on single-frame detection methods; lack of cross-frame consistency in segmentation results; and lack of in-depth utilization of temporal information, making it difficult to achieve true individual-level trajectory tracking and quantity statistics.

[0005] Disadvantages of existing technologies: Traditional bird count identification methods suffer from low statistical efficiency, high risk of duplicate counting, and difficulty in achieving true individual-level trajectory tracking and count. Summary of the Invention

[0006] The present invention provides a video bird count identification method based on segmented images, which can improve the efficiency and accuracy of bird count statistics.

[0007] To achieve the above objectives, the present invention provides a video bird count recognition method based on segmented images, comprising the following steps:

[0008] Step 1: Construct a video bird count recognition system based on segmented images. This video bird count recognition system is equipped with a bird video acquisition module, a bird target detection module, an image segmentation module, a target association and trajectory generation module, and a bird count statistics module connected in sequence.

[0009] Step 2: The bird video acquisition module acquires continuous video data in the monitoring area in real time, performs frame sampling on the continuous video data in the time dimension, extracts the frame image set a, and sends it to the bird target detection module and the image segmentation module;

[0010] Step 3: The bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, outputs a high-precision detection box list b, calculates the center point coordinates of each detection box, and then passes the center point coordinates, bird category label and confidence score of each detection box to the image segmentation module.

[0011] Step 4: The image segmentation module uses the center point coordinates of the detection box as prompt points, calls the image segmentation model, performs instance-level segmentation in the corresponding frame image, and performs cross-frame consistency tracking in the frame image set a. Through joint modeling, it ensures the continuity of the segmentation results in the time dimension; and outputs the segmentation mask and ID tracking results to the target association and trajectory generation module.

[0012] Step 5: The target association and trajectory generation module performs bird individual matching and trajectory generation between multiple frame images based on the segmentation mask and ID tracking results, establishes the association of the same bird individual between different frame images to avoid duplicate counting, and finally outputs cross-frame trajectory information to the bird count module.

[0013] Step 6: The bird population statistics module performs bird population statistics based on the unique identifier trajectory of the same bird individual in the cross-frame trajectory information, and finally outputs the total number of birds and their movement trajectory data for each time period. The final output results can be used to support scientific research and bird population management needs.

[0014] Through the above design, this invention fully utilizes the high-resolution segmentation and video tracking capabilities of efficient target detection and image segmentation models to achieve deep integration of spatial detection and temporal tracking, greatly improving the accuracy and stability of bird identification and population statistics. This will provide more precise and efficient technical support for ecological monitoring, bird behavior research, and related fields. Simultaneously, the end-to-end intelligent detection also significantly improves the efficiency of bird identification and population statistics while reducing labor and time costs.

[0015] Preferably, the bird video acquisition module is a multi-view video acquisition device arrayed in the monitoring area, and the video sampling method of the multi-view video acquisition device is one or more combinations of frame rate adaptive sampling, motion sensing sampling and uniform time sampling.

[0016] To reduce computational load, improve processing efficiency, and ensure that key bird behavioral features are not missed, this invention performs frame sampling processing on the original video. This involves extracting keyframe images from the original continuous frames according to a predefined strategy, forming an image sequence for subsequent recognition. The sampling method includes the following three strategies, which can be flexibly selected or combined according to the application scenario:

[0017] The frame rate adaptive sampling dynamically sets the sampling interval based on the frame rate of the input video, ensuring that the sampling interval remains consistent across the time dimension. For example, for a 30 fps video, one frame is sampled every 5 frames, i.e., every 0.2 seconds; for a 60 fps video, one frame is sampled every 10 frames. This strategy achieves a uniform temporal resolution across videos with different frame rates, facilitating subsequent cross-video processing and trajectory comparison.

[0018] Motion-aware sampling employs optical flow or inter-frame difference calculation techniques to evaluate the amount of image change between adjacent frames. When significant motion is detected in a local area, such as a bird suddenly taking off or flying quickly, the sampling density is automatically increased to ensure the capture of high-speed dynamic behavior. When there is little motion or the background is static, the sampling frequency is reduced to save resources. This strategy is suitable for scenarios where bird activity is irregular or sudden.

[0019] Uniform temporal sampling involves evenly distributing sampling points along the timeline of the entire video segment and extracting image frames at fixed time intervals, such as every 0.5 seconds, ensuring full coverage of the video and avoiding missed detections. This strategy is suitable for analyzing periodically occurring bird activities, ensuring the balance and representativeness of overall observations.

[0020] The bird video acquisition module can acquire videos in common formats such as MP4, AVI, and MOV, with a resolution of no less than 720p and a frame rate range of 25fps to 60fps, ensuring that the video has good image quality and continuity, and is suitable for subsequent image segmentation and target detection tasks.

[0021] After frame sampling is completed, all frame images are uniformly converted to standard formats such as JPEG and PNG, their original timestamps are preserved, and they are numbered and stored in chronological order to provide standardized input data for subsequent object detection, image segmentation, and cross-frame trajectory tracking.

[0022] Preferably, in step 3, the bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, including the following steps:

[0023] Step 3.1: The target detection model performs uniform scaling and adjustment on each frame image in the frame image set a, and then performs bird target detection. The detection outputs a list of candidate detection boxes for one or more bird targets in each frame image. The information of each detection box includes the bird category label, the detection box position coordinates and the confidence score.

[0024] This invention improves detection efficiency by uniformly scaling and adjusting the size of each frame of image to ensure consistent network input size.

[0025] Among them, the bird category label is the bird species or the unified bird identifier, the detection box position coordinates are the pixel values ​​of the upper left and lower right corners of the detection box, and the confidence score is the probability that the target corresponding to the current detection box exists.

[0026] Step 3.2: Subsequently, the bird target detection module performs standardization, low-confidence filtering, and deduplication / overlapping processing on all candidate detection box lists in sequence to obtain the final detection box list b;

[0027] Step 3.3: The bird target detection module uses a center point calculation method to calculate the center point coordinates of each detection box in the detection box list b. The expression is:

[0028] ;

[0029] ;

[0030] in, The coordinates of the top left corner of the detection box. The coordinates of the bottom right corner of the detection box. Here are the coordinates of the center point, and i is the index of the detection box in the current frame;

[0031] The above method for calculating center point coordinates is simple and efficient, and can accurately locate the center position of each detected target in pixel space. The coordinates can be retained as the original pixel position, or they can be normalized to a ratio of the image width and height, ranging from 0 to 1, to adapt to the input requirements of subsequent segmentation models.

[0032] Step 3.4: The bird target detection module outputs a set of center point cueings containing rich attributes. For each detected bird target, i.e., the detection box, the output information includes the center point coordinates, bird category label, and confidence score.

[0033] In order to efficiently convert the target detection results into prompts for subsequent image segmentation modules such as SAM2, this invention calculates the center point of the detection box and accurately converts each detection box into a prompt point in spatial coordinates.

[0034] Preferably, in step 4, the image segmentation module performs instance-level segmentation and cross-frame consistency tracking on each frame image, as follows:

[0035] The image segmentation module receives the current frame image and the center point coordinates of the corresponding detection box, then uses the center point coordinates as prompts to generate K sets of candidate mask segmentation results, and then selects the optimal segmentation mask based on a preset selection strategy.

[0036] Meanwhile, the image segmentation module associates the optimal segmentation mask with the target IDs assigned in historical frames, achieving unique identification and cross-frame consistency of the segmentation results, providing accurate single-frame segmentation basis and ID annotation information for subsequent multi-frame tracking and statistical analysis.

[0037] Preferably, the image segmentation module generates the following expression for the K groups of candidate mask segmentation results:

[0038] ;

[0039] in, For the k-th candidate segmentation mask, Let K be the confidence score of the k-th candidate segmentation mask, K be the total number of candidate segmentation masks, and k be the index of the candidate segmentation mask, where k∈[1,K]. For image segmentation models, For the current frame image, This is the set of prompts used for conditional splitting. For time, For the set of candidate segmentation masks;

[0040] The preset selection strategy is either the highest confidence score strategy or the maximum intersection-union ratio (IoU) strategy with the detection box;

[0041] The highest confidence score strategy is: the confidence score of all candidate segmentation masks. Sort the masks and select the one with the highest confidence as the optimal segmentation mask. Its expression is:

[0042] ;

[0043] in, For optimal segmentation mask, This is a function used to return the index of the maximum value in an array;

[0044] The highest confidence score strategy focuses on the model's confidence in the segmentation results and does not require additional spatial prior information.

[0045] The maximum intersection-union (IoU) strategy for detection boxes is as follows: First, calculate the IoU between each candidate mask and the input detection box. The intersection-union ratio (IoU) is expressed as:

[0046] ;

[0047] Then, the candidate mask with the highest IoU is selected as the optimal segmentation mask. The expression is:

[0048] ;

[0049] in, For the k-th candidate segmentation mask, This is the detection frame.

[0050] The maximum intersection-union (IoU) strategy of the detection boxes utilizes the spatial position constraints of the detection boxes to ensure maximum consistency between the mask segmentation results and the detection positions.

[0051] Preferably, the image segmentation module associates the optimal segmentation mask with the target IDs already assigned in historical frames, as follows:

[0052] First, define the set of allocated masks for historical frames and their IDs as follows: Then calculate the mask for the current frame. With historical frame mask Similarity measurement The expression is:

[0053] ;

[0054] in, Represents the feature similarity function, , These are the weighting coefficients. for and The intersection-union ratio, where j is the index of the detection box in the historical frame;

[0055] Finally, IDs are assigned based on similarity, expressed as:

[0056] ;

[0057] in, This is the similarity threshold.

[0058] Preferably, in step 5, the target association and trajectory generation module generates cross-frame trajectory information as follows:

[0059] Step 5.1: Feature Extraction: The target association and trajectory generation module obtains the segmentation mask of the current frame image, as well as the segmentation masks cached in historical frame images and their assigned ID information. This input provides contextual information for cross-frame matching, ensuring temporal continuity modeling. Then, the temporal tracking module built into the segmentation model extracts cross-frame continuous features from the mask region. These features include appearance features, motion features, and mask morphology features. The feature extraction expression is:

[0060] ;

[0061] in, Let be the temporal feature vector of the i-th mask in the current frame image. This provides the mask pixel information for the i-th target in the current frame. For local motion vector fields, For feature encoder;

[0062] The definition is: using the optimal segmentation mask of the current frame. For the original sampled frame image After spatial cropping, a colored local image patch retains the bird's texture and color. The formula can be expressed as follows: ,in This indicates element-wise multiplication.

[0063] Appearance features: The static information of birds, such as shape, color, and texture, is encoded by a convolutional neural network (CNN) to generate a high-dimensional feature vector.

[0064] Motion characteristics: Calculate the displacement vector (Δx, Δy), motion direction (angle), and motion speed (pixels / second) of the mask center point between adjacent frames. Optionally, this can be combined with the local motion vector field estimated by optical flow. .

[0065] Mask morphological features may include the aspect ratio of the circumscribed rectangle, the rate of change of area, and the similarity of the outline, which are used to further distinguish individuals with similar morphologies.

[0066] Step 5.2: Matching strategy: For each mask in the current frame, first calculate its overall similarity with the unmatched masks in the historical frame buffer, then select the historical ID with the highest overall similarity to associate with it. If the highest overall similarity exceeds the set threshold τ, then inherit the corresponding historical ID; if no historical mask has an overall similarity exceeding the threshold, then assign a new global ID to the current mask.

[0067] Calculate the combined similarity between the masks of historical frames and the current frame. The expression is:

[0068] Appearance similarity: ;

[0069] Motion similarity: ;

[0070] Morphological similarity: ;

[0071] Overall similarity: ;

[0072] in, , , These are the weight coefficients corresponding to the similarity features. ; Let be the temporal feature vector of the j-th mask in the tk-th frame of history. The cosine similarity function is used. For the current frame number The local motion vector field corresponding to each target History The first frame The local motion vector field corresponding to each target For the current frame number Binary segmentation mask corresponding to each target, For the first time in history The first frame Binary segmentation mask corresponding to each target, for The L2 norm;

[0073] Step 5.3: ID Management and Track Update:

[0074] Added ID management: Assign a unique global ID to newly appearing unmatched targets and initialize their trajectory records;

[0075] ID retention and deletion: For IDs that are not matched for N consecutive frames, such as 3 to 5 frames, they are marked as "disappeared" and are still retained in the cache; when the "disappeared" state exceeds a preset time threshold, such as 10 frames, they are deleted from the cache to avoid redundant calculations.

[0076] Track update: For each global ID, record its time series information, including timestamps. Center coordinates and area That is, ID:[( Optionally, the trajectory points can be smoothed using a Kalman filter to reduce noise-induced jumps.

[0077] Step 5.4: Output results: The target association and trajectory generation module outputs cross-frame trajectory information to the bird count module; the cross-frame trajectory information includes the current frame segmentation mask with a global ID tag and its attribute information, as well as the updated trajectory library, which includes the complete motion path and time information of all active targets.

[0078] Through the above design, the unique identifier and continuous tracking of the motion trajectory of the same bird individual in a multi-frame video sequence are achieved, avoiding duplicate counting and improving statistical accuracy. The target association and trajectory generation module establishes a correlation between mask segmentation results in the time dimension, enabling consistent modeling and trajectory recording of the same target in consecutive frames.

[0079] Preferably, the bird count module treats each unique identifier trajectory in the cross-frame trajectory information as a unique bird individual, as follows:

[0080] First, the bird count module receives the trajectory data set output after cross-frame processing, the expression of which is:

[0081] ;

[0082] Where T represents the set of all bird trajectories that have undergone cross-frame association processing and have been assigned globally unique IDs. Let m be the trajectory of the m-th individual. The total number of individual trajectories;

[0083] Since the aforementioned cross-frame association and ID management have linked the detection results of the same bird individual across multiple frames into a complete trajectory, each trajectory is regarded as a unique bird individual.

[0084] Within a set time window or global monitoring range, the number of different IDs in the trajectory set is counted to obtain the corresponding bird individual quantity information. The statistical formula is:

[0085] ;

[0086] That is, the total number of all valid tracks is counted and considered as the final number of individual birds.

[0087] When segmented by time window, the bird count module can further filter the movement trajectory times appearing within a set time range based on the timestamp information recorded in the trajectory, as expressed by:

[0088] ;

[0089] in, This is used to set a time window to count the number of unique trajectory IDs that appear within that time period.

[0090] Preferably, the target detection model uses the YOLO target detection model, and the image segmentation model uses the SAM2 image segmentation model. By combining the YOLO target detection model and the SAM2 segmentation and tracking model, an organic fusion of spatial detection and temporal correlation is achieved.

[0091] The beneficial effects of this invention are as follows: By introducing multiple optimizations into the video processing workflow, including efficient target detection, segmentation cue generation, cross-frame correlation tracking, and statistical analysis, this invention effectively improves the accuracy and robustness of bird population identification in complex environments and large-scale scenes. Compared to traditional manual statistics, this invention effectively reduces labor and time costs. Compared to traditional single-frame detection methods, this invention avoids the problem of repeated counting, achieving automated tracking and high-precision identification of individual bird trajectories across frames. Finally, by combining spatial segmentation and temporal consistency analysis, this invention significantly improves the accuracy of population statistics, meeting the needs of ecological monitoring and bird behavior research for long-term, large-scale, and automated observation. Attached Figure Description

[0092] Figure 1 This is a flowchart of the method of the present invention;

[0093] Figure 2 This is a flowchart of image segmentation based on cue points in the embodiment;

[0094] Figure 3 This is a flowchart of cross-frame target association and trajectory generation in the embodiment. Detailed Implementation

[0095] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following embodiments or drawings are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0096] like Figure 1 As shown, a method for identifying the number of birds in a video based on segmented images includes the following steps:

[0097] Step 1: Construct a video bird count recognition system based on segmented images. This video bird count recognition system is equipped with a bird video acquisition module, a bird target detection module, an image segmentation module, a target association and trajectory generation module, and a bird count statistics module connected in sequence.

[0098] Step 2: The bird video acquisition module acquires continuous video data in the monitoring area in real time, performs frame sampling on the continuous video data in the time dimension, extracts the frame image set a, and sends it to the bird target detection module and the image segmentation module;

[0099] Step 3: The bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, outputs a high-precision detection box list b, calculates the center point coordinates of each detection box, and then passes the center point coordinates, bird category label and confidence score of each detection box to the image segmentation module.

[0100] Step 4: The image segmentation module uses the center point coordinates of the detection box as prompt points, calls the image segmentation model, performs instance-level segmentation in the corresponding frame image, and performs cross-frame consistency tracking in the frame image set a. Through joint modeling, it ensures the continuity of the segmentation results in the time dimension; and outputs the segmentation mask and ID tracking results to the target association and trajectory generation module.

[0101] Step 5: The target association and trajectory generation module performs bird individual matching and trajectory generation between multiple frame images based on the segmentation mask and ID tracking results, establishes the association of the same bird individual between different frame images, and outputs cross-frame trajectory information to the bird count module.

[0102] Step 6: The bird count module performs bird count based on the unique identifier trajectory of the same bird individual in the cross-frame trajectory information, and finally outputs the total number of birds and their movement trajectory data for each time period.

[0103] This invention proposes a video bird count identification method based on segmented images. It combines object detection models such as YOLO with video segmentation and tracking models such as SAM2, leveraging the complementary advantages of detection and segmentation to achieve accurate segmentation and cross-frame correlation tracking of individual birds in videos. By establishing consistent segmentation and individual-level trajectory information in the temporal dimension, this invention effectively avoids the problem of double-counting and improves the accuracy, efficiency, and robustness of identification in complex environments. This invention provides efficient and intelligent technical support for fields such as ecological monitoring, bird behavior research, and wildlife conservation, meeting the practical needs of large-scale, long-term, and automated monitoring.

[0104] The bird video acquisition module is a multi-view video acquisition device arrayed in the monitoring area. The video sampling method of the multi-view video acquisition device is one or more combinations of frame rate adaptive sampling, motion sensing sampling, and uniform time sampling.

[0105] To reduce computational load, improve processing efficiency, and ensure that key bird behavioral features are not missed, this invention performs frame sampling processing on the original video. This involves extracting keyframe images from the original continuous frames according to a predefined strategy, forming an image sequence for subsequent recognition. The sampling method includes the following three strategies, which can be flexibly selected or combined according to the application scenario:

[0106] The frame rate adaptive sampling dynamically sets the sampling interval based on the frame rate of the input video, ensuring that the sampling interval remains consistent across the time dimension. For example, for a 30 fps video, one frame is sampled every 5 frames, i.e., every 0.2 seconds; for a 60 fps video, one frame is sampled every 10 frames. This strategy achieves a uniform temporal resolution across videos with different frame rates, facilitating subsequent cross-video processing and trajectory comparison.

[0107] Motion-aware sampling employs optical flow or inter-frame difference calculation techniques to evaluate the amount of image change between adjacent frames. When significant motion is detected in a local area, such as a bird suddenly taking off or flying quickly, the sampling density is automatically increased to ensure the capture of high-speed dynamic behavior. When there is little motion or the background is static, the sampling frequency is reduced to save resources. This strategy is suitable for scenarios where bird activity is irregular or sudden.

[0108] Uniform temporal sampling involves evenly distributing sampling points along the timeline of the entire video segment and extracting image frames at fixed time intervals, such as every 0.5 seconds, ensuring full coverage of the video and avoiding missed detections. This strategy is suitable for analyzing periodically occurring bird activities, ensuring the balance and representativeness of overall observations.

[0109] The bird video acquisition module can acquire videos in common formats such as MP4, AVI, and MOV, with a resolution of no less than 720p and a frame rate range of 25fps to 60fps, ensuring that the video has good image quality and continuity, and is suitable for subsequent image segmentation and target detection tasks.

[0110] After frame sampling is completed, all frame images are uniformly converted to standard formats such as JPEG and PNG, their original timestamps are preserved, and they are numbered and stored in chronological order to provide standardized input data for subsequent object detection, image segmentation, and cross-frame trajectory tracking.

[0111] In step 3, the bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, including the following steps:

[0112] Step 3.1: The target detection model performs uniform scaling and adjustment on each frame image in the frame image set a, and then performs bird target detection. The detection outputs a list of candidate detection boxes for one or more bird targets in each frame image. The information of each detection box includes the bird category label, the detection box position coordinates and the confidence score.

[0113] Among them, the bird category label is the bird species or the unified bird identifier, the detection box position coordinates are the pixel values ​​of the upper left and lower right corners of the detection box, and the confidence score is the probability that the target corresponding to the current detection box exists.

[0114] Step 3.2: Subsequently, the bird target detection module performs standardization, low-confidence filtering, and deduplication / overlapping processing on all candidate detection box lists in sequence to obtain the final detection box list b;

[0115] Step 3.3: The bird target detection module uses a center point calculation method to calculate the center point coordinates of each detection box in the detection box list b. The expression is:

[0116] ;

[0117] ;

[0118] in, The coordinates of the top left corner of the detection box. The coordinates of the bottom right corner of the detection box. Here are the coordinates of the center point, and i is the index of the detection box in the current frame;

[0119] Step 3.4: The bird target detection module outputs the center point coordinates, bird category label, and confidence score of each detection box to the image segmentation module.

[0120] like Figure 2 As shown, in step 4, the image segmentation module performs instance-level segmentation and cross-frame consistency tracking on each frame image, as follows:

[0121] The image segmentation module receives the current frame image and the center point coordinates of the corresponding detection box, then uses the center point coordinates as prompts to generate K sets of candidate mask segmentation results, and then selects the optimal segmentation mask based on a preset selection strategy.

[0122] Meanwhile, the image segmentation module associates the optimal segmentation mask with the target IDs assigned in historical frames, achieving unique identification and cross-frame consistency of the segmentation results, providing accurate single-frame segmentation basis and ID annotation information for subsequent multi-frame tracking and statistical analysis.

[0123] The image segmentation module generates the following expressions for K groups of candidate mask segmentation results:

[0124] ;

[0125] in, For the k-th candidate segmentation mask, Let K be the confidence score of the k-th candidate segmentation mask, K be the total number of candidate segmentation masks, and k be the index of the candidate segmentation mask, where k∈[1,K]. For image segmentation models, For the current frame image, This is the set of prompts used for conditional splitting. For time, For the set of candidate segmentation masks;

[0126] The preset selection strategy is either the highest confidence score strategy or the maximum intersection-union ratio (IoU) strategy with the detection box;

[0127] The highest confidence score strategy is: the confidence score of all candidate segmentation masks. Sort the masks and select the one with the highest confidence as the optimal segmentation mask. Its expression is:

[0128] ;

[0129] in, For optimal segmentation mask, This is a function used to return the index of the maximum value in an array;

[0130] The highest confidence score strategy focuses on the model's confidence in the segmentation results and does not require additional spatial prior information.

[0131] The maximum intersection-union (IoU) strategy for the detection boxes is as follows: First, calculate the IoU between each candidate mask and the input detection box. The intersection-union ratio (IoU) is expressed as:

[0132] ;

[0133] Then, the candidate mask with the highest IoU is selected as the optimal segmentation mask. The expression is:

[0134] ;

[0135] in, For the k-th candidate segmentation mask, This is the detection frame.

[0136] The maximum intersection-union (IoU) strategy of the detection boxes utilizes the spatial position constraints of the detection boxes to ensure maximum consistency between the segmentation results and the detection positions.

[0137] The image segmentation module associates the optimal segmentation mask with the target IDs already assigned in historical frames, as follows:

[0138] First, define the set of allocated masks for historical frames and their IDs as follows: Then calculate the mask for the current frame. With historical frame mask Similarity measurement The expression is:

[0139] ;

[0140] in, Represents the feature similarity function, , These are the weighting coefficients. for and The intersection-union ratio, where j is the index of the detection box in the historical frame;

[0141] Finally, IDs are assigned based on similarity, expressed as:

[0142] ;

[0143] in, This is the similarity threshold.

[0144] like Figure 3 As shown, in step 5, the steps by which the target association and trajectory generation module generates cross-frame trajectory information are as follows:

[0145] Step 5.1: Feature Extraction: The target association and trajectory generation module obtains the segmentation mask of the current frame image, as well as the segmentation masks cached in historical frame images and their assigned ID information. Then, it uses the temporal tracking module built into the segmentation model to extract continuous cross-frame features from the mask region. The features include appearance features, motion features, and mask morphology features. The feature extraction expression is:

[0146] ;

[0147] in, Let be the temporal feature vector of the i-th mask in the current frame image. This provides the mask pixel information for the i-th target in the current frame. For local motion vector fields, For feature encoder;

[0148] Appearance features: The static information of birds, such as shape, color, and texture, is encoded by a convolutional neural network (CNN) to generate a high-dimensional feature vector.

[0149] Motion characteristics: Calculate the displacement vector (Δx, Δy), motion direction (angle), and motion speed (pixels / second) of the mask center point between adjacent frames. Optionally, this can be combined with the local motion vector field estimated by optical flow. .

[0150] Mask morphological features may include the aspect ratio of the circumscribed rectangle, the rate of change of area, and the similarity of the outline, which are used to further distinguish individuals with similar morphologies.

[0151] Step 5.2: Matching strategy: For each mask in the current frame, first calculate its overall similarity with the unmatched masks in the historical frame buffer, then select the historical ID with the highest overall similarity to associate with it. If the highest overall similarity exceeds the set threshold τ, then inherit the corresponding historical ID; if no historical mask has an overall similarity exceeding the threshold, then assign a new global ID to the current mask.

[0152] Calculate the combined similarity between the masks of historical frames and the current frame. The expression is:

[0153] Appearance similarity: ;

[0154] Motion similarity: ;

[0155] Morphological similarity: ;

[0156] Overall similarity: ;

[0157] in, , , These are the weight coefficients corresponding to the similarity features. ; Let be the temporal feature vector of the j-th mask in the tk-th frame of history. The cosine similarity function is used. For the current frame number The local motion vector field corresponding to each target History The first frame The local motion vector field corresponding to each target For the current frame number Binary segmentation mask corresponding to each target, For the first time in history The first frame Binary segmentation mask corresponding to each target, for The L2 norm;

[0158] Step 5.3: ID Management and Track Update:

[0159] Added ID management: Assign a unique global ID to newly appearing unmatched targets and initialize their trajectory records;

[0160] ID retention and deletion: For IDs that are not matched for N consecutive frames, such as 3 to 5 frames, they are marked as "disappeared" and are still retained in the cache; when the "disappeared" state exceeds a preset time threshold, such as 10 frames, they are deleted from the cache to avoid redundant calculations.

[0161] Track update: For each global ID, record its time series information, including timestamps. Center coordinates and area That is, ID:[( Optionally, the trajectory points can be smoothed using a Kalman filter to reduce noise-induced jumps.

[0162] Step 5.4: Output results: The target association and trajectory generation module outputs cross-frame trajectory information to the bird count module; the cross-frame trajectory information includes the current frame segmentation mask with a global ID tag and its attribute information, as well as the updated trajectory library, which includes the complete motion path and time information of all active targets.

[0163] The bird count module treats each unique identifier trajectory in the cross-frame trajectory information as a unique bird individual, as follows:

[0164] First, the bird count module receives the trajectory data set output after cross-frame processing, the expression of which is:

[0165] ;

[0166] Where T represents the set of all bird trajectories that have undergone cross-frame association processing and have been assigned globally unique IDs. Let m be the trajectory of the m-th individual. The total number of individual trajectories;

[0167] Within a set time window or global monitoring range, the number of different IDs in the trajectory set is counted to obtain the corresponding bird individual quantity information. The statistical formula is:

[0168] ;

[0169] That is, the total number of all valid tracks is counted and considered as the final number of individual birds.

[0170] When segmented by time window, the bird count module can further filter the movement trajectory times appearing within a set time range based on the timestamp information recorded in the trajectory, as expressed by:

[0171] ;

[0172] in, This is used to set a time window to count the number of unique trajectory IDs that appear within that time period.

[0173] The target detection model uses the YOLO target detection model, and the image segmentation model uses the SAM2 image segmentation model.

[0174] This embodiment uses publicly available target detection models such as YOLO pre-trained weights, combined with labeled bird target data, for fine-tuning training to adapt to the morphology, posture, and environmental characteristics of birds in specific monitoring scenarios. The training data covers diverse bird species and different shooting conditions, improving the model's generalization ability and robustness.

[0175] The object detection model employs an end-to-end convolutional neural network architecture. The input image undergoes multi-layer feature extraction, fusion, and multi-scale prediction, outputting a list of candidate detection boxes. Redundant boxes are removed using confidence thresholding and non-maximum suppression (NMS), preserving high-quality detection results. Finally, the detection box information is standardized, low-confidence results are filtered, and duplicate and overlapping results are removed, outputting a high-precision list of detection boxes, providing input for subsequent steps such as center point calculation and segmentation cue point generation.

[0176] This invention acquires continuous video within a monitored area using a bird video acquisition module. Then, it employs strategies such as frame rate adaptation, motion perception, and uniform temporal sampling to extract keyframes. This reduces computational load while ensuring no key bird behavioral features are missed, providing high-quality input for subsequent processing and improving detection efficiency. Finally, it uses a target detection model such as YOLO to detect birds on the sampled frames, outputting high-precision bounding boxes containing location, confidence level, and category labels. Combined with the calculation of the bounding box center point, the spatial location information is transformed into cue point input for the SAM2 model, achieving precise integration of target detection and segmentation. Leveraging the instance-level segmentation and cross-frame tracking capabilities of the SAM2 model, the highest confidence score or the largest bounding box is used for target detection. IoU is used as a selection strategy to screen the optimal segmentation mask. Cross-frame target ID association is achieved through similarity calculation of appearance features, motion features, and mask morphology features, ensuring the continuity and uniqueness of segmentation results in the temporal dimension. Based on the cross-frame association results, a bird trajectory dataset with global IDs is generated, recording information such as individual motion coordinates, timestamps, and mask areas. Combined with trajectory updates and ID lifecycle management, automated bird count statistics and individual behavior analysis are achieved, overcoming the limitations of traditional manual statistics which are inefficient and prone to duplicate counting in single-frame detection. Finally, bird count and motion trajectory data for each time period are output, providing high-precision and intelligent technical support for fields such as ecological monitoring, bird behavior research, and wildlife conservation.

[0177] Compared to traditional bird population counting methods, this invention has the following advantages:

[0178] (1) Multi-strategy frame sampling optimization: Compared with traditional video processing methods, this method uses a combination of three strategies: frame rate adaptive sampling, motion-aware sampling, and uniform time sampling to extract key frames from continuous videos. By dynamically adjusting the sampling density, the computational load is reduced while ensuring that key features of bird activities are not missed. This solves the contradiction between the redundancy of the original video data and the capture of key information, and provides high-quality standardized input for subsequent processing.

[0179] (2) Deep fusion of detection and segmentation: This invention proposes a segmentation cueing mechanism based on the center point of the detection box, which transforms the high-precision detection box output by the target detection into spatial cue points of the SAM2 model. The optimal segmentation mask is selected by the highest confidence score or the largest detection box IoU strategy, so as to achieve accurate connection between target detection and instance segmentation, solve the problem of matching deviation between single frame segmentation results and target position, and improve segmentation accuracy.

[0180] (3) Cross-frame target association and trajectory tracking: By utilizing the cross-frame consistency tracking capability of the SAM2 model and combining multi-dimensional similarity calculation of appearance features, motion features and mask morphology features, cross-frame ID association of individual birds is achieved. Through ID lifecycle management such as adding, retaining, deleting and smooth trajectory updates, a continuous trajectory dataset with global IDs is generated, which effectively avoids duplicate counting and solves the technical problem of individual recognition and trajectory coherence across multiple frames.

[0181] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying the number of birds in a video based on segmented images, characterized in that, Includes the following steps: Step 1: Construct a video bird count recognition system based on segmented images. This video bird count recognition system is equipped with a bird video acquisition module, a bird target detection module, an image segmentation module, a target association and trajectory generation module, and a bird count statistics module connected in sequence. Step 2: The bird video acquisition module acquires continuous video data in the monitoring area in real time, performs frame sampling on the continuous video data in the time dimension, extracts the frame image set a, and sends it to the bird target detection module and the image segmentation module; Step 3: The bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, outputs a list of detection boxes b, calculates the center point coordinates of each detection box, and then passes the center point coordinates, bird category label and confidence score of each detection box to the image segmentation module. Step 4: The image segmentation module uses the center point coordinates of the detection box as prompt points, calls the image segmentation model, performs instance-level segmentation in the corresponding frame image, performs cross-frame consistency tracking in the frame image set a, and outputs the segmentation mask and ID tracking results to the target association and trajectory generation module. Step 5: Based on the segmentation mask and ID tracking results, the target association and trajectory generation module performs bird individual matching and trajectory generation between multiple frame images, establishes the association between the same bird individual in the current frame and historical frame images, and outputs cross-frame trajectory information to the bird count module. Step 6: The bird count module performs bird count based on the unique identifier trajectory of the same bird individual in the cross-frame trajectory information, and finally outputs the total number of birds and their movement trajectory data for each time period.

2. The video bird count recognition method based on segmented images according to claim 1, characterized in that: The bird video acquisition module is a multi-view video acquisition device arrayed in the monitoring area. The video sampling method of the multi-view video acquisition device is one or more combinations of frame rate adaptive sampling, motion sensing sampling, and uniform time sampling.

3. The video bird count recognition method based on segmented images according to claim 1, characterized in that: In step 3, the bird target detection module uses a target detection model to automatically detect bird targets in the frame image set a, including the following steps: Step 3.1: The target detection model performs uniform scaling and adjustment on each frame image in the frame image set a, and then performs bird target detection. The detection outputs a list of candidate detection boxes for one or more bird targets in each frame image. The information of each detection box includes the bird category label, the detection box position coordinates and the confidence score. Step 3.2: Subsequently, the bird target detection module performs standardization, low-confidence filtering, and deduplication / overlapping processing on all candidate detection box lists in sequence to obtain the final detection box list b; Step 3.3: The bird target detection module uses a center point calculation method to calculate the center point coordinates of each detection box in the detection box list b. The expression is: ; ; in, The coordinates of the top left corner of the detection box. The coordinates of the bottom right corner of the detection box. Here are the coordinates of the center point, and i is the index of the detection box in the current frame; Step 3.4: The bird target detection module outputs the center point coordinates, bird category label, and confidence score of each detection box to the image segmentation module.

4. The video bird count recognition method based on segmented images according to claim 1, characterized in that: In step 4, the image segmentation module performs instance-level segmentation and cross-frame consistency tracking on each frame image, as follows: The image segmentation module receives the current frame image and the center point coordinates of the corresponding detection box, then uses the center point coordinates as prompts to generate K sets of candidate mask segmentation results, and then selects the optimal segmentation mask based on a preset selection strategy. Simultaneously, the image segmentation module associates the optimal segmentation mask with the target IDs already assigned in historical frames, thus completing the ID labeling of the optimal segmentation mask for the current frame.

5. The video bird count recognition method based on segmented images according to claim 4, characterized in that: The image segmentation module generates the following expressions for K groups of candidate mask segmentation results: ; in, For the k-th candidate segmentation mask, Let K be the confidence score of the k-th candidate segmentation mask, K be the total number of candidate segmentation masks, and k be the index of the candidate segmentation mask, where k∈[1,K]. For image segmentation models, For the current frame image, This is the set of prompts used for conditional splitting. For the current time, For the set of candidate segmentation masks; The preset selection strategy is either the highest confidence score strategy or the maximum intersection-union ratio (IoU) strategy with the detection box; The highest confidence score strategy is: the confidence score of all candidate segmentation masks. Sort the masks and select the one with the highest confidence as the optimal segmentation mask. Its expression is: ; in, For optimal segmentation mask, This is a function used to return the index of the maximum value in an array; The maximum intersection-union (IoU) strategy for detection boxes is as follows: First, calculate the IoU between each candidate mask and the input detection box. The intersection-union ratio (IoU) is expressed as: ; Then, the candidate mask with the highest IoU is selected as the optimal segmentation mask. The expression is: ; in, For the k-th candidate segmentation mask, This is the detection frame.

6. The video bird count recognition method based on segmented images according to claim 4, characterized in that: The image segmentation module associates the optimal segmentation mask with the target IDs already assigned in historical frames, as follows: First, define the set of allocated masks for historical frames and their IDs as follows: Then calculate the mask for the current frame. With historical frame mask Similarity measurement The expression is: ; in, Represents the feature similarity function, , These are the weighting coefficients. for and The intersection-union ratio, where j is the index of the detection box in the historical frame; Finally, IDs are assigned based on similarity, expressed as: ; in, This is the similarity threshold.

7. The video bird count recognition method based on segmented images according to claim 1, characterized in that: In step 5, the target association and trajectory generation module generates cross-frame trajectory information in the following steps: Step 5.1: Feature Extraction: The target association and trajectory generation module obtains the segmentation mask of the current frame image, as well as the segmentation masks cached in historical frame images and their assigned ID information. Then, it uses the temporal tracking module built into the segmentation model to extract continuous cross-frame features from the mask region. The features include appearance features, motion features, and mask morphology features. The feature extraction expression is: ; in, Let be the temporal feature vector of the i-th mask in the current frame image. This provides the mask pixel information for the i-th target in the current frame. For local motion vector fields, For feature encoder; Step 5.2: Matching strategy: For each mask in the current frame, first calculate its overall similarity with the unmatched masks in the historical frame buffer, then select the historical ID with the highest overall similarity to associate with it. If the highest overall similarity exceeds the set threshold τ, then inherit the corresponding historical ID; if no historical mask has an overall similarity exceeding the threshold, then assign a new global ID to the current mask. Calculate the combined similarity between the masks of historical frames and the current frame. The expression is: Appearance similarity: ; Motion similarity: ; Morphological similarity: ; Overall similarity: ; in, , , These are the weight coefficients corresponding to the similarity features. ; Let be the temporal feature vector of the j-th mask in the tk-th frame of history. The cosine similarity function is used. For the current frame number The local motion vector field corresponding to each target History The first frame The local motion vector field corresponding to each target For the current frame number Binary segmentation mask corresponding to each target, For the first time in history The first frame Binary segmentation mask corresponding to each target, for The L2 norm; Step 5.3: ID Management and Track Update: Added ID management: Assign a unique global ID to newly appearing unmatched targets and initialize their trajectory records; ID retention and deletion: For IDs that are not matched for N consecutive frames, they are marked as "disappeared" and are still retained in the cache; when the "disappeared" state exceeds a preset time threshold, they are deleted from the cache. Track update: For each global ID, record its time series information, including timestamps. Center coordinates and area That is, ID:[( ]; Step 5.4: Output results: The target association and trajectory generation module outputs cross-frame trajectory information to the bird count module; the cross-frame trajectory information includes the current frame segmentation mask with a global ID tag and its attribute information, as well as the updated trajectory library, which includes the complete motion path and time information of all active targets.

8. The video bird count recognition method based on segmented images according to claim 1, characterized in that: The bird count module treats each unique identifier trajectory in the cross-frame trajectory information as a unique bird individual, as follows: First, the bird count module receives the trajectory data set output after cross-frame processing, the expression of which is: ; Where T represents the set of all bird trajectories that have undergone cross-frame association processing and have been assigned globally unique IDs. Let m be the trajectory of the m-th individual. The total number of individual trajectories; Within a set time window or global monitoring range, the number of different IDs in the trajectory set is counted to obtain the corresponding bird individual quantity information. The statistical formula is: ; When segmented by time window, the bird count module filters out the movement trajectory timestamps recorded in the trajectory based on the timestamp information, as expressed by: ; in, To set a time window.

9. The video bird count recognition method based on segmented images according to claim 1, characterized in that: The target detection model uses the YOLO target detection model, and the image segmentation model uses the SAM2 image segmentation model.