Wulong goose target detection and counting system based on improved YOLO11

By improving the YOLO11 model, introducing the SEAM attention mechanism and depthwise separable convolution, and combining it with an active learning strategy, the problem of goose occlusion detection was solved, achieving efficient and accurate detection and counting of goose flocks, supporting breeding management decisions, and promoting the modernization and intelligentization of goose farming.

CN120876828APending Publication Date: 2025-10-31DEZHOU UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510973192.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing goose flock target detection and counting systems have low detection accuracy under occlusion conditions, lack public datasets, and traditional methods are inefficient, making it difficult to meet the needs of intensive farming.

Method used

An improved YOLO11 model is adopted, and the SEAM attention mechanism module is introduced. Combining deep separable convolution and active learning strategies, the model is optimized through a composite loss function to achieve compensation for occluded faces and enhancement of unoccluded faces. By combining deep learning and active learning training methods, a complete modular system is designed for object detection and counting.

Benefits of technology

It significantly improves the accuracy and efficiency of target detection in dense goose flocks, reduces the workload of data labeling, provides real-time and accurate detection and counting capabilities, supports breeding management decisions, and promotes the modernization and intelligent development of the goose farming industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876828A_ABST
    Figure CN120876828A_ABST
Patent Text Reader

Abstract

The invention discloses a Wulong goose target detection and counting system based on improved YOLO11, and relates to the technical field of computer vision target detection, and the technical key points comprise the following steps: S1, a video collection module, which is used for obtaining the monitoring video data of a designated area through an arranged camera device; s2, a video frame extraction module which is used for performing frame extraction processing, quality inspection and image zooming on the video data; the algorithm training module is used for training a deep learning model based on YOLO11 through an active learning strategy, the technical effects are that by introducing the SEAM attention mechanism module, the response loss of the shielded face is effectively compensated, the response of the non-shielded face is enhanced, the target detection capability of the model under the shielding condition is remarkably improved, and the detection accuracy of the model under the shielding condition is improved. The detection problem caused by shielding when the geese are dense is solved, and the detection result is more accurate and comprehensive.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision target detection technology, specifically to a five-goose target detection and counting system based on an improved YOLO11. Background Technology

[0002] Geese are a distinctive industry in my country. Stocking density is a key factor influencing large-scale livestock production and animal welfare, and it is also one of the most important feeding and management factors affecting goose growth, development, and production performance. Currently, traditional manual counting methods in goose farming are inefficient, labor-intensive, and inaccurate, prone to double counting and omissions. Traditional meat goose farming in my country is characterized by farmers breeding and raising their own geese, seasonal breeding, and small-group grazing. This results in long breeding cycles, low feed conversion efficiency, small industry scale, and low levels of commercialization and industrialization, hindering the development of the goose industry. In recent years, my country's goose industry has gradually shifted from traditional extensive farming to intensive farming, with the emergence of more advanced methods such as thick-litter floor rearing, net-lined floor rearing, and cage rearing, significantly increasing the scale and intensification of farming. However, under intensive and industrialized production models, the excessive pursuit of economic benefits often leads to excessively high stocking densities. High-density rearing conditions can damage the living environment of geese, leading to increased bedding moisture, higher concentrations of harmful gases such as ammonia, and a greater number of pathogenic microorganisms. This can directly or indirectly negatively impact goose production and the health and welfare of the geese. Therefore, real-time monitoring of stocking density is crucial for the healthy breeding of geese.

[0003] The key to improving breeding efficiency lies in real-time monitoring of stocking density and rational allocation of flock space. Essentially, the density of a goose flock depends on the size of its effective activity space and the size of the population. Given the increasing limitations on breeding area, the main factor affecting this is the number of geese. Therefore, we will focus our research on goose flock counting.

[0004] With the development of technology, monitoring equipment plays a significant role in animal husbandry. Various methods exist for monitoring individual animal behavior, such as inserting chips to record physiological data, using wearable sensors, and (thermal) imaging technology. Some methods employ wearable sensors attached to birds' feet to measure their activity, but this can have additional effects on the monitored animals. Especially in commercial environments, technological limitations and high costs make such methods less feasible. Therefore, optical flow-based video assessment is an ideal method for monitoring poultry behavior and physiology. Initially, much surveillance video was observed manually, which was inefficient, relied on the experience and judgment of staff, and lacked standards. However, in recent years, the advent of the big data era and the rapid development of computer graphics cards have continuously improved computing power, accelerating the development of artificial intelligence. Research related to artificial intelligence is increasing, and the application of computer vision in animal detection is becoming more widespread.

[0005] Through in-depth research, the inventors discovered three main problems with current AI methods and systems for detecting and counting geese. First, most existing systems focus on chickens and ducks, lacking research on geese. Second, occlusion is a significant issue; in densely packed flocks of geese, occlusion is severe. Because two mutually occluding targets belong to the same category, their features are similar, making it difficult for the target detection algorithm to accurately locate them. Therefore, setting the threshold too high or too low will negatively impact the detection results. Finally, there is a lack of publicly available datasets for studying geese; we must collect video data ourselves and manually label it for subsequent research. Therefore, there is an urgent need to develop new technologies and systems to overcome these challenges, improving the accuracy of goose target detection and counting and its practical application in goose farms. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a five-goose target detection and counting system based on an improved YOLO11. The model incorporates a SEAM attention mechanism module to compensate for the response loss of occluded faces by enhancing the response of unoccluded faces. This module design aims to enhance the network's attention to and capture ability of occluded facial features through meticulous processing of spatial dimensions and channels, thereby achieving better target detection results.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a five-goose target detection and counting system based on an improved YOLOv11 (11 represents the version of the image processing tool), comprising the following steps:

[0008] S1, Video Acquisition Module: Used to acquire surveillance video data of a designated area through deployed camera equipment;

[0009] S2, Video frame extraction module: used for frame extraction processing, quality inspection, and image scaling of video data;

[0010] S3, Algorithm Training Module: Used to train a YOLO11-based deep learning model through an active learning strategy. The model includes an improved backbone network and neck architecture to enhance feature extraction capabilities. Two depthwise separable convolutional layers are added to the classification detection head. The bounding box regression loss and classification loss are balanced through a composite loss function.

[0011] S4, Target Detection Module: Used to perform feature recognition on frame-by-frame images based on the trained model, generate target detection boxes through non-maximum suppression algorithm, and determine the detected target based on the actual intersection-union ratio (IoU_M) and the standard intersection-union ratio (IoU_0);

[0012] S5, Counting and Statistics Module: Used to count the number of geese passing through a specified area based on the direction of coordinate change of the detected target;

[0013] S6, Target Tracking Module: Used for real-time tracking and occlusion recovery of detected targets using deep learning algorithms;

[0014] S7, Video Processing Module: Used to group, compress, and upload video files according to preset video group durations;

[0015] S8, Central Control Module: Used to monitor the online status of camera equipment and trigger alarms for equipment malfunctions.

[0016] Preferably, the video frame extraction module performs the following operations:

[0017] S21. Preset frame extraction frequency to extract frames from the video of the online camera device;

[0018] S22. Perform quality inspection on the recognized images and retain only those that meet the preset resolution and lighting standards;

[0019] S23. Scale the qualified image to fit the model input size.

[0020] Preferably, the training method of the algorithm training module is as follows:

[0021] S31. Adopt an active learning strategy and optimize model training by combining a small amount of labeled data with unlabeled data.

[0022] S32. Design a composite loss function, which is a weighted combination of bounding box regression loss and classification loss;

[0023] S33. Optimize the model's detection performance in densely occluded scenes by learning occlusion relationships.

[0024] Preferably, the target detection module determines the target through the following steps:

[0025] S41. Calculate the IoU_M of the recognized image and compare it with IoU_0;

[0026] S42. If IoU_M≥IoU_0, it is determined that a target exists; otherwise, the image is discarded.

[0027] S43. Determine whether the target detected in adjacent frames is the same target by using interactive floating values. If the difference exceeds the floating value, create a new IP and update the count.

[0028] Preferably, the target detection module determines the target through the following steps:

[0029] S41. Calculate the IoU_M of the recognized image and compare it with IoU_0;

[0030] S42. If IoU_M≥IoU_0, it is determined that a target exists; otherwise, the image is discarded.

[0031] S43. Determine whether the target detected in adjacent frames is the same target by using interactive floating values. If the difference exceeds the floating value, create a new IP and update the count.

[0032] Preferably, the calculation of the interactive floating value is dynamically adjusted based on the actual intersection-union ratio of historical frames, and the coordinates of the detected target are determined by the mass points of the target detection box.

[0033] Preferably, the counting and statistics module counts the number of geese in the following way:

[0034] S51. Determine whether the target's direction of motion is consistent with the specified direction based on the direction of coordinate change (x-axis coordinate increase or decrease);

[0035] S52. If the directions are the same, the count is incremented by 1; otherwise, the count is decremented by 1.

[0036] S53. For detected targets whose coordinates remain unchanged for multiple consecutive frames, predict the motion trend to update the count.

[0037] Preferably, the target tracking module performs real-time tracking and occlusion recovery through the following process:

[0038] S61. Receive the valid recognition image and its corresponding target coordinate information from the target detection module;

[0039] S62. A pre-trained deep learning-based target tracking algorithm continuously tracks the position of detected targets in a video sequence and generates motion trajectory data.

[0040] S63. When the target is lost due to occlusion or rapid movement, extract the target's appearance features and motion law parameters, and estimate the target's position in subsequent frames through a prediction model to restore tracking.

[0041] S64. Update the position coordinates of the detected target in real time, and feed the updated coordinates and motion trajectory back to the counting and statistics module;

[0042] S65. Generate a record of the motion state changes of the detected target based on the tracking results, which is used to correct the counting logic of the counting and statistics module and the trajectory analysis of the video processing module.

[0043] Preferably, the video processing module performs the following operations:

[0044] S71. Group the video files according to the preset duration, with the end time of each group being an integer multiple of the video group duration;

[0045] S72. Video files whose duration is not fully covered by the group should be moved to the next group;

[0046] S73. Upload video group files to the backend server using the breakpoint resume mechanism.

[0047] Preferably, the processing flow of the central control module is as follows:

[0048] S81, the central control module sends a start recording signal to the camera device and simultaneously starts a timer to record the feedback duration;

[0049] S82. Receive the confirmation recording signal returned by the camera device, stop the timer and obtain the current feedback duration;

[0050] S83. Compare the feedback duration with a preset maximum feedback duration threshold;

[0051] S84. If the feedback duration exceeds the maximum feedback duration threshold, the camera device is determined to be offline, and a device abnormality alarm is triggered.

[0052] S85. If the feedback duration is less than or equal to the maximum feedback duration threshold, the camera device is determined to be online, and subsequent video acquisition and processing operations are allowed.

[0053] Compared with existing technologies, this invention provides a target detection and counting system for five-dragon geese based on an improved YOLO11, which has the following beneficial effects: By introducing a SEAM attention mechanism module, the response loss of occluded faces is effectively compensated, and the response of unoccluded faces is enhanced, thereby significantly improving the model's target detection capability under occlusion conditions. This solves the detection problem caused by occlusion when geese are densely packed, making the detection results more accurate and comprehensive. Secondly, the model structure is optimized by using depthwise separable convolution (DWConv), reducing the amount of computation and parameters. While ensuring detection accuracy, the model's running efficiency is improved, making it more suitable for real-time detection and counting in actual breeding scenarios. Thirdly, the designed composite loss function can better balance the bounding box regression loss and classification loss, further optimizing the model's prediction performance and improving the accuracy of bounding box localization, which helps to more accurately count the number of geese. In addition, the training method that integrates deep learning and active learning can achieve good training results even with only a small amount of labeled data, greatly reducing the workload and cost of data labeling and improving the operability and practicality of the system. Finally, the system's overall architecture is well-designed, with each module working collaboratively. From video acquisition, frame extraction, target detection and tracking, counting and statistics to video processing and uploading, it forms a complete solution that enables real-time and accurate detection and counting of the Wulong goose flock. The detection results are effectively stored and managed in the form of video files, providing farmers with a powerful decision support tool. This helps improve breeding efficiency and optimize feeding management, and is of great significance to promoting the modernization and intelligent development of the goose farming industry. Attached Figure Description

[0054] Figure 1 This is a flowchart of the target detection and counting process of the present invention;

[0055] Figure 2 This is a flowchart of the video frame extraction, quality inspection, and processing process in this invention;

[0056] Figure 3 This is a flowchart of the goose detection process in this invention;

[0057] Figure 4 This is a flowchart of the goose tracking process in this invention;

[0058] Figure 5 This is a flowchart of the goose counting process in this invention. Detailed Implementation

[0059] In this invention, unless otherwise stated, the embodiments of the invention are described below through specific examples. Those skilled in the art can easily understand other advantages and effects of the invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the invention, and not all of them. The invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0060] Please see Figures 1-5 This invention provides a technical solution for a five-goose target detection and counting system based on an improved YOLO11, specifically including the following steps:

[0061] S1, Video Acquisition Module: Used to acquire surveillance video data of a designated area through deployed camera equipment;

[0062] In this embodiment, the video acquisition module is used to acquire video of a designated area. The video acquisition module includes several camera devices. The camera devices are used to acquire video of the designated area according to the start video acquisition signal and send a confirmation video acquisition signal when the video acquisition starts.

[0063] S2, Video frame extraction module: used for frame extraction processing, quality inspection, and image scaling of video data;

[0064] In this embodiment, the video captured by the video acquisition module is subjected to frame extraction processing and quality inspection. The video frame extraction module includes a video frame extraction unit and a video quality inspection unit. The video frame extraction unit is used to preset the frame extraction frequency and extract frames from the video captured by the camera device according to the frame extraction frequency to obtain recognition images. The video quality inspection unit is used to inspect the recognition images and discard images that fail the quality inspection directly.

[0065] S3, Algorithm Training Module: Used to train a YOLO11-based deep learning model through an active learning strategy. The model includes an improved backbone network and neck architecture to enhance feature extraction capabilities. Two depthwise separable convolutional layers are added to the classification detection head. The bounding box regression loss and classification loss are balanced through a composite loss function.

[0066] In this embodiment, the recognition algorithm in the target detection module is trained and optimized. This module collects a large number of video samples related to a specified region, including positive samples containing the target and negative samples not containing the target. The target in the samples is precisely labeled using an annotation tool, with annotations including the target's location, shape, size, and other feature information. Then, the labeled sample data is input into the YOLO11 deep learning framework, and the model parameters of the recognition algorithm are trained and adjusted using the backpropagation algorithm. During training, the model is continuously iterated and optimized, and its performance, such as accuracy, recall, and intersection-over-union ratio, is evaluated using validation and test sets. Based on the evaluation results, the model is fine-tuned until it achieves the expected recognition effect. The trained recognition algorithm model is then deployed to the target detection module to perform feature recognition and target detection on the images obtained by the video frame extraction module, thereby improving the accuracy and reliability of target detection. Simultaneously, the algorithm training module periodically collects new sample data to update and iteratively train the model to adapt to changes in the target and the influence of environmental factors, ensuring the long-term effectiveness and accuracy of the recognition algorithm.

[0067] S4, Target Detection Module: Used to perform feature recognition on frame-by-frame images based on the trained model, generate target detection boxes through non-maximum suppression algorithm, and determine the detected target based on the actual intersection-union ratio (IoU_M) and the standard intersection-union ratio (IoU_0);

[0068] S5, Counting and Statistics Module: Used to count the number of geese passing through a specified area based on the direction of coordinate change of the detected target;

[0069] In this embodiment, a preset specified motion direction is used to determine whether the motion direction of the detected target is consistent with the specified motion direction based on the coordinate record of the detected target, and to record the number of detected targets passing through the specified area based on the determination result, and to sort the valid identification images of the detected targets with the same IP according to the shooting time order and generate a video file of the detected targets of the IP.

[0070] S6, Target Tracking Module: Used for real-time tracking and occlusion recovery of detected targets using deep learning algorithms;

[0071] In this embodiment, the module receives valid recognized images and related coordinate information from the target detection module, and uses an advanced deep learning target tracking algorithm to continuously track the position of the detected target in the video sequence. During the tracking process, when the target is temporarily lost due to occlusion, rapid movement, or other reasons, the target tracking module can use the target's appearance features, motion patterns, and other information to predict and recover the tracking. Simultaneously, the module updates the target's position information in real time and feeds the tracking results back to the counting and statistics module to accurately record the target's motion trajectory and state changes, providing accurate data for subsequent counting and statistics and video processing.

[0072] S7, Video Processing Module: Used to group, compress, and upload video files according to preset video group durations.

[0073] In this embodiment, several start and end times for video groups are set according to the duration of each video group. Video files are grouped based on their end times and the end times of the video files. When the end time of a video file is less than or equal to the end time of the video group, the video file is added to the video group. When the end time of a video file is greater than the end time of the video group, the video group is considered to have been generated. The videos are then compressed into a single video group file and added to the video group for the next time period. The video processing module also uploads the video group file to the backend and checks whether the upload was successful. If the upload fails, the video group file is re-uploaded. If the upload is successful, the module checks whether there are any new video group files that have not yet been uploaded. If so, the upload continues; otherwise, the module waits for a new video group file to be generated.

[0074] S8, Central Control Module: Used to monitor the online status of camera equipment and trigger alarms for equipment malfunctions;

[0075] In this embodiment, a system is configured to send a start recording signal to the camera device and start recording the feedback duration, receive a confirmation recording signal and stop recording the feedback duration, preset a maximum feedback duration, and determine whether the camera device is online based on the maximum feedback duration and the feedback duration. When the feedback duration is greater than the maximum feedback duration, the central control module determines that the camera device is offline and issues an alarm for device abnormality. When the feedback duration is less than or equal to the maximum feedback duration, the central control module determines that the camera device is online.

[0076] The video frame extraction module performs the following operations:

[0077] S21. Preset frame extraction frequency to extract frames from the video of the online camera device;

[0078] S22. Perform quality inspection on the recognized images and retain only those that meet the preset resolution and lighting standards;

[0079] S23. Scale the qualified image to fit the model input size;

[0080] The training method for the algorithm training module is as follows:

[0081] S31. Adopt an active learning strategy and optimize model training by combining a small amount of labeled data with unlabeled data.

[0082] S32. Design a composite loss function, which is a weighted combination of bounding box regression loss and classification loss. The specific formula is as follows:

[0083] L total =α·L box +β·L cls

[0084] In the formula, L box For bounding box regression loss, the CIoU loss function is used to calculate the positional deviation between the predicted box and the ground truth box, L. cls For classification loss, the cross-entropy loss function is used to calculate the target category prediction error. α and β are dynamic weight coefficients with a value range of [0,1][0,1]. The initial values ​​are set to α = 0.5 and β = 0.5, and are automatically adjusted according to the loss ratio during training.

[0085] The dynamic weighting coefficients are adjusted as follows:

[0086]

[0087] Specifically, when the proportion of classification loss increases, α is reduced to decrease the bounding box regression weight, thus preventing the model from focusing too much on localization and ignoring classification.

[0088] S33. Optimize the model's detection performance in densely occluded scenes by learning occlusion relationships;

[0089] The target detection module determines the target by following these steps:

[0090] S41. Calculate the IoU_M of the recognized image and compare it with IoU_0;

[0091] S42. If IoU_M≥IoU_0, it is determined that a target exists; otherwise, the image is discarded.

[0092] S43. Determine whether the target detected in adjacent frames is the same target by interactive floating value. If the difference exceeds the floating value, create a new IP and update the count.

[0093] A preset interaction float value is used. Based on the interaction float value, the actual cross-union ratio of the previous valid recognition image, and the actual cross-union ratio of the valid recognition image, it is determined whether the detected target in the valid recognition image is the same as the detected target in the previous valid recognition image.

[0094] When (actual cross-union ratio of the previous valid recognition image - interactive floating value) is less than or equal to the actual cross-union ratio of the valid recognition image, or when the actual cross-union ratio of the valid recognition image is less than or equal to (actual cross-union ratio of the previous valid recognition image + interactive floating value), it is determined that the detection target of the valid recognition image and the detection target of the previous valid recognition image are the same target. The IP of the detection target of the valid recognition image is set to the IP of the detection target of the previous valid recognition image, and the coordinates of the detection target of the valid recognition image are obtained.

[0095] When (actual crossover ratio of the previous valid recognition image - crossover and floating value) is greater than the actual crossover ratio of the valid recognition image, or when the actual crossover ratio of the valid recognition image is greater than (actual crossover ratio of the previous valid recognition image + crossover and floating value), it is determined that the detected target of the valid recognition image is a different target from the detected target of the previous valid recognition image. It is determined that the detected target of the previous valid recognition image passes through the specified area. At the same time, a new IP is created for the detected target of the valid recognition image and the coordinates of the detected target are obtained.

[0096] The calculation of the interactive floating value is dynamically adjusted based on the actual intersection-over-union ratio of historical frames, and the coordinates of the detected target are determined by the mass points of the target detection box;

[0097] The counting and statistics module counts the number of geese in the following ways:

[0098] Based on the shooting time sequence of the effectively recognized images, the direction of coordinate change for targets with the same IP address is determined. The coordinates of the target at shooting time t1 are (x1, y1), and the coordinates of the target at shooting time t2 are (x2, y2). t1 is less than t2.

[0099] When x1 is less than x2, it is determined that the movement direction of the detected target is consistent with the specified direction, and the number of detected targets passing through the specified area is incremented by 1.

[0100] When x1 is greater than x2, it is determined that the movement direction of the detected target is consistent with the specified direction, and the number of detected targets passing through the specified area is decremented by 1.

[0101] When x1 equals x2, obtain the new coordinates (xn, yn) of the detected target at time tn, compare the magnitudes of x1 and xn, tn is greater than t2, which is greater than t1, n = 3, 4, ..., m;

[0102] S51. Determine whether the target's direction of motion is consistent with the specified direction based on the direction of coordinate change (x-axis coordinate increase or decrease);

[0103] S52. If the directions are the same, the count is incremented by 1; otherwise, the count is decremented by 1.

[0104] S53. Predict the motion trend of detected targets whose coordinates do not change for multiple consecutive frames to update the count;

[0105] The target tracking module performs real-time tracking and occlusion recovery through the following process:

[0106] S61. Receive the valid recognition image and its corresponding target coordinate information from the target detection module;

[0107] S62. A pre-trained deep learning-based target tracking algorithm continuously tracks the position of detected targets in a video sequence and generates motion trajectory data.

[0108] S63. When the target is lost due to occlusion or rapid movement, extract the target's appearance features and motion law parameters, and estimate the target's position in subsequent frames through a prediction model to restore tracking.

[0109] S64. Update the position coordinates of the detected target in real time, and feed the updated coordinates and motion trajectory back to the counting and statistics module;

[0110] S65. Generate a record of the motion state change of the detected target based on the tracking results, which is used to correct the counting logic of the counting and statistics module and the trajectory analysis of the video processing module.

[0111] The video processing module performs the following operations:

[0112] S71. Group the video files according to the preset duration, with the end time of each group being an integer multiple of the video group duration;

[0113] S72. Video files whose duration is not fully covered by the group should be moved to the next group;

[0114] S73. Upload video group files to the backend server using a breakpoint resume mechanism;

[0115] The system presets the duration of each video group, groups video files according to the duration, generates video group files, uploads these video group files to the backend, and checks whether the video group files were uploaded successfully. If the upload was successful, the system checks if any video group files were not uploaded and continues uploading them or waits for them to be uploaded. If the upload failed, the system re-uploads the video group files.

[0116] The processing flow of the central control module is as follows:

[0117] S81, the central control module sends a start recording signal to the camera device and simultaneously starts a timer to record the feedback duration;

[0118] S82. Receive the confirmation video signal returned by the camera device, stop the timer and obtain the current feedback duration;

[0119] S83. Compare the feedback duration with the preset maximum feedback duration threshold;

[0120] S84. If the feedback time exceeds the maximum feedback time threshold, the camera device is determined to be offline and an abnormal device alarm is triggered.

[0121] S85. If the feedback duration is less than or equal to the maximum feedback duration threshold, the camera device is determined to be online, and subsequent video acquisition and processing operations are allowed.

[0122] The above are merely specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent substitutions, or modifications made based on the present invention to solve essentially the same technical problems and achieve essentially the same technical effects are all covered within the protection scope of the present invention.

Claims

1. A target detection and counting system for five dragon geese based on an improved YOLOv11, characterized in that, Specifically, the following steps are included: S1, Video Acquisition Module: Used to acquire surveillance video data of a designated area through deployed camera equipment; S2, Video frame extraction module: used for frame extraction processing, quality inspection, and image scaling of video data; S3, Algorithm Training Module: Used to train a YOLO11-based deep learning model through an active learning strategy. The model includes an improved backbone network and neck architecture to enhance feature extraction capabilities. Two depthwise separable convolutional layers are added to the classification detection head. The bounding box regression loss and classification loss are balanced through a composite loss function. S4, Target Detection Module: Used to perform feature recognition on frame-by-frame images based on the trained model, generate target detection boxes through non-maximum suppression algorithm, and determine the detected target based on the actual intersection-union ratio (IoU_M) and the standard intersection-union ratio (IoU_0); S5, Counting and Statistics Module: Used to count the number of geese passing through a specified area based on the direction of coordinate change of the detected target; S6, Target Tracking Module: Used for real-time tracking and occlusion recovery of detected targets using deep learning algorithms; S7, Video Processing Module: Used to group, compress, and upload video files according to preset video group durations; S8, Central Control Module: Used to monitor the online status of camera equipment and trigger alarms for equipment malfunctions.

2. The five-goose target detection and counting system based on improved YOLOv11 according to claim 1, characterized in that: The video frame extraction module performs the following operations: S21. Preset frame extraction frequency to extract frames from the video of the online camera device; S22. Perform quality inspection on the recognized images and retain only those that meet the preset resolution and lighting standards; S23. Scale the qualified image to fit the model input size.

3. The five-goose target detection and counting system based on the improved YOLOv11 according to claim 2, characterized in that: The training method for the algorithm training module is as follows: S31. Adopt an active learning strategy and optimize model training by combining a small amount of labeled data with unlabeled data. S32. Design a composite loss function, which is a weighted combination of bounding box regression loss and classification loss, with the following formula: L total =α·L box +β·L cls In the formula, L box For bounding box regression loss, L cls For classification loss, α and β are dynamic weight coefficients with values ​​ranging from [0,1] to [0,1]. The initial values ​​are set to α = 0.5 and β = 0.5, and they are automatically adjusted according to the loss ratio during training. S33. Optimize the model's detection performance in densely occluded scenes by learning occlusion relationships.

4. The five-goose target detection and counting system based on improved YOLOv11 according to claim 3, characterized in that: The target detection module determines the target through the following steps: S41. Calculate the actual intersection-union ratio (IoU_M) of the recognized image and compare it with the standard intersection-union ratio (IoU_0); S42. If IoU_M≥IoU_0, it is determined that a target exists; otherwise, the image is discarded. S43. Determine whether the target detected in adjacent frames is the same target by using interactive floating values. If the difference exceeds the floating value, create a new IP and update the count.

5. The five-goose target detection and counting system based on the improved YOLO11 according to claim 4, characterized in that: The calculation of the interactive floating value is dynamically adjusted based on the actual intersection-over-union ratio of historical frames, and the coordinates of the detected target are determined by the mass points of the target detection box.

6. The five-goose target detection and counting system based on the improved YOLOv11 according to claim 4, characterized in that: The counting and statistics module counts the number of geese in the following way: S51. Based on the direction of coordinate change of the detected target, determine whether its direction of movement is consistent with the specified direction; S52. If the directions are the same, the count is incremented by 1; otherwise, the count is decremented by 1. S53. For detected targets whose coordinates remain unchanged for multiple consecutive frames, predict the motion trend to update the count.

7. The five-goose target detection and counting system based on improved YOLOv11 according to claim 6, characterized in that: The target tracking module performs real-time tracking and occlusion recovery through the following process: S61. Receive the valid recognition image and its corresponding target coordinate information from the target detection module; S62. A pre-trained deep learning-based target tracking algorithm continuously tracks the position of detected targets in a video sequence and generates motion trajectory data. S63. When the target is lost due to occlusion or rapid movement, extract the target's appearance features and motion law parameters, and estimate the target's position in subsequent frames through a prediction model to restore tracking. S64. Update the position coordinates of the detected target in real time, and feed the updated coordinates and motion trajectory back to the counting and statistics module; S65. Generate a record of the motion state changes of the detected target based on the tracking results, which is used to correct the counting logic of the counting and statistics module and the trajectory analysis of the video processing module.

8. The five-goose target detection and counting system based on improved YOLOv11 according to claim 7, characterized in that: The video processing module performs the following operations: S71. Group the video files according to the preset duration, with the end time of each group being an integer multiple of the video group duration; S72. Video files whose duration is not fully covered by the group should be moved to the next group; S73. Upload video group files to the backend server using the breakpoint resume mechanism.

9. A five-goose target detection and counting system based on an improved YOLOv11 according to claim 8, characterized in that: The processing flow of the central control module is as follows: S81, the central control module sends a start recording signal to the camera device and simultaneously starts a timer to record the feedback duration; S82. Receive the confirmation recording signal returned by the camera device, stop the timer and obtain the current feedback duration; S83. Compare the feedback duration with a preset maximum feedback duration threshold; S84. If the feedback duration exceeds the maximum feedback duration threshold, the camera device is determined to be offline, and a device abnormality alarm is triggered. S85. If the feedback duration is less than or equal to the maximum feedback duration threshold, the camera device is determined to be online, and subsequent video acquisition and processing operations are allowed.