Enhanced Method and System for Small Target Detection and Tracking in Swimming Venues with Multi-Strategy Fusion

By adopting a multi-strategy fusion method in swimming venues, the sensitivity and accuracy of the target detection model are improved, and combined with multi-threshold detection and compensation tracking strategies, the problem of difficulty in detecting and tracking small targets in traditional technologies is solved, achieving higher detection accuracy and tracking stability.

CN119741342BActive Publication Date: 2025-06-10巨岩智能科技(杭州)有限公司 +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510243575.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-10
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

In swimming venues, traditional object detection algorithms are difficult to effectively extract the characteristics of small targets, resulting in poor detection results, especially in the case of complex backgrounds, dense targets and high-speed movement.

Method used

Using a multi-strategy fusion method, by acquiring the initial image and inputting the object detection model, identifying potential small target areas, cropping and data enhancement, and improving the sensitivity and accuracy of the object detection model. Then, the detected ROI area is cropped and amplified, combining multi-threshold detection and compensation tracking strategies to ensure stable tracking of the target.

Benefits of technology

It significantly improves the detection accuracy and tracking stability of small targets, reduces misidentification and loss, and improves the accuracy and robustness of overall tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741342B_ABST
    Figure CN119741342B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for enhancing the detection and tracking of small targets in a swimming pool venue through multi-strategy fusion. The method includes: acquiring an image to be detected and tracked to obtain an initial image; inputting the initial image into an object detection model for object detection to obtain an ROI region; wherein, the object detection model is obtained by cropping and data augmenting the images with targets in the swimming pool venue and using them as a sample set to train a pre-trained detection model; cropping and magnifying the ROI region, and inputting the processed ROI region into the object detection model again for object detection to obtain a detection result; mapping the detection result back to the coordinate system corresponding to the initial image to obtain a mapping result; performing multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result; and outputting the tracking result. By implementing the method of the present invention, it is possible to improve the detection accuracy of small targets and achieve more stable target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically to a method and system for enhancing the detection and tracking of small targets in a swimming pool venue through multi-strategy fusion. Background Art

[0002] In recent years, computer vision technology has made remarkable progress in multiple application fields. Especially in object detection and tracking, deep learning algorithms have been widely adopted. However, in practical applications, the detection and tracking of small targets still face many challenges. Especially when the target size is small, the background is complex, or the targets are dense, problems such as missed detection or false detection often occur. These problems are particularly prominent in indoor swimming pool venues.

[0003] Swimming pool venues have some unique characteristics, which exacerbate the difficulty of small target detection and tracking. Specifically, the surveillance cameras in swimming pool venues usually cover the entire pool area, which makes the targets in the pool, such as the heads and bodies of swimmers, relatively small in the image. Traditional object detection algorithms are difficult to effectively extract the features of such small-sized targets, thus affecting the detection effect. Background factors in the swimming pool environment, such as water surface reflection, water ripples, lane lines, and buoys, all have dynamic changes, and these factors are likely to interfere with the recognition of small targets by object detection algorithms. Swimmers move at a relatively fast speed and have irregular trajectories, which pose higher requirements for real-time performance and robustness of object tracking algorithms. Especially in the case of high-speed movement, the tracking accuracy may be affected. During the peak swimming period, there may be a large number of swimmers in the pool, and the targets may be blocked or overlapped with each other, which increases the difficulty of small target detection and tracking.

[0004] In response to these problems, the main challenges faced by current surveillance requirements include: difficulties in extracting features of small targets, insufficient integration of detection and tracking, and easy loss of tracking in dynamic scenarios.

[0005] Therefore, it is necessary to design a new method that can improve the detection accuracy of small targets and achieve more stable object tracking. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the prior art and provide a method and system for enhancing the detection and tracking of small targets in a swimming pool venue through multi-strategy fusion.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions: A method for enhancing the detection and tracking of small targets in a swimming pool venue through multi-strategy fusion, including:

[0008] Obtain the image to be detected and tracked to obtain the initial image;

[0009] Input the initial image into the target detection model for target detection to obtain the ROI region; wherein, the target detection model is obtained by cropping and data augmenting the images with targets in the swimming stadium and using them as a sample set to train the pre-trained detection model;

[0010] Crop and magnify the ROI region, and input the processed ROI region into the target detection model again for target detection to obtain the detection result;

[0011] Map the detection result back to the coordinate system corresponding to the initial image to obtain the mapping result;

[0012] Perform multi-threshold detection and compensation tracking on the mapping result to obtain the tracking result;

[0013] Output the tracking result.

[0014] Its further technical solution is: the target detection model is obtained by cropping and data augmenting the images with targets in the swimming stadium and using them as a sample set to train the pre-trained detection model, including:

[0015] Obtain the images with targets in the swimming stadium;

[0016] Crop the image using a dynamic sliding window to obtain the cropping result;

[0017] Perform label annotation on the cropping result to generate a training set;

[0018] Perform data augmentation on the training set to obtain a sample set;

[0019] Construct a pre-trained detection model and a loss function;

[0020] Use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0021] Its further technical solution is: the cropping of the image using a dynamic sliding window to obtain the cropping result includes:

[0022] Crop the image using different sliding windows according to the lane area and non-critical areas to obtain the cropping result.

[0023] Its further technical solution is: the label annotation of the cropping result to generate a training set includes:

[0024] Perform target annotation on the cropping result and record the position of the target to obtain a training set.

[0025] Its further technical solution is: performing data augmentation on the training set to obtain a sample set, including:

[0026] Identifying and annotating the head and body features of the swimmers in the training set, and adjusting the scaling ratio to enlarge the extracted area of the training set to obtain an enlarged image;

[0027] Using the target tiling algorithm to splice the enlarged images according to the distribution of targets in the actual scene, and updating the corresponding annotation information to obtain a synthesis result;

[0028] Performing weighted processing on the small targets in the synthesis result to obtain a weighted processing result;

[0029] Using a pre-trained model to screen out small target instances with confidence meeting the requirements from the weighted processing result to obtain a sample set.

[0030] Its further technical solution is: performing multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result, including:

[0031] Setting a first threshold and a second threshold, where the first threshold is greater than the second threshold, the first threshold is used for multi-target tracking; the second threshold is used to compensate for lost target boxes;

[0032] Performing multi-target tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is lost in the current frame, starting a compensation mechanism, and using IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-target tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result.

[0033] Its further technical solution is: performing multi-target tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is lost in the current frame, starting a compensation mechanism, and using IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-target tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result, including:

[0034] Using a multi-target tracking algorithm to perform multi-target tracking on the results in the mapping result that are greater than the first threshold to obtain a multi-target tracking result, where the multi-target tracking result includes a list of human tracking boxes with tracking IDs;

[0035] Checking the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, using IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-target tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result.

[0036] Its further technical solution is as follows: frame by frame, check the status of each tracking object. When it is found that the tracking ID of a certain swimmer disappears in the current frame and the swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find a matching target box from the results below the second threshold in the multi-object tracking results to compensate for the lost human tracking box, so as to obtain the tracking result, including:

[0037] Frame by frame, check the status of each tracking object. When it is found that the tracking ID of a certain swimmer disappears in the current frame and the swimmer does not swim out of the range, compare the human tracking box corresponding to the disappeared human tracking box in the multi-object tracking results with all the human tracking boxes below the second threshold in the current frame, and calculate the intersection over union between the two human tracking boxes to obtain the IoU matrix;

[0038] Construct a cost matrix according to the IoU matrix;

[0039] Apply the Hungarian algorithm to the cost matrix to match all the human tracking boxes below the second threshold in the current frame and the disappeared human tracking box, so as to obtain the matching result;

[0040] Add the matching result to the multi-object tracking result of the current frame to obtain the tracking result.

[0041] The present invention also provides a multi-strategy fusion swimming pool small target detection and tracking enhancement system, including:

[0042] An image acquisition unit, configured to acquire an image to be detected and tracked to obtain an initial image;

[0043] A target detection unit, configured to input the initial image into a target detection model for target detection to obtain an ROI region; wherein, the target detection model is obtained by cropping and data augmenting the images with targets in the swimming pool as a sample set to train a pre-trained detection model;

[0044] A re-detection unit, configured to crop and enlarge the ROI region, and input the processed ROI region into the target detection model again for target detection to obtain a detection result;

[0045] A mapping unit, configured to map the detection result back to the coordinate system corresponding to the initial image to obtain a mapping result;

[0046] A tracking unit, configured to perform multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result;

[0047] An output unit, configured to output the tracking result.

[0048] Its further technical solution is: It further includes a training unit for:

[0049] Obtain an image with a target in the swimming stadium;

[0050] Crop the image using a dynamic sliding window to obtain a cropping result;

[0051] Label the cropping result to generate a training set;

[0052] Perform data augmentation on the training set to obtain a sample set;

[0053] Construct a pre-trained detection model and a loss function;

[0054] Use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0055] The beneficial effects of the present invention compared with the prior art are: By obtaining an initial image and inputting it into the target detection model, potential small target areas are identified. Through cropping and data augmentation of the training data set, the sensitivity and accuracy of the target detection model for small targets are improved. The detected ROI areas are cropped and enlarged to further enhance the detailed features of small targets and improve the accuracy of subsequent detections. Map the detection results back to the coordinate system of the initial image to ensure the accuracy of the target position. Combine multi-threshold detection and compensation tracking strategies to stably track the target, reduce misidentifications and losses, and thus improve the overall tracking stability and accuracy.

[0056] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 It is a flowchart of the multi-strategy fusion method for enhancing the detection and tracking of small targets in a swimming stadium provided by an embodiment of the present invention;

[0059] Figure 2 It is a sub-flowchart of the multi-strategy fusion method for enhancing the detection and tracking of small targets in a swimming stadium provided by an embodiment of the present invention Figure 1 ;

[0060] Figure 3Schematic diagram of the sub - process of the enhanced small - target detection and tracking method for swimming venues with multi - strategy fusion provided by the embodiments of the present invention Figure 2 ;

[0061] Figure 4 Schematic diagram of the sub - process of the enhanced small - target detection and tracking method for swimming venues with multi - strategy fusion provided by the embodiments of the present invention Figure 3 ;

[0062] Figure 5 Schematic diagram of the sub - process of the enhanced small - target detection and tracking method for swimming venues with multi - strategy fusion provided by the embodiments of the present invention Figure 4 ;

[0063] Figure 6 Schematic diagram of the sub - process of the enhanced small - target detection and tracking method for swimming venues with multi - strategy fusion provided by the embodiments of the present invention Figure 5 ;

[0064] Figure 7 Schematic diagram of the multi - threshold detection and compensation tracking process of the enhanced small - target detection and tracking method for swimming venues with multi - strategy fusion provided by the embodiments of the present invention;

[0065] Figure 8 Schematic block diagram of the enhanced small - target detection and tracking system for swimming venues with multi - strategy fusion provided by the embodiments of the present invention;

[0066] Figure 9 Schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0068] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0069] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0070] It should be further understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0071] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the multi-strategy fusion method for enhancing small target detection and tracking in a swimming pool venue provided by an embodiment of the present invention. The multi-strategy fusion method for enhancing small target detection and tracking in a swimming pool venue is applied to a server, which interacts with a terminal and a camera for data. By cropping and data enhancing the images in the swimming pool venue, diverse training samples are generated to improve the recognition ability of the target detection model for small targets. The ROI region in the initial image is cropped and enlarged to enhance the details of small targets and further improve the detection accuracy. The detection result is mapped back to the initial image coordinate system to ensure the accuracy of the target position. Combining the multi-threshold detection and compensation tracking method ensures stable tracking of small targets and reduces the phenomena of loss and false detection.

[0072] Figure 1 is a schematic flowchart of the multi-strategy fusion method for enhancing small target detection and tracking in a swimming pool venue provided by an embodiment of the present invention. As Figure 1 shown, the method includes the following steps S110 to S160.

[0073] S110. Obtain the image to be detected and tracked to obtain the initial image.

[0074] In this embodiment, the initial image refers to the original image first obtained from the swimming pool venue monitoring system or other image acquisition devices in the target detection and tracking task. These images contain the targets (such as swimmers) to be detected and tracked and serve as the basis for subsequent processing. The initial image is usually the original input image without any processing or enhancement, and is mainly used to extract the potential regions (ROIs) of the targets and subsequent processing steps before target recognition by the target detection model.

[0075] S120. Input the initial image into the target detection model for target detection to obtain the ROI region; wherein, the target detection model is obtained by cropping and data enhancing the images with targets in the swimming pool venue as the sample set to train the pre-trained detection model.

[0076] In this embodiment, the ROI region refers to the region with the target of the swimmer.

[0077] Specifically, through the target detection model, small target regions in the image are identified, and the corresponding ROI regions are extracted; the image with an original resolution of 1920×1080 is segmented by a pre-made mask to determine a specific in-pool area and filter out areas such as the shore that do not need attention.

[0078] In one embodiment, refer to Figure 2 , the above target detection model is obtained by cropping and data augmenting the images with targets in the swimming venue as the sample set to train the pre-trained detection model, including: steps S121~S126.

[0079] S121. Obtain images with targets in the swimming venue.

[0080] In this embodiment, video streams or picture libraries from the swimming venue monitoring system are collected to ensure that the images cover various scenarios and lighting conditions in the pool area. These images will serve as the basis for subsequent processing.

[0081] S122. Crop the image using a dynamic sliding window to obtain a cropping result.

[0082] In this embodiment, the cropping result refers to the result obtained by cropping the image according to different sliding windows.

[0083] Specifically, the image is cropped using different sliding windows according to the lane area and non-critical areas to obtain a cropping result. According to the density and distribution of targets in the image, the size and step length of the sliding window are dynamically adjusted, rather than using fixed window sizes and overlap ratios. Ensure more effective coverage of small targets when the target density is high in the pool area.

[0084] According to the characteristics of the pool environment (such as lane layout) and the distribution of targets in the image, use a sliding window algorithm with dynamically adjusted size and step length to crop the input image. The initial window size is set to 256×256 pixels, and the default overlap ratio is 50%, but the window size and overlap ratio will be adaptively adjusted according to the actual target density. Specifically, the window size is reduced in the lane area to increase the coverage of small targets, and the window is enlarged between lanes to reduce interference from irrelevant backgrounds. The position information of each cropping window is recorded for subsequent processing. The specific cropping formula is: . Where represents the step length of each slide of the sliding window, represents the window size, represents the proportion of the overlapping area between adjacent windows, which determines the overlapping degree of two consecutive windows when sliding.

[0085] Optimize the window clipping strategy considering the lane layout in the swimming venue and the typical movement trajectories of swimmers. Specifically, use smaller windows within the lane area to improve the coverage of small targets, while use larger windows at the lane intervals to reduce the clipping of irrelevant backgrounds.

[0086] S123. Perform label annotation on the clipping result to generate a training set.

[0087] In this embodiment, perform object annotation on the clipping result and record the position of the object to obtain a training set.

[0088] Specifically, perform accurate object annotation on each clipped image segment, especially the head and body positions of swimmers, and mark the position of each clipping window relative to the original image. This step creates a new and rich training set of small target instances.

[0089] Perform detailed object annotation on each small clipped image, while recording the position information of each clipping window relative to the original image. This step ensures that all small target instances are correctly marked, thus providing a high-quality data basis for subsequent steps. After completing clipping and annotation, integrate these data to form a new training set. This training set pays special attention to covering more small target instances, enhancing the model's ability to recognize small targets.

[0090] S124. Perform data augmentation on the training set to obtain a sample set.

[0091] In this embodiment, the sample set refers to the data set used to train and verify the performance of the trained model.

[0092] Data augmentation: To improve the generalization ability and robustness of the model, use multiple methods to expand the training data. For example:

[0093] Scaling and translation: Enlarge the image by a certain ratio (e.g., 2 times) and re-annotate the positions of small targets.

[0094] Tiling and stitching: Based on the actual layout of the swimming pool, randomly select multiple images containing small targets and stitch them into a larger composite image.

[0095] Object insertion: Simulate real movement trajectories, insert additional small targets into the new image, and update the corresponding annotation information.

[0096] Weighted annotation: Give higher weights to objects that meet the definition of small targets, especially at the head position, to emphasize their importance during training.

[0097] In one embodiment, please refer to Figure 3 , the above step S124 may include steps S1241 to S1244.

[0098] S1241. Identify and label the head and body features of the swimmers in the training set, and adjust the scaling ratio to enlarge the area extracted from the training set to obtain an enlarged image.

[0099] In this embodiment, the enlarged image refers to the area where the head and body features of the swimmer are identified and marked, and then the ratio is adjusted again to enlarge the area where the head and body features are located.

[0100] Combined with the characteristics of the pool scene, use a pre-trained pool detection model to identify and focus on labeling the positions of the heads and bodies of the swimmers. Extract features using the typical postures of the swimmers to improve the model's understanding ability of specific postures.

[0101] Considering that the size of small targets in the original image is small, the bilinear interpolation algorithm is used to enlarge the original image proportionally, with the default being 2 times. The positions of the small targets are re-labeled in the enlarged image and used in the training stage to better capture the details of the small targets and improve the detection rate. The specific implementation uses the following formula, where represents the transformed horizontal axis coordinate, represents the transformed vertical axis coordinate, represents the width of the transformed target box, represents the height of the transformed target box, represents the image scaling ratio. , where x, y, w, and h correspond to the horizontal axis coordinate, vertical axis coordinate, width of the target box, and height of the target box before transformation respectively.

[0102] S1242. Use the target tiling algorithm to splice the enlarged images according to the distribution of targets in the actual scene, and update the corresponding annotation information to obtain a synthesis result.

[0103] In this embodiment, the synthesis result refers to the result formed by splicing several images with small targets selected from the above-mentioned enlarged images.

[0104] Specifically, considering the actual layout of the pool, such as the lane distribution and edge positions, randomly select 4 images containing small targets from the existing training set, specifically the enlarged images, and splice them into a new image according to a 2×2 layout. This method ensures the realism of the synthesized image and avoids unreasonable target arrangements. The updated annotation information ensures that the annotation boxes of the original image can be correctly mapped to the new image, increasing the diversity of training samples. The mapping formula is:

[0105] , where represents the mapped horizontal axis coordinate, Represents the mapped vertical axis, Represents the horizontal axis coordinate before mapping, Represents the horizontal axis coordinate after mapping, 、 Respectively represent the offsets of the horizontal axis coordinate and the vertical axis coordinate.

[0106] Based on the movement trajectories of swimmers and the layout of the swimming pool, a target insertion algorithm is designed. This algorithm first crops out small target objects, then inserts them into a new background image, and at the same time adjusts the size to match the scale of the current image. Finally, the annotation information is updated to ensure that the information of the inserted target is accurately added to the annotation file of the new image.

[0107] S1243. Perform weighted processing on the small targets in the synthesis result to obtain a weighted processing result.

[0108] In this embodiment, a standard is set to define which targets belong to "small targets". For example, at a resolution of 1920x1080, targets with a width and height both less than 40 pixels are regarded as small targets; or more generally, the small target threshold can be set to 5% of the width and height of the image resolution.

[0109] To make the model pay more attention to small targets, especially the heads of swimmers, higher weights are given to these targets. By adjusting the weight parameters in the loss function, the small target detection task occupies a larger proportion in the training process, thereby strengthening the model's attention and detection accuracy for small targets.

[0110] To sum up, the above measures act together on the data preparation and model training processes, aiming to improve the detection performance of small targets (such as distant or smaller swimmers) in the swimming pool monitoring system, and ensure that targets can be efficiently and accurately identified and tracked even in complex environments.

[0111] S1244. Use a pre-trained model to screen out small target instances with confidence levels meeting the requirements from the weighted processing result to obtain a sample set.

[0112] In this embodiment, a pre-trained detection model is used to perform inference on a large amount of swimming pool monitoring data, and high-confidence results containing small targets are screened out. The screened images are manually verified to ensure the accuracy of the annotations, thereby constructing a dedicated small target training set to further optimize the model performance.

[0113] S125. Construct a pre-trained detection model and a loss function.

[0114] In this embodiment, a deep learning architecture suitable for processing small targets is developed or selected as the basis for the pre-trained detection model, such as a deep learning network, etc. An appropriate loss function is designed or selected, which can identify and give more attention to small targets, and this is achieved by setting weight parameters.

[0115] S126. Use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0116] In this embodiment, the pre-trained detection model constructed above is trained using the sample set after data augmentation. During the training process, the model parameters are continuously optimized until satisfactory performance indicators are achieved. Finally, the fully trained model will have stronger small target detection capabilities and can be used to accurately identify and track swimmers in the actual application environment, especially those small targets that may be difficult to capture.

[0117] In summary, this series of steps aims to improve the detection accuracy and reliability of small targets (such as distant swimmers) in the swimming pool monitoring system through fine data preprocessing, effective data augmentation strategies, and targeted model training.

[0118] S130. Crop and magnify the ROI region, and input the processed ROI region into the target detection model again for target detection to obtain the detection result.

[0119] In this embodiment, the detection result refers to the region where the small target is located.

[0120] Specifically, first, an image with an original resolution of 1920×1080 is segmented using a pre-made mask image. This mask is used to accurately locate the area inside the swimming pool, thereby effectively filtering out the shore or other irrelevant background parts. This step ensures that subsequent processing only focuses on the activities inside the swimming pool.

[0121] Within the region determined by the above mask, use the target detection model to identify and locate all small target regions of interest (ROI). For each detected ROI, perform a cropping operation and moderately magnify these regions, by default, twice the size. This process not only helps to enhance the feature expression of small targets but also reduces the influence of background information on the detection result.

[0122] The cropped and magnified image is input into the YOLOv5 detection model, which has been pre-trained using a dataset customized for the private swimming pool scenario. This pre-training in a specific scenario improves the generalization ability and accuracy of the model in actual applications, especially the recognition ability for small targets.

[0123] To ensure that the model can identify global targets within a large range and accurately capture local small targets, the cropped image is combined with the original image. Specifically, the ROI after being enlarged and further processed is remapped back to the original image position to form a new image with rich details. In addition, an effective algorithm is used to remove duplicates from the results of multi-object detection, avoiding double counting the same object and ensuring the accuracy of the detection results.

[0124] S140. Map the detection results back to the coordinate system corresponding to the initial image to obtain a mapping result.

[0125] In this embodiment, the mapping result refers to the result obtained by mapping all the coordinates of the detection results to the coordinate system of the original initial image.

[0126] Finally, for all the detection results obtained in the image after Crop, their coordinate positions are accurately mapped back to the corresponding positions in the original image. Mapping according to the above-mentioned coordinate mapping formula can achieve the conversion of coordinates from the cropped image to the original image coordinates. This step is crucial because it ensures that the finally output detection boxes can correctly reflect the true positions of the targets in the original scene.

[0127] In summary, through the effective preprocessing of the pool scene image, the targeted training of the object detection model, and the reasonable image fusion strategy, the accuracy and efficiency of small object detection have been significantly improved, providing strong technical support for swimming pool safety management.

[0128] S150. Perform multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result.

[0129] In this embodiment, the tracking result refers to the tracking trajectory of small targets, etc.

[0130] In one embodiment, please refer to Figure 4 , the above step S150 may include steps S151 to S152.

[0131] S151. Set a first threshold and a second threshold, where the first threshold is greater than the second threshold, the first threshold is used for multi-object tracking; the second threshold is used to compensate for lost target boxes.

[0132] In this embodiment, first, two different levels of thresholds are set for detection, including: a first threshold (high threshold) and a second threshold (low threshold). Among them, the high threshold is used to ensure high-quality object detection results and is applicable to multi-object tracking; while the low threshold relaxes the detection criteria to facilitate capturing more small targets or blurred targets that may be ignored, serving as the basis for compensation tracking.

[0133] S152. Perform multi-object tracking on the results greater than the first threshold in the mapping result. When the tracking is lost in the current frame, start the compensation mechanism, and use IoU matching and the Hungarian algorithm to find matching target boxes from the results lower than the second threshold in the multi-object tracking results to compensate for the lost human tracking box, so as to obtain the tracking result.

[0134] In this embodiment, then, for all detection results higher than the first threshold, an advanced multi-object tracking algorithm (such as ByteTrack) is used for tracking to generate a list of human tracking boxes containing tracking IDs. Whenever the tracking ID of a certain swimmer disappears in the current frame, the compensation mechanism will be started, that is, to find a matching candidate human tracking box from the results lower than the second threshold to compensate for the lost human tracking box.

[0135] In one embodiment, please refer to Figure 5 , the above step S152 may include steps S1521 to S1522.

[0136] S1521. Use a multi-object tracking algorithm to perform multi-object tracking on the results greater than the first threshold in the mapping result to obtain a multi-object tracking result, where the multi-object tracking result includes a list of human tracking boxes with tracking IDs.

[0137] In this embodiment, the multi-object tracking result contains the tracking ID of each swimmer and its corresponding list of human tracking boxes. These human tracking boxes represent the position information of each successfully tracked target in the current frame.

[0138] S1522. Check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find a matching target box from the results lower than the second threshold in the multi-object tracking result to compensate for the lost human tracking box, so as to obtain the tracking result.

[0139] In this embodiment, use a multi-object tracking algorithm to process all detection results exceeding the first threshold to obtain a multi-object tracking result including the tracking ID of each swimmer. Check the status of each tracking object in each frame of the image. If it is found that the tracking ID of a certain swimmer disappears in the current frame, and it is judged according to the mask corresponding to the preset pool range that the swimmer does not swim out of the camera coverage range, it is considered that tracking failure has occurred, and at this time, the compensation mechanism needs to be started.

[0140] Specifically, relatively reliable detection results are screened out through the set high threshold, and these results are passed to the multi-object tracking module based on the ByteTrack algorithm. ByteTrack is an efficient online multi-object tracking algorithm that can handle the object association problem in real-time video streams, thereby generating a stable and accurate list of human tracking bounding boxes.

[0141] In one embodiment, referring to Figure 6 , the above step S1522 may include steps S15221 to S15224.

[0142] S15221. Check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, compare the human tracking bounding box corresponding to the disappeared human tracking bounding box in the multi-object tracking result with all human tracking bounding boxes below the second threshold in the current frame, and calculate the intersection over union (IoU) between the two human tracking bounding boxes to obtain an IoU matrix.

[0143] In this embodiment, as Figure 7 shown, for each frame of image, the algorithm checks whether the tracking ID of the swimmers present in the previous frame still exists in the current frame. If the tracking ID of a certain swimmer disappears, but it is analyzed based on the mask corresponding to the preset pool range and the position information of the front and back frames that the swimmer has not swum out of the camera coverage range, it is inferred that the tracking fails. This kind of failure usually occurs when small targets at a long distance are difficult to be accurately tracked, such as due to environmental factors like reflection or water splash interference.

[0144] Once it is confirmed that tracking failure occurs, the compensation tracking mechanism is activated. At this time, calculate the IoU (intersection over union) between the stable human tracking bounding box previously generated by the high threshold corresponding to the disappeared human tracking bounding box and all human tracking bounding boxes below the low threshold in the current frame, and construct an IoU matrix. IoU is a standard metric for measuring the overlapping degree of two bounding boxes, which helps to identify potential best matching candidate bounding boxes.

[0145] Specifically, the stable human tracking bounding box previously generated by the high threshold corresponding to the disappeared human tracking bounding box and all human detection bounding boxes below the low threshold in the current frame are passed into the detection compensation tracking algorithm module. First, process the human bounding box with tracking failure and perform IoU-based matching with all human detection bounding boxes below the low threshold in the current frame ( Figure 7 in to ) to obtain an IoU matrix , and the specific calculation formula is , where Results indicating the IoU size between the high-threshold human tracking boxes and each low-threshold human detection box, indicating the IoU size between the high-threshold human tracking boxes and each low-threshold human detection box, indicating the area of the high-threshold human tracking boxes, indicating the area of each low-threshold human detection box.

[0146] S15222. Construct a cost matrix based on the IoU matrix.

[0147] In this embodiment, the IoU matrix will be converted into a cost matrix. In this process, a higher IoU value indicates a lower cost (i.e., a better match). This step is to prepare for the subsequent application of the Hungarian algorithm. The specific conversion formula is , where represents a lost human tracking box, represents a low-threshold human detection box.

[0148] S15223. Apply the Hungarian algorithm to the cost matrix to match all the human tracking boxes below the second threshold and the disappeared human tracking boxes in the current frame to obtain a matching result.

[0149] In this embodiment, the matching result refers to the human tracking boxes in the current frame that match the disappeared human tracking boxes among all the human tracking boxes below the second threshold.

[0150] Specifically, use the Hungarian algorithm to perform an optimal matching search on the cost matrix. This algorithm aims to solve the assignment problem on a bipartite graph, and its goal is to minimize the total cost. Therefore, through the Hungarian algorithm, the best matching scheme that minimizes the total cost can be found, that is, a most likely corresponding low-threshold detection box is found for the lost human box .

[0151] S15224. Add the matching result to the multi-object tracking result of the current frame (represented as a set in the figure to ) to obtain a tracking result.

[0152] In this embodiment, finally, apply the best matching relationship found by the Hungarian algorithm to the multi-object tracking result of the current frame to supplement those tracking boxes lost for various reasons. This step not only restores continuity but also improves the robustness and accuracy of the entire tracking system, especially optimized for small target situations where tracking loss is prone to occur.

[0153] In summary, the multi-threshold detection compensation tracking algorithm effectively solves the problem of small target tracking in complex environments by introducing a high-low dual-threshold strategy, combining advanced multi-object tracking technology and an effective compensation mechanism, providing strong technical support for monitoring in specific scenarios such as swimming pools.

[0154] The human tracking box mentioned above is actually a target detection box.

[0155] S160. Output the tracking result.

[0156] Output the tracking result to the terminal for display.

[0157] The method of this embodiment makes full use of the cutting-edge technologies in the fields of computer vision and image processing, aiming to solve the detection and tracking problems of small targets (such as swimmers) in swimming venues caused by factors such as blurring due to distance, water surface reflection, and water splashes. Through the application of this innovative solution, the recall rate for small target detection has been significantly improved, and the common ID switching (id_switch) phenomenon during the tracking process has been effectively reduced.

[0158] This method enhances the ability to identify small targets, ensuring that the position information of swimmers can be accurately captured even under adverse conditions. The optimized algorithm can more stably maintain continuous tracking of each individual, avoiding misidentification of identities caused by environmental interference or target occlusion. It provides reliable data input for the anti-drowning intelligent system, thus assisting in making more accurate safety judgments and emergency responses.

[0159] It can be used in the construction of smart swimming venues. As part of the intelligent infrastructure, this technology helps to build a more efficient and safe service system, promoting the digital transformation of swimming venues. Through the effective monitoring of swimmers' activities, managers can timely discover potential risks and take preventive measures, reducing the probability of drowning accidents. Stable and reliable detection and tracking not only ensure safety but also provide a more comfortable and reassuring sports environment for users.

[0160] In summary, the enhanced method for small target detection and tracking through multi-strategy fusion plays an irreplaceable role in promoting the modernization construction of swimming venues and improving the level of public safety guarantee. It not only improves the operation efficiency of facilities and service quality but also brings a higher sense of security and satisfaction to users. It effectively improves the detection accuracy of small targets and the tracking stability. This method can accurately identify the positions of swimmers in complex pool environments and continuously track their movement trajectories, providing technical support for the safety monitoring, intelligent management, and drowning risk analysis of swimming venues.

[0161] The above multi-strategy fusion method for enhancing small target detection and tracking in swimming stadiums obtains an initial image and inputs it into a target detection model to identify potential small target areas. By cropping and data augmentation of the training dataset, the sensitivity and accuracy of the target detection model for small targets are improved. The detected ROI areas are cropped and enlarged to further enhance the detailed features of small targets and improve the accuracy of subsequent detections. The detection results are mapped back to the coordinate system of the initial image to ensure the accuracy of the target position. Combining multi-threshold detection and compensation tracking strategies, the target is stably tracked to reduce misidentifications and losses, thereby improving the overall tracking stability and accuracy.

[0162] Figure 8 FIG. 4 is a schematic block diagram of a multi-strategy fusion system 300 for enhancing small target detection and tracking in a swimming stadium provided by an embodiment of the present invention. As Figure 8 shown, corresponding to the above multi-strategy fusion method for enhancing small target detection and tracking in a swimming stadium, the present invention also provides a multi-strategy fusion system 300 for enhancing small target detection and tracking in a swimming stadium. The multi-strategy fusion system 300 for enhancing small target detection and tracking in a swimming stadium includes units for executing the above multi-strategy fusion method for enhancing small target detection and tracking in a swimming stadium. The system can be configured in a server. Specifically, please refer to Figure 8 , the multi-strategy fusion system 300 for enhancing small target detection and tracking in a swimming stadium includes an image acquisition unit 301, a target detection unit 302, a re-detection unit 303, a mapping unit 304, a tracking unit 305, and an output unit 306.

[0163] The image acquisition unit 301 is configured to acquire an image to be detected and tracked to obtain an initial image. The target detection unit 302 is configured to input the initial image into a target detection model for target detection to obtain ROI areas. Wherein, the target detection model is obtained by cropping and data augmentation of images with targets in a swimming stadium as a sample set to train a pre-trained detection model. The re-detection unit 303 is configured to crop and enlarge the ROI areas, and input the processed ROI areas into the target detection model again for target detection to obtain detection results. The mapping unit 304 is configured to map the detection results back to the coordinate system corresponding to the initial image to obtain a mapping result. The tracking unit 305 is configured to perform multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result. The output unit 306 is configured to output the tracking result.

[0164] In an embodiment, the multi-strategy fusion system 300 for enhancing small target detection and tracking in a swimming stadium further includes a training unit for:

[0165] Obtain an image with a target in a swimming venue; crop the image using a dynamic sliding window to obtain a cropping result; perform label annotation on the cropping result to generate a training set; perform data augmentation on the training set to obtain a sample set; construct a pre-trained detection model and a loss function; use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0166] In one embodiment, the training unit is further configured to: crop the image using different sliding windows according to the lane area and the non-critical area to obtain a cropping result.

[0167] In one embodiment, the training unit is further configured to: perform target annotation on the cropping result and record the position of the target to obtain a training set.

[0168] In one embodiment, the training unit is further configured to: identify and annotate the head and body features of the swimmers in the training set, and adjust the scaling ratio to enlarge the area extracted by the training set to obtain an enlarged image; use the target tiling algorithm to splice the enlarged images according to the distribution of the targets in the actual scene and update the corresponding annotation information to obtain a synthesis result; perform weighted processing on the small targets in the synthesis result to obtain a weighted processing result; use the pre-trained model to screen out small target instances with confidence meeting the requirements from the weighted processing result to obtain a sample set.

[0169] In one embodiment, the tracking unit 305 includes a setting subunit and a compensation subunit.

[0170] The setting subunit is configured to set a first threshold and a second threshold, where the first threshold is greater than the second threshold, and the first threshold is used for multi-target tracking; the second threshold is used to compensate for the lost target box; the compensation subunit is configured to perform multi-target tracking on the results greater than the first threshold in the mapping result, and when the tracking is lost in the current frame, start a compensation mechanism, and use the IoU matching and the Hungarian algorithm to find a matching target box from the results lower than the second threshold in the multi-target tracking results to compensate for the lost human tracking box to obtain a tracking result.

[0171] In one embodiment, the compensation subunit includes:

[0172] The multi-object tracking module is used to perform multi-object tracking on the results greater than the first threshold in the mapping result by using a multi-object tracking algorithm to obtain a multi-object tracking result, where the multi-object tracking result includes a list of human tracking boxes with tracking IDs; the compensation tracking module is used to check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, the Hungarian algorithm is used to find a matching target box from the results lower than the second threshold in the multi-object tracking result to compensate for the lost human tracking box to obtain a tracking result.

[0173] In one embodiment, the compensation tracking module includes:

[0174] The intersection over union (IoU) calculation sub-module is used to check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, the human tracking box corresponding to the disappeared human tracking box in the multi-object tracking result is compared with all human tracking boxes lower than the second threshold in the current frame, and the intersection over union between the two human tracking boxes is calculated to obtain an IoU matrix; the cost matrix construction sub-module is used to construct a cost matrix according to the IoU matrix; the matching sub-module is used to apply the Hungarian algorithm to the cost matrix to match all human tracking boxes lower than the second threshold in the current frame and the disappeared human tracking box to obtain a matching result; the adding matching sub-module is used to add the matching result to the multi-object tracking result of the current frame to obtain a tracking result.

[0175] It should be noted that those skilled in the art can clearly understand the specific implementation processes of the above-mentioned multi-strategy fusion small target detection and tracking enhancement system 300 for swimming pools and each unit, which can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and conciseness of description, they will not be repeated here.

[0176] The above-mentioned multi-strategy fusion small target detection and tracking enhancement system 300 for swimming pools can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 9 shown.

[0177] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.

[0178] Refer to Figure 9, the computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.

[0179] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to execute an enhanced method for detecting and tracking small targets in a swimming pool venue with multi-strategy fusion.

[0180] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0181] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, it can cause the processor 502 to execute an enhanced method for detecting and tracking small targets in a swimming pool venue with multi-strategy fusion.

[0182] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0183] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps:

[0184] Obtain the image to be detected and tracked to obtain the initial image; input the initial image into the target detection model for target detection to obtain the ROI region; wherein, the target detection model is obtained by cropping and data augmenting the images with targets in the swimming pool venue and using them as a sample set to train the pre-trained detection model; crop and enlarge the ROI region, and input the processed ROI region into the target detection model again for target detection to obtain the detection result; map the detection result back to the coordinate system corresponding to the initial image to obtain the mapping result; perform multi-threshold detection and compensation tracking on the mapping result to obtain the tracking result; output the tracking result.

[0185] In an embodiment, when the processor 502 implements the step that the target detection model is obtained by cropping and data augmenting the images with targets in the swimming pool venue and using them as a sample set to train the pre-trained detection model, the specific implementation steps are as follows:

[0186] Obtain an image with a target in a swimming venue; crop the image using a dynamic sliding window to obtain a cropping result; perform label annotation on the cropping result to generate a training set; perform data augmentation on the training set to obtain a sample set; construct a pre-trained detection model and a loss function; use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0187] In one embodiment, when the processor 502 implements the step of cropping the image using a dynamic sliding window to obtain a cropping result, the following steps are specifically implemented:

[0188] Crop the image using different sliding windows according to the lane area and non-critical areas to obtain a cropping result.

[0189] In one embodiment, when the processor 502 implements the step of performing label annotation on the cropping result to generate a training set, the following steps are specifically implemented:

[0190] Perform target annotation on the cropping result and record the position of the target to obtain a training set.

[0191] In one embodiment, when the processor 502 implements the step of performing data augmentation on the training set to obtain a sample set, the following steps are specifically implemented:

[0192] Identify and annotate the head and body features of the swimmers in the training set, and adjust the scaling ratio to enlarge the extracted area of the training set to obtain an enlarged image; use the target tiling algorithm to splice the enlarged images according to the distribution of the targets in the actual scene and update the corresponding annotation information to obtain a synthesis result; perform weighted processing on the small targets in the synthesis result to obtain a weighted processing result; use the pre-trained model to screen out small target instances with confidence meeting the requirements from the weighted processing result to obtain a sample set.

[0193] In one embodiment, when the processor 502 implements the step of performing multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result, the following steps are specifically implemented:

[0194] Set a first threshold and a second threshold, where the first threshold is greater than the second threshold, the first threshold is used for multi-object tracking; the second threshold is used to compensate for lost target boxes; perform multi-object tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is currently lost in the current frame, start a compensation mechanism, and use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain the tracking result.

[0195] In one embodiment, when the processor 502 implements the step of performing multi-object tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is currently lost in the current frame, start a compensation mechanism, and use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain the tracking result, the specific implementation is as follows:

[0196] Use a multi-object tracking algorithm to perform multi-object tracking on the results in the mapping result that are greater than the first threshold to obtain a multi-object tracking result, where the multi-object tracking result includes a list of human tracking boxes with tracking IDs; check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain the tracking result.

[0197] In one embodiment, when the processor 502 implements the step of checking the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain the tracking result, the specific implementation is as follows:

[0198] Check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, compare the human tracking box corresponding to the disappeared human tracking box in the multi-object tracking result with all the human tracking boxes in the current frame that are lower than the second threshold, and calculate the intersection over union between the two human tracking boxes to obtain an IoU matrix; construct a cost matrix according to the IoU matrix; apply the Hungarian algorithm to the cost matrix to match all the human tracking boxes in the current frame that are lower than the second threshold and the disappeared human tracking box to obtain a matching result; add the matching result to the multi-object tracking result of the current frame to obtain the tracking result.

[0199] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0201] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor executes the following steps:

[0202] Obtain the image to be detected and tracked to obtain an initial image; input the initial image into a target detection model for target detection to obtain an ROI region; wherein, the target detection model is obtained by cropping and data augmentation of the images with targets in the swimming pool hall as a sample set to train a pre-trained detection model; crop and enlarge the ROI region, and input the processed ROI region into the target detection model again for target detection to obtain a detection result; map the detection result back to the coordinate system corresponding to the initial image to obtain a mapping result; perform multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result; output the tracking result.

[0203] In one embodiment, when the processor executes the computer program to implement the step that the target detection model is obtained by cropping and data augmentation of the images with targets in the swimming pool hall as a sample set to train a pre-trained detection model, the following steps are specifically implemented:

[0204] Obtain an image with a target in a swimming venue; crop the image using a dynamic sliding window to obtain a cropping result; perform label annotation on the cropping result to generate a training set; perform data augmentation on the training set to obtain a sample set; construct a pre-trained detection model and a loss function; use the sample set to train the pre-trained detection model, and determine the target detection model in combination with the loss function and the trained model.

[0205] In one embodiment, when the processor executes the computer program to implement the step of cropping the image using a dynamic sliding window to obtain a cropping result, the following steps are specifically implemented:

[0206] Crop the image using different sliding windows according to the lane area and non-critical area to obtain a cropping result.

[0207] In one embodiment, when the processor executes the computer program to implement the step of performing label annotation on the cropping result to generate a training set, the following steps are specifically implemented:

[0208] Perform target annotation on the cropping result and record the position of the target to obtain a training set.

[0209] In one embodiment, when the processor executes the computer program to implement the step of performing data augmentation on the training set to obtain a sample set, the following steps are specifically implemented:

[0210] Identify and annotate the head and body features of the swimmers in the training set, and adjust the scaling ratio to enlarge the area extracted from the training set to obtain an enlarged image; use the target tiling algorithm to splice the enlarged images according to the distribution of the targets in the actual scene and update the corresponding annotation information to obtain a synthesis result; perform weighted processing on the small targets in the synthesis result to obtain a weighted processing result; use the pre-trained model to screen out small target instances with confidence meeting the requirements from the weighted processing result to obtain a sample set.

[0211] In one embodiment, when the processor executes the computer program to implement the step of performing multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result, the following steps are specifically implemented:

[0212] Set a first threshold and a second threshold, where the first threshold is greater than the second threshold, the first threshold is used for multi-object tracking; the second threshold is used to compensate for lost target boxes; perform multi-object tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is currently lost in the current frame, start a compensation mechanism, and use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result.

[0213] In one embodiment, when the processor executes the computer program to implement the step of performing multi-object tracking on the results in the mapping result that are greater than the first threshold, and when the tracking is currently lost in the current frame, start a compensation mechanism, and use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result, the specific implementation is as follows:

[0214] Use a multi-object tracking algorithm to perform multi-object tracking on the results in the mapping result that are greater than the first threshold to obtain a multi-object tracking result, where the multi-object tracking result includes a list of human tracking boxes with tracking IDs; check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result.

[0215] In one embodiment, when the processor executes the computer program to implement the step of checking the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, use IoU matching and the Hungarian algorithm to find matching target boxes from the results in the multi-object tracking result that are lower than the second threshold to compensate for the lost human tracking box to obtain a tracking result, the specific implementation is as follows:

[0216] Check the status of each tracking object frame by frame. When it is found that the tracking ID of a certain swimmer disappears in the current frame and a certain swimmer does not swim out of the range, compare the human tracking box corresponding to the disappeared human tracking box in the multi-object tracking result with all the human tracking boxes in the current frame that are lower than the second threshold, and calculate the intersection over union between the two human tracking boxes to obtain an IoU matrix; construct a cost matrix according to the IoU matrix; apply the Hungarian algorithm to the cost matrix to match all the human tracking boxes in the current frame that are lower than the second threshold and the disappeared human tracking box to obtain a matching result; add the matching result to the multi-object tracking result of the current frame to obtain a tracking result.

[0217] The storage medium may be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disk, or other computer-readable storage media that can store program codes.

[0218] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0219] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0220] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0221] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0222] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multi-strategy fusion method for small target detection and tracking enhancement in swimming pools, characterized in that: include: Acquire the image to be detected and tracked to obtain an initial image; Input the initial image into the target detection model for target detection to obtain the ROI area; wherein the target detection model is obtained by obtaining images with targets in the swimming pool, cropping them, and performing data enhancement to train the pre-trained detection model as a sample set; The ROI region is cropped and enlarged, and the processed ROI region is input into the target detection model again for target detection to obtain a detection result; Mapping the detection result back to the coordinate system corresponding to the initial image to obtain a mapping result; Using multi-threshold detection and compensation tracking on the mapping result to obtain a tracking result; Outputting the tracking result; The mapping result is subjected to multi-threshold detection and compensation tracking to obtain a tracking result, including: Setting a first threshold and a second threshold, wherein the first threshold is greater than the second threshold, the first threshold is used for multi-target tracking; and the second threshold is used to compensate for lost target frames; Perform multi-target tracking on the results greater than the first threshold in the mapping results, and when the tracked target is lost in the current frame, start the compensation mechanism, use IoU matching and the Hungarian algorithm to find a matching target frame from the results less than the second threshold in the multi-target tracking results, so as to compensate for the lost human tracking frame, so as to obtain a tracking result; The multi-target tracking is performed on the results greater than the first threshold in the mapping results, and when the tracking target is lost in the current frame, a compensation mechanism is started, and a matching target frame is found from the results less than the second threshold in the multi-target tracking results by using IoU matching and the Hungarian algorithm to compensate for the lost human tracking frame, so as to obtain a tracking result, including: Performing multi-target tracking on the results greater than the first threshold in the mapping results using a multi-target tracking algorithm to obtain a multi-target tracking result, wherein the multi-target tracking result includes a list of human body tracking frames of tracking IDs; The status of each tracked object is checked frame by frame. When it is found that the tracking ID of a swimmer disappears in the current frame and a swimmer has not swam out of the range, the IoU matching and Hungarian algorithm are used to find the matching target frame from the results below the second threshold in the multi-target tracking results to compensate for the lost human tracking frame and obtain the tracking result.

2. According to the multi-strategy fusion swimming pool small target detection and tracking enhancement method of claim 1, it is characterized in that: The target detection model is obtained by obtaining images with targets in the swimming pool, cropping them, and performing data enhancement as a sample set to train a pre-trained detection model, including: Acquire images with targets in the swimming pool; Cropping the image using a dynamic sliding window to obtain a cropping result; Labeling the cropping results to generate a training set; Performing data enhancement on the training set to obtain a sample set; Build a pre-trained detection model and loss function; The pre-trained detection model is trained using the sample set, and the target detection model is determined by combining the loss function and the trained model.

3. The multi-strategy fusion swimming pool small target detection and tracking enhancement method according to claim 2 is characterized in that: The step of cropping the image using a dynamic sliding window to obtain a cropping result includes: The image is cropped using different sliding windows according to the lane area and the non-critical area to obtain a cropping result.

4. The multi-strategy fusion swimming pool small target detection and tracking enhancement method according to claim 2 is characterized in that: The step of labeling the cropping results to generate a training set includes: The cropping results are labeled with objects, and the positions of the objects are recorded to obtain a training set.

5. The multi-strategy fusion swimming pool small target detection and tracking enhancement method according to claim 2 is characterized in that: The data enhancement of the training set to obtain a sample set includes: Identifying and labeling the head and body features of the swimmers in the training set, and adjusting the scaling to magnify the region extracted from the training set to obtain an enlarged image; Using a target tiling algorithm to stitch the enlarged images according to the distribution of targets in the actual scene, and updating corresponding annotation information to obtain a synthesis result; Performing weighted processing on the small targets in the synthesis result to obtain a weighted processing result; The pre-trained model is used to filter out small target instances whose confidence meets the requirements from the weighted processing results to obtain a sample set.

6. The multi-strategy fusion swimming pool small target detection and tracking enhancement method according to claim 5 is characterized in that: The state of each tracked object is checked frame by frame. When it is found that the tracking ID of a swimmer disappears in the current frame and the swimmer has not swam out of the range, the matching target frame is found from the results below the second threshold in the multi-target tracking results by using IoU matching and the Hungarian algorithm to compensate for the lost human tracking frame, so as to obtain the tracking result, including: Check the status of each tracked object frame by frame. When it is found that the tracking ID of a swimmer disappears in the current frame and the swimmer has not swam out of the range, compare the corresponding human tracking frame in the multi-target tracking result corresponding to the disappeared human tracking frame with all human tracking frames in the current frame that are lower than the second threshold, and calculate the intersection over union ratio between the two human tracking frames to obtain the IoU matrix. Construct a cost matrix based on the IoU matrix; Applying the Hungarian algorithm to the cost matrix to match all human tracking frames below the second threshold and disappeared human tracking frames in the current frame to obtain a matching result; The matching result is added to the multi-target tracking result of the current frame to obtain the tracking result.

7. The multi-strategy fusion small target detection and tracking enhancement system for swimming pools is characterized by: The system uses the method according to any one of claims 1 to 6, including: An image acquisition unit, used to acquire an image to be detected and tracked to obtain an initial image; A target detection unit, used for inputting the initial image into a target detection model for target detection to obtain a ROI area; wherein the target detection model is obtained by obtaining an image with a target in a swimming pool, cropping it, and performing data enhancement to train a pre-trained detection model as a sample set; A re-detection unit, used for cropping and enlarging the ROI area, and inputting the processed ROI area into the target detection model again for target detection to obtain a detection result; A mapping unit, used for mapping the detection result back to the coordinate system corresponding to the initial image to obtain a mapping result; A tracking unit, used for applying multi-threshold detection and compensation tracking to the mapping result to obtain a tracking result; An output unit is used to output the tracking result.

8. The multi-strategy fusion swimming pool small target detection and tracking enhancement system according to claim 7 is characterized in that: Also included are training units for: Acquire images with targets in the swimming pool; Cropping the image using a dynamic sliding window to obtain a cropping result; Labeling the cropping results to generate a training set; Performing data enhancement on the training set to obtain a sample set; Build a pre-trained detection model and loss function; The pre-trained detection model is trained using the sample set, and the target detection model is determined by combining the loss function and the trained model.

Citation Information

Patent Citations

  • Digital pathological image analysis method, system and device and storage medium

    CN112419253A

  • Cross-scene target automatic identification and tracking method and application thereof

    CN112801018A

  • Multi-target tracking method for natatorium scene

    CN116342645A

  • Multi-target detection tracking statistical method and device based on edge calculation

    CN118968044A

  • KR20240172475A