Swimming pool anti-drowning head tracking methods, devices, computer equipment, and storage media
By combining a deep learning network with FAM to form a tracking model, the problem of decreased detection accuracy caused by perspective distortion in pool images was solved, and accurate detection and tracking of small targets such as swimmers' heads were achieved, thus improving detection and tracking accuracy.
Patent Information
- Application Number
- CN202310125387.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-02-16
AI Technical Summary
In existing technologies, perspective distortion in underwater camera images within swimming pools leads to a decrease in the accuracy of head target detection and tracking, making it difficult to accurately detect and track small targets such as swimmers' heads.
A tracking model combining a deep learning network with FAM is adopted. By using a single-stage detector and a detection box-level attention mechanism, combined with a multi-scale attention mechanism of multi-scale features and semantic segmentation, the problem of image perspective distortion is solved, and the accurate detection and tracking of human heads is achieved.
It effectively solves the problem of image perspective distortion, enables accurate detection and tracking of small targets such as swimmers' heads, and improves the accuracy of target detection and tracking.
Smart Images

Figure CN116402847B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a target tracking method, and more specifically to a method, apparatus, computer equipment, and storage medium for tracking heads on the surface of a swimming pool to prevent drowning. Background Technology
[0002] With the development of sports, people's enthusiasm for participating in sports activities is increasing. However, swimming, as one of the most popular sports, has become the sport with the highest incidence of safety accidents.
[0003] Existing technologies use underwater cameras installed around the perimeter and bottom of swimming pools to determine, through algorithms, whether swimmers are swimming normally or struggling and drowning. However, the wide-angle lenses of these pool cameras produce images with severe perspective distortion, posing a significant challenge to real-time detection and tracking. Severe perspective distortion causes geometric misalignment of targets in the image, resulting in blurring. Information and details may be lost due to resolution changes or too much information crammed into a single pixel. When detecting heads on the surface of a pool, severe perspective distortion can easily lead to errors in head target information, blurred head pixels, and loss of crucial head information, thus affecting the accuracy of target detection and tracking.
[0004] Therefore, it is necessary to design a new method to solve the perspective distortion problem of images and achieve accurate detection and tracking of small targets such as swimmers' heads. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, device, computer equipment and storage medium for tracking heads on the surface of swimming pools to prevent drowning.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a swimming pool anti-drowning head tracking method, comprising:
[0007] Acquire images captured by cameras inside the swimming pool;
[0008] The image is input into the tracking model for head detection and tracking to obtain the tracking result;
[0009] The tracking model is obtained by combining a deep learning network with a sample set of images containing true bounding boxes of human heads.
[0010] The further technical solution is as follows: the tracking model is a model formed by combining a single-stage detector and a detection box-level attention mechanism.
[0011] The further technical solution is as follows: the tracking model includes five detection layers, and each detection layer has a detection frame of a corresponding size.
[0012] The further technical solution is as follows: the aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062 mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
[0013] The further technical solution is as follows: during the training of the tracking model, when the ratio of the IOU between the detection box and one of the real head boxes is not less than a first threshold, the detection box is used for head prediction; when the ratio of the IOU between the detection box and all real head boxes is less than a second threshold, the detection box is determined to be a background detection box, and the remaining detection boxes do not participate in the training of the tracking model.
[0014] The further technical solution is as follows: the tracking model adds a FAM branch to the deep learning network used for head detection; the labeled value of the FAM branch is the filling result of the head detection box, and a hierarchical attention image is used.
[0015] The further technical solution is as follows: the loss value in the training process of the tracking model includes the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
[0016] The present invention also provides a swimming pool anti-drowning head tracking device, characterized in that it includes:
[0017] An image acquisition unit is used to acquire images captured by a camera inside the swimming pool.
[0018] The tracking unit is used to input the image into the tracking model for head detection and tracking to obtain the tracking result. The tracking model is obtained by combining a deep learning network with a sample set of images with ground truth bounding box labels for heads.
[0019] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the above-described method.
[0020] The present invention also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0021] The beneficial effects of this invention compared with the prior art are as follows: This invention uses a tracking model to track images captured by a camera in a swimming pool. The tracking model is formed by training a deep learning network combined with FAM. The entire tracking model has multi-scale features, multi-scale detection boxes, and a multi-scale attention mechanism based on semantic segmentation, which can solve the perspective distortion problem of the image and achieve accurate detection and tracking of small targets such as swimmers' heads.
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A schematic diagram illustrating an application scenario of the swimming pool anti-drowning head tracking method provided in an embodiment of the present invention;
[0025] Figure 2 A schematic flowchart of the swimming pool anti-drowning head tracking method provided in an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of the tracking model provided in an embodiment of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of the attention function provided in an embodiment of the present invention;
[0028] Figure 5 A schematic block diagram of a swimming pool anti-drowning head tracking device provided in an embodiment of the present invention;
[0029] Figure 6 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0032] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0033] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0034] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the swimming pool anti-drowning head tracking method provided in an embodiment of the present invention. Figure 2 This is a schematic flowchart illustrating the swimming pool head tracking method for preventing drowning provided in this embodiment of the invention. The method is applied in a server. The server interacts with the terminal, employing a deep learning model and fusing lens distortion and spherical distortion correction modules to remove distortion, achieving accurate detection and tracking of small targets such as swimmers' heads. This addresses the perspective distortion problem in images and enables precise detection and tracking of small targets such as swimmers' heads.
[0035] Figure 2 This is a flowchart illustrating the swimming pool head tracking method for preventing drowning provided in an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S120.
[0036] S110, Acquire images captured by the camera inside the swimming pool.
[0037] In this embodiment, the image refers to a photograph taken by a camera positioned inside the swimming pool.
[0038] S120. Input the image into the tracking model for head detection and tracking to obtain the tracking result.
[0039] In this embodiment, a robust tracking model is first learned from a set of image samples with ground truth bounding box labels using a deep network. This learned tracking model is then used to track a person's head and obtain a tracking route. This embodiment will be described in two parts: first, the tracking model will be introduced, followed by an explanation of the training method.
[0040] For specific tracking models, please refer to Figure 4 The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism. The bounding box-level attention mechanism serves three purposes: to achieve multi-scale face detection; to emphasize face regions in the image while downplaying background regions; and to generate more occluded faces for training.
[0041] This tracking model incorporates two attention mechanisms: LAM (Line Attention Module) and FAM (Face Attention Module). Through adaptive learning of the sample set, it transforms the features of lines and heads respectively to compensate for information loss caused by distortions such as linear distortion and spherical distortion in the image. Both attention mechanisms use the same network structure; the difference lies in that LAM takes the swimmer's body features in the water as input, while FAM takes the swimmer's facial features as input.
[0042] Taking FAM as an example, this embodiment utilizes feature pyramids for multi-scale face detection. The specific implementation of FAM involves applying attention modules to image features at each scale to enhance the features of the face region and improve the accuracy of face region recognition.
[0043] In addition, the tracking model includes five detection layers, each with a detection box of a corresponding size.
[0044] In this embodiment, the aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062 mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
[0045] The training process of the tracking model described in this embodiment is as follows: In calculating the loss between the model's predicted bounding box and the ground truth bounding box, if the ratio of the IOU between the detected bounding box and one of the ground truth bounding boxes of a human head is not less than a first threshold, then the detected bounding box is used for human head prediction; if the ratio of the IOU between the detected bounding box and all ground truth bounding boxes of a human head is less than a second threshold, then the detected bounding box is determined to be a background detected bounding box, and the remaining detected bounding boxes do not participate in the training of the tracking model.
[0046] In FAM, this embodiment uses the filled results of head detection boxes as the ground truth and the prediction results of the tracking model to calculate the loss. Detection boxes of corresponding scales are set for each of the five detection layers. The aspect ratio of each detection box is set to either 1 or 1.5, because the aspect ratio of a frontal face is close to 1, and the aspect ratio of a side face is close to 1.5. The size of the detection boxes is set between 162 and 4062, and the size of the detection boxes in each layer is increased by a factor of 21 / 3. This dense set of detection boxes ensures that each ground truth box has a corresponding detection box with an IOU greater than 0.6.
[0047] If a detection box has the largest IOU with a ground truth bounding box that is greater than 0.5, then that detection box is assigned to predict the face. If the largest IOU between a detection box and all ground truth bounding boxes is less than 0.4, then it is designated as a background detection box and is not assigned to face prediction. The remaining detection boxes are not included in the training process.
[0048] In this embodiment, please refer to Figure 4 The loss values in the training process of the tracking model include the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
[0049] Loss function: Where k represents the level of the pyramid, ; Indicates the pyramid layers. The detection frame is set; , These represent the confidence level of each predicted detection box containing a face and its actual label, respectively. This represents the learned and labeled coordinate values of each detection box.
[0050] The loss consists of the classification loss of each detection box, the smooth L1 loss of the coordinates of the positive detection boxes, and the sigmoid cross-entropy loss per pixel, which is used as the loss for attention learning.
[0051] The entire process and details of adding LAM to the deep learning network described above can be found in the process of adding FAM to the deep learning network described above, and will not be repeated here.
[0052] The above-mentioned swimming pool anti-drowning head tracking method tracks images captured by cameras in the pool using a tracking model. The tracking model is trained using a deep learning network combined with FAM. The entire tracking model has multi-scale features, multi-scale detection boxes, and a multi-scale attention mechanism based on semantic segmentation, which can solve the perspective distortion problem of the image and achieve accurate detection and tracking of small targets such as swimmers' heads.
[0053] Figure 5This is a schematic block diagram of a swimming pool anti-drowning head tracking device 300 provided in an embodiment of the present invention. Figure 5 As shown, corresponding to the above-described swimming pool anti-drowning head tracking method, the present invention also provides a swimming pool anti-drowning head tracking device 300. This swimming pool anti-drowning head tracking device 300 includes a unit for performing the above-described swimming pool anti-drowning head tracking method, and the device can be configured in a server. Specifically, please refer to... Figure 5 The swimming pool anti-drowning head tracking device 300 includes an image acquisition unit 301 and a tracking unit 302.
[0054] Image acquisition unit 301 is used to acquire images captured by a camera in the swimming pool; tracking unit 302 is used to input the images into a tracking model for head detection and tracking to obtain tracking results, wherein the tracking model is obtained by combining a deep learning network of FAM with a sample set of images with ground truth bounding box labels for heads.
[0055] The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism.
[0056] The tracking model includes five detection layers, each with a detection box of a corresponding size.
[0057] The aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062 mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
[0058] During the training of the tracking model, if the IOU ratio between the detection box and one of the ground truth head boxes is not less than a first threshold, the detection box is used for head prediction; if the IOU ratio between the detection box and all ground truth head boxes is less than a second threshold, the detection box is determined to be a background detection box, and the remaining detection boxes are not included in the training of the tracking model.
[0059] The tracking model adds a FAM branch to the deep learning network used for head detection; the labeled value of the FAM branch is the filling result of the head detection box, and a hierarchical attention image is used.
[0060] The loss values in the training process of the tracking model include the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
[0061] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned swimming pool anti-drowning head tracking device 300 and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0062] The aforementioned swimming pool anti-drowning head tracking device 300 can be implemented as a computer program, which can, for example... Figure 6 It runs on the computer device shown.
[0063] Please see Figure 6 , Figure 6 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0064] See Figure 6 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0065] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a swimming pool anti-drowning head tracking method.
[0066] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0067] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a swimming pool anti-drowning water surface head tracking method.
[0068] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0069] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:
[0070] Acquire images captured by cameras inside the swimming pool; input the images into a tracking model for head detection and tracking to obtain tracking results;
[0071] The tracking model is obtained by combining a deep learning network with a sample set of images containing ground truth bounding boxes of human heads.
[0072] The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism.
[0073] The tracking model includes five detection layers, each with a detection box of a corresponding size.
[0074] The aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062 mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
[0075] During the training of the tracking model, if the IOU ratio between the detection box and one of the ground truth head boxes is not less than a first threshold, the detection box is used for head prediction; if the IOU ratio between the detection box and all ground truth head boxes is less than a second threshold, the detection box is determined to be a background detection box, and the remaining detection boxes are not included in the training of the tracking model.
[0076] The tracking model adds a FAM branch to the deep learning network used for head detection; the labeled value of the FAM branch is the filling result of the head detection box, and a hierarchical attention image is used.
[0077] The loss values in the training process of the tracking model include the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
[0078] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0079] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0080] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:
[0081] Acquire images captured by cameras inside the swimming pool; input the images into a tracking model for head detection and tracking to obtain tracking results;
[0082] The tracking model is obtained by combining a deep learning network with a sample set of images containing ground truth bounding boxes of human heads.
[0083] The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism.
[0084] The tracking model includes five detection layers, each with a detection box of a corresponding size.
[0085] The aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062 mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
[0086] During the training of the tracking model, if the IOU ratio between the detection box and one of the ground truth head boxes is not less than a first threshold, the detection box is used for head prediction; if the IOU ratio between the detection box and all ground truth head boxes is less than a second threshold, the detection box is determined to be a background detection box, and the remaining detection boxes are not included in the training of the tracking model.
[0087] The tracking model adds a FAM branch to the deep learning network used for head detection; the labeled value of the FAM branch is the filling result of the head detection box, and a hierarchical attention image is used.
[0088] The loss values in the training process of the tracking model include the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
[0089] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0090] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0091] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0092] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0093] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0094] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for tracking heads on the surface of a swimming pool to prevent drowning, characterized in that, include: Acquire images captured by cameras inside the swimming pool; The image is input into the tracking model for head detection and tracking to obtain the tracking result; The tracking model is obtained by combining a deep learning network called FAM with a sample set of images containing real-world head bounding boxes. The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism. The tracking model includes five detection layers, each with a detection box of a corresponding size. The aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
2. The swimming pool anti-drowning head tracking method according to claim 1, characterized in that, During the training of the tracking model, when the ratio of the Intersection over Union (IOU) between the detection box and one of the ground truth head boxes is not less than a first threshold, the detection box is used for head prediction. If the IOU ratio between the detected bounding box and all ground truth bounding boxes is less than the second threshold, the detected bounding box is determined to be a background detected bounding box, and the remaining detected bounding boxes are not included in the training of the tracking model.
3. The swimming pool anti-drowning head tracking method according to claim 1, characterized in that, The tracking model adds a FAM branch to the deep learning network used for head detection; the labeled value of the FAM branch is the filling result of the head detection box, and a hierarchical attention image is used.
4. The swimming pool anti-drowning head tracking method according to claim 1, characterized in that, The loss values in the training process of the tracking model include the classification loss value of each detection box, the coordinate loss value of the positive detection box, and the sum of the sigmoid cross-entropy loss for each pixel.
5. A swimming pool anti-drowning water surface head tracking device, characterized in that, include: An image acquisition unit is used to acquire images captured by a camera inside the swimming pool. The tracking unit is used to input the image into the tracking model for head detection and tracking to obtain the tracking result. The tracking model is obtained by combining a deep learning network with a sample set of images with ground truth bounding box labels of heads. The tracking model is a combination of a single-stage detector and a bounding box-level attention mechanism. The tracking model includes five detection layers, each with a detection box of a corresponding size. The aspect ratio of the detection frame is set to 1 and 1.5; the size of the detection frame is 162mm. 2 Up to 4062mm 2 The size of the detection frame for each detection layer is increased by a factor of 2. 1 / 3 .
6. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Swimming pool drowning prevention supervision method and device, computer equipment and storage medium
CN114022910A
Pedestrian detection method of Faster R-CNN network based on improved clustering algorithm
CN114332921A