Swimming pool drowning detection method based on human skeleton key points
Through the drowning detection method based on key points of human skeletons, combined with human bounding box detection, multi-object tracking and behavior recognition model, the problems of low recognition accuracy and misjudgment in computer vision drowning monitoring methods are solved, and high accuracy and efficient drowning detection are achieved.
Patent Information
- Application Number
- CN202510184648.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
AI Technical Summary
The drowning monitoring method based on computer vision has problems of low recognition accuracy and misjudgment, especially when facing problems such as underwater human body occlusion and turbid water quality in indoor swimming pools.
A drowning detection method based on key points of human bones is adopted, and videos are obtained through underwater monitoring equipment and frame images are extracted. The human bounding box and bone key point detection model are used to detect human body position and bone key point key points. Combined with multi-objective tracking algorithm and behavior recognition model, the key point position flow and motion information flow are analyzed to judge drowning situation.
It improves the accuracy and efficiency of drowning detection, reduces misjudgment and misjudgment, is suitable for various swimming scenarios, and can run unattended for a long time.
Smart Images

Figure CN120108002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision monitoring, and in particular to a swimming pool drowning detection method based on key points of human skeleton. Background Art
[0002] Drowning accidents during swimming are safety accidents caused by many unforeseen factors such as accidental collisions, sudden illnesses or physical overdrafts. By monitoring the behavior of swimmers and timely detecting drowning behavior, drowning accidents can be effectively avoided. Existing methods for monitoring drowning behavior mainly use manual monitoring methods. Since this method requires safety officers to observe whether swimmers show signs of drowning through naked eyes or monitoring equipment, it is a challenge for safety officers to maintain high concentration for a long time. In complex multi-person scenes, safety officers need to monitor multiple swimmers or multiple screens at the same time, making it difficult to ensure that every swimmer is paid attention to. Therefore, this method has many uncertainties and often leads to missed detections. For this reason, a monitoring method based on computer vision is proposed. It sends a single frame image into the target detection model to identify the swimmer's position to generate a bounding box, and obtains the swimmer's current posture through the posture estimation model. It detects whether the posture meets the drowning characteristics to determine whether the swimmer is drowning, thereby realizing automatic monitoring of drowning. This method can be integrated with the existing surveillance camera system, has the characteristics of high efficiency and automation, and can significantly improve the accuracy and timeliness of drowning monitoring without increasing additional labor costs. Although the drowning monitoring method based on computer vision has many advantages, the current technology still faces the following problems: First, drowning is a gradual process. It is difficult to accurately judge the behavior status of the swimmer by relying on a single frame or a few frames of images, which is prone to missed or misjudgment. Second, the current target detection model has limitations, especially when facing underwater human occlusion and turbid water problems common in indoor swimming pools, the recognition accuracy of the swimmer's movement status is low. Summary of the invention
[0003] The present invention aims to solve the problem of low recognition accuracy and misjudgment in drowning monitoring methods based on computer vision, and provides a swimming pool drowning detection method based on key points of human skeleton.
[0004] To solve the above problems, the present invention is achieved through the following technical solutions:
[0005] A method for detecting drowning in a swimming pool based on key points of human skeleton, comprising the following steps:
[0006] Step 1: Use the underwater monitoring equipment of the swimming pool to obtain underwater video, and extract a predetermined number of frame images from the underwater video at fixed time intervals, and pre-process these frame images to obtain underwater images, wherein each underwater image has a timestamp;
[0007] Step 2: Each underwater image with a timestamp is sent to a human body bounding box and skeleton key point detection model, which outputs a detection result of each underwater image with a timestamp, wherein the detection result includes a human body bounding box and confidence, and a skeleton key point and the confidence of the corresponding point;
[0008] Step 3: Based on the human body bounding box in the detection result of each underwater image with a timestamp, a multi-target tracking algorithm is used to perform target allocation and trajectory tracking, and a tracking ID is assigned to each detected target; then, the skeleton key points in the detection result of each underwater image with a timestamp are associated with the tracking ID, and organized into a target time sequence key point sequence according to the time order of the corresponding timestamp;
[0009] Step 4: Process each target time-series key point sequence into a key point position stream and a motion information stream, and simultaneously feed the key point position stream and the motion information stream into the behavior recognition model. The behavior recognition model outputs the classification result to obtain the swimmer's final dangerous or normal behavior classification probability.
[0010] The preprocessing of each image in the above step 1 includes resizing and filling processing, color and channel processing, and dimension and type adjustment processing.
[0011] In the above step 2, the human body bounding box and skeleton key point detection model consists of a backbone network, a neck network and a head network.
[0012] The above-mentioned backbone network includes 1 convolution layer, 4 lightweight convolution modules, 1 spatial pyramid pooling module, and 1 feature extraction module fused with hybrid local channel attention; wherein 1 convolution layer, 4 lightweight convolution modules, 1 spatial pyramid pooling module and 1 feature extraction module fused with hybrid local channel attention are connected in sequence; the input of the convolution layer forms the input of the backbone network, that is, the input of the human body bounding box and skeleton key point detection model; the output of the feature extraction module fused with hybrid local channel attention forms the first output of the backbone network, the output of the third lightweight convolution module forms the second output of the backbone network, and the output of the second lightweight convolution module forms the third output of the backbone network.
[0013] The above-mentioned neck network includes 4 splicing layers, 5 grouped spatial convolution modules, 4 cross-stage partial network modules, and 2 upsampling layers; wherein the input of the first grouped spatial convolution module forms the first input of the neck network, which is connected to the first output of the backbone network; an input of the first splicing layer forms the second input of the neck network, which is connected to the second output of the backbone network; an input of the second splicing layer forms the third input of the neck network, which is connected to the third output of the backbone network; the output of the first grouped spatial convolution module is connected to another input of the first splicing layer via the first upsampling layer, the output of the first splicing layer is connected to the input of the second grouped spatial convolution module via the first cross-stage partial network module, the output of the second grouped spatial convolution module is connected to another input of the second splicing layer via the second upsampling layer, and the output of the second splicing layer is connected to the second cross-stage via the third grouped spatial convolution module. The input of the partial network module, the output of the second cross-stage partial network module is connected to the input of the fourth grouped spatial convolution module; the output of the fourth grouped spatial convolution module and the output of the second grouped spatial convolution module are simultaneously connected to the input of the third grouped spatial convolution module, the output of the third grouped spatial convolution module is connected to the input of the third cross-stage partial network module, and the output of the third cross-stage partial network module is connected to the input of the fifth grouped spatial convolution module; the output of the fifth grouped spatial convolution module and the output of the first grouped spatial convolution module are simultaneously connected to the input of the fourth splicing layer, and the output of the fourth splicing layer is connected to the input of the fourth cross-stage partial network module; the output of the fourth cross-stage partial network module forms the first output of the neck network, the output of the third cross-stage partial network module forms the second output of the neck network, and the output of the second cross-stage partial network module forms the third output of the neck network.
[0014] The head network includes three detection heads, one bounding box selection module and one key point selection module; the first detection head forms the first input end of the head network and is connected to the first output of the neck network; the second detection head forms the second input end of the head network and is connected to the second output of the neck network; the third detection head forms the third input end of the head network and is connected to the third output of the neck network; the bounding box outputs of the three detection heads are simultaneously connected to the input of the bounding box selection module, and the output of the bounding box selection module forms the bounding box output of the head network, that is, the bounding box output of the human body bounding box and the skeleton key point detection model, and the key point outputs of the three detection heads are simultaneously connected to the input of the key point selection module, and the output of the key point selection module forms the key point output of the head network, that is, the key point output of the human body bounding box and the skeleton key point detection model.
[0015] The feature extraction module fused with mixed local channel attention includes 2 convolutional layers, N pyramid slice attention modules, and 1 splicing layer, where N is a set value. Each pyramid slice attention module includes 1 mixed local channel attention module, 2 convolutional layers, and 2 addition layers; wherein the input of the mixed local channel attention module forms the input of the pyramid slice attention module, the input and output of the fused mixed local channel attention are simultaneously connected to the input of the first addition layer, the output of the first addition layer is connected to the input of the first convolutional layer, the output of the first convolutional layer is connected to the input of the second convolutional layer, the output of the second convolutional layer and the input of the first convolutional layer are simultaneously connected to the input of the second addition layer, and the output of the second addition layer forms the output of the pyramid slice attention module. The input of the first convolutional layer forms the input of the feature extraction module fused with mixed local channel attention, the N pyramid slice attention modules are connected in sequence, the input of the first pyramid slice attention module is connected to the output of the first convolutional layer, the output of the last pyramid slice attention module and the output of the first convolutional layer are simultaneously connected to the input of the splicing layer, the output of the splicing layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer forms the output of the feature extraction module fused with mixed local channel attention.
[0016] In the above step 2, there are 17 bone key points, namely nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle.
[0017] In the above step 3, the multi-target tracking algorithm is the ByteTrack multi-target tracking algorithm, that is:
[0018] The detection results of underwater images of all timestamps are processed in order from the front to the back:
[0019] The following operations are performed on the detection results of the underwater image at the initial timestamp:
[0020] ① Extract bounding boxes and confidences from the detection results of the underwater image at the initial timestamp, and divide the bounding boxes into two categories by comparing them with the preset confidence threshold: when the confidence of the bounding box is greater than or equal to the confidence threshold, it is defined as a high-confidence bounding box; when the confidence of the bounding box is less than the confidence threshold, it is defined as a low-confidence bounding box;
[0021] ② Initialize each high-confidence bounding box in the underwater image at the initial timestamp as a tracking instance of each target trajectory, assign a unique tracking ID to each, and initialize the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the initial timestamp;
[0022] The following operations are performed on the detection results of the underwater images with the remaining time stamps after the initial time stamp:
[0023] ① The current Kalman filter predicts the position of each target track's tracking instance in the underwater image at the current timestamp based on the tracking instance state of each target track in the underwater image at the previous timestamp, and obtains the predicted bounding box of each target track;
[0024] ② Extract bounding boxes and confidences from the detection results of the underwater image at the current timestamp, and divide the bounding boxes into two categories by comparing them with the preset confidence threshold: when the confidence of the bounding box is greater than or equal to the confidence threshold, it is defined as a high-confidence bounding box; when the confidence of the bounding box is less than the confidence threshold, it is defined as a low-confidence bounding box;
[0025] ③ For each high-confidence bounding box in the underwater image of the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of each target trajectory of the underwater image of the current timestamp: if the high-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated high-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory; if the high-confidence bounding box is not successfully associated with any predicted bounding box and its confidence is higher than the preset tracking threshold, it is initialized as a new tracking instance of the target trajectory and is assigned a new tracking ID;
[0026] ④ For all low-confidence bounding boxes in the underwater image at the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of the target trajectory that has not been successfully associated with the underwater image at the current timestamp: if the low-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated low-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory;
[0027] ⑤ Update the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the current timestamp;
[0028] ⑥ If the predicted bounding box of the target track is not associated with a high confidence bounding box or a low confidence bounding box in a preset number of consecutive frames, the target track is removed from the tracking list.
[0029] In step 4 above, the key point position stream is the coordinates (x, y) of each key point at time t t , the motion information flow is the motion change of the key point between consecutive frames, that is, the coordinates (x, y) of each key point at time t+1 t+1 Subtract the coordinates (x,y) of each key point at time t t .
[0030] In the above step 4, the behavior recognition model includes 2 batch normalization layers, 18 spatiotemporal graph convolution modules (ST-GCN), 2 global average pooling layers, 1 splicing layer, 1 fully connected layer, and 1 activation function layer. Among the 18 spatiotemporal graph convolution modules, 9 spatiotemporal graph convolution modules are connected in sequence to form the first spatiotemporal graph convolution group, and the other 9 spatiotemporal graph convolution modules are connected in sequence to form the second spatiotemporal graph convolution group; the input of the first batch normalization layer forms the key point position stream input of the behavior recognition model, and the input of the second batch normalization layer forms the motion information stream input of the behavior recognition model; the output of the first batch normalization layer is connected to the input of the first spatiotemporal graph convolution group, and the output of the second batch normalization layer is connected to the input of the second spatiotemporal graph convolution group; the output of the first spatiotemporal graph convolution group and the output of the second spatiotemporal graph convolution group are connected to the input of the splicing layer, the output of the splicing layer is connected to the input of the fully connected layer, the output of the fully connected layer is connected to the input of the activation function layer, and the output of the activation function layer forms the output of the behavior recognition model.
[0031] Compared with the prior art, the invention has the following characteristics:
[0032] 1. By collecting underwater human body images and using the SSM-YOLO11-pose model to detect targets and key points, it not only improves detection speed and efficiency, but also ensures high accuracy;
[0033] 2. Using the ByteTrack multi-target tracking algorithm, multiple targets are distinguished, and the key points of each swimmer are matched with the tracking ID to achieve character tracking, and the target ID and key point time sequence are obtained. Different from the traditional model based on single-frame images, the accuracy of drowning behavior recognition is improved by analyzing the time sequence state sequence;
[0034] 3. Using the TS-STG model, the drowning situation can be accurately judged based on the key points of the human skeleton and the characteristics of the motion vector. Without the need for complex key point judgment logic, combined with timing analysis, it can identify abnormal limb swings, stillness, twitching and other behaviors in emergency situations, showing better model generalization;
[0035] 4. It does not require the person being tested to wear additional hardware. It is easy to deploy and can run unattended for a long time. It is suitable for various swimming scenarios and can achieve drowning detection with high accuracy, strong real-time performance and high anti-interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A flowchart of a method for detecting drowning in a swimming pool based on key points of human skeleton;
[0037] Figure 2It is a structural diagram of the human body bounding box and skeleton key point detection model (SSM-YOLO11-pose);
[0038] Figure 3 It is a schematic diagram of the structure of the lightweight convolution module (ShuffleNetV2);
[0039] Figure 4 It is a schematic diagram of the structure of the spatial pyramid pooling module (SPPF);
[0040] Figure 5 It is a schematic diagram of the structure of the feature extraction module (C2PSA_MLCA) fused with hybrid local channel attention;
[0041] Figure 6 This is a schematic diagram of the structure of the pyramid slice attention module (PSABlock);
[0042] Figure 7 It is a schematic diagram of the structure of the hybrid local channel attention module (MCLA);
[0043] Figure 8 It is a schematic diagram of the structure of the grouped spatial convolution module (GSConv);
[0044] Fig. 9 It is a schematic diagram of the structure of the cross-stage partial network module (VoVGSCSP);
[0045] Fig.10 It is a schematic diagram of the structure of the group shuffle bottleneck module (GSBottleneck);
[0046] Fig.11 It is the principle diagram of the multi-target tracking (ByteTrack) algorithm;
[0047] Fig.12 It is a structural diagram of the behavior recognition model (TS-STG);
[0048] Fig.13 Schematic diagram of the structure of the spatiotemporal graph convolution module (ST-GCN). DETAILED DESCRIPTION
[0049] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.
[0050] A swimming pool drowning detection method based on human skeleton key points, such as Figure 1 As shown, the steps include:
[0051] (1) Data acquisition and preprocessing: Underwater video is acquired using the swimming pool underwater monitoring equipment. A predetermined number of frame images are extracted from the underwater video at fixed time intervals. These frame images are preprocessed to obtain underwater images. Each underwater image has a timestamp.
[0052] The specific process of preprocessing each frame image includes:
[0053] First, the frame image is resized and padded to ensure that the image is scaled to the size required by the model input (640×640), while maintaining the image ratio and filling the blank area with gray;
[0054] Next, the frame image is processed in color and channels, the image is converted from BGR format to RGB format, and the pixel value is normalized to the [0.0, 1.0] interval;
[0055] Finally, the dimension and type of the frame image are adjusted, the image shape is adjusted to [N, C, H, W], and converted to FP16 or FP32 format as needed to adapt to the model input requirements.
[0056] (2) Human body bounding box and skeleton key point detection: The human body bounding box and skeleton key point detection model outputs the detection results of each underwater image with a timestamp, where the detection results include the human body bounding box and confidence, skeleton key points and the confidence of the corresponding points.
[0057] See also Figure 2 ,The human bounding box and skeleton key point detection model (SSM-YOLO11-pose) consists of three parts: the backbone network, the neck network and the head network.
[0058] First, the underwater image with a timestamp is input into the backbone network to generate a primary feature map, and three feature maps of different scales are generated through multi-level feature extraction to represent the multi-level information of the underwater human body. Then, the feature maps of three different scales are input into the neck network optimized based on Slim-Neck for feature fusion, which further improves the computational efficiency of the model by reducing redundant features and maintaining the ability to fusion multi-scale information. Finally, the fused feature map is sent to the detection head for prediction, and each human body bounding box and skeleton key points, as well as the confidence of the human body bounding box and skeleton key points, are marked in the underwater image with a timestamp.
[0059] The above backbone network includes 1 convolution layer (Conv), 4 lightweight convolution modules (ShuffleNetV2), 1 spatial pyramid pooling module (SPPF), and 1 feature extraction module fused with hybrid local channel attention (C2PSA_MLCA). 1 convolution layer, 4 lightweight convolution modules, 1 spatial pyramid pooling module and 1 feature extraction module fused with hybrid local channel attention are connected in sequence. The input of the convolution layer forms the input of the backbone network, that is, the input of the human body bounding box and skeleton key point detection model. The output of the feature extraction module fused with hybrid local channel attention forms the first output of the backbone network, the output of the third lightweight convolution module forms the second output of the backbone network, and the output of the second lightweight convolution module forms the third output of the backbone network.
[0060] The backbone network gradually extracts features from shallow to deep layers through three core steps:
[0061] Initial feature extraction: ShuffleNetV2 (such as Figure 3 As shown in the figure, the input feature map is first processed into two paths: the first path is downsampled through 3×3 deep convolution, and the number of channels is adjusted by 1×1 convolution to retain the global information and structural continuity of the input features; the second path first linearly transforms the input channel through 1×1 convolution, then downsamples the feature map using 3×3 deep convolution to extract spatial features, and then maps the downsampled features to the target number of channels through 1×1 convolution again, thereby extracting more detailed features and capturing high-order information in the input features. Finally, the results of the two paths are fused through splicing and channel shuffling to efficiently complete the downsampling operation while minimizing the computational overhead and the number of parameters.
[0062] Multi-scale enhancement: SPPF (such as Figure 4 As shown in Figure 1, the spatial pyramid pooling technique is used to perform multi-scale pooling operations on the input feature map. The feature map is pooled using pooling kernels of different sizes, and then the obtained multi-scale features are concatenated or added, thereby enhancing the model's ability to recognize the target.
[0063] Attention Enhancement: C2PSA_MLCA (such as Figure 5 The CNN (as shown in Figure 1) is a module based on convolution and attention mechanism. The input feature map is first compressed by a 1×1 convolution layer and divided into two parts, a and b. Part a is directly reserved for the final feature fusion, while part b is sent to the stacked PSABlock to enhance its channel features and spatial features through the attention mechanism. The processed feature map b is concatenated with the directly retained feature a, restored to the dimension of the input feature through a 1×1 convolution, and the final enhanced feature is output.
[0064] See also Figure 5 The feature extraction module of fused hybrid local channel attention includes 2 convolutional layers (Conv), N pyramid slice attention modules (PSABlock), and 1 concatenation layer (Concat), where N is a set value. The input of the first convolutional layer forms the input of the feature extraction module of fused hybrid local channel attention, the N pyramid slice attention modules are connected in sequence, the input of the first pyramid slice attention module is connected to the output of the first convolutional layer, the output of the last pyramid slice attention module and the output of the first convolutional layer are simultaneously connected to the input of the concatenation layer, the output of the concatenation layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer forms the output of the feature extraction module of fused hybrid local channel attention.
[0065] See also Figure 6 Each pyramid slice attention module includes a hybrid local channel attention module (MCLA), two convolutional layers (Conv), and two addition layers; the input of the hybrid local channel attention module forms the input of the pyramid slice attention module, the input and output of the fused hybrid local channel attention are simultaneously connected to the input of the first addition layer, the output of the first addition layer is connected to the input of the first convolutional layer, the output of the first convolutional layer is connected to the input of the second convolutional layer, the output of the second convolutional layer and the input of the first convolutional layer are simultaneously connected to the input of the second addition layer, and the output of the second addition layer forms the output of the pyramid slice attention module. The hybrid local channel attention module is as follows: Figure 7 shown.
[0066] The neck network includes 4 concatenation layers (Concat), 5 grouped spatial convolution modules (GSConv), 4 cross-stage partial network modules (VoVGSCSP), and 2 upsampling layers (Upsample). The input of the first grouped spatial convolution module forms the first input of the neck network, which is connected to the first output of the backbone network; an input of the first concatenation layer forms the second input of the neck network, which is connected to the second output of the backbone network; an input of the second concatenation layer forms the third input of the neck network, which is connected to the third output of the backbone network. The output of the first grouped spatial convolution module is connected to another input of the first splicing layer via the first upsampling layer, the output of the first splicing layer is connected to the input of the second grouped spatial convolution module via the first cross-stage partial network module, the output of the second grouped spatial convolution module is connected to another input of the second splicing layer via the second upsampling layer, the output of the second splicing layer is connected to the input of the second cross-stage partial network module via the third grouped spatial convolution module, and the output of the second cross-stage partial network module is connected to the input of the fourth grouped spatial convolution module; the output of the fourth grouped spatial convolution module and the output of the second grouped spatial convolution module are simultaneously connected to the input of the third grouped spatial convolution module, the output of the third grouped spatial convolution module is connected to the input of the third cross-stage partial network module, and the output of the third cross-stage partial network module is connected to the input of the fifth grouped spatial convolution module; the output of the fifth grouped spatial convolution module and the output of the first grouped spatial convolution module are simultaneously connected to the input of the fourth splicing layer, and the output of the fourth splicing layer is connected to the input of the fourth cross-stage partial network module. The output of the fourth cross-stage partial network module forms the first output of the neck network, the output of the third cross-stage partial network module forms the second output of the neck network, and the output of the second cross-stage partial network module forms the third output of the neck network.
[0067] The neck network adopts the Slim-Neck structure and focuses on the fusion and processing of features. Feature maps of different scales are first processed by GSConv, and then combined with feature maps of other scales through upsampling and splicing operations. After this series of processing, the feature maps are refined again by GSConv, and finally VoVGSCSP is used to further extract and fuse features to prepare for target detection by the detection head.
[0068] GSConv (such as Figure 8The CNN (as shown in Figure 1) combines standard convolution with depthwise separable convolution to achieve efficient and lightweight feature extraction. The input channel is first divided into two parts: the first part is processed by standard convolution to extract basic features; the second part is subjected to deep convolution on the features extracted in the first part to enhance the local receptive field. Subsequently, the two parts of features are concatenated together and the channel order is rearranged through feature reorganization operations, thereby significantly enhancing the expressiveness of features while reducing the computational burden.
[0069] VoVGSCSP (such as Fig. 9 The CNN (as shown in Figure 1) is an efficient feature fusion architecture. The initial convolutional layer first divides the input features into two branches: the main branch extracts deep features through multiple grouping and shuffling bottleneck modules, and uses residual connections to adjust features to enhance feature representation; the other branch directly transmits low-level features to maintain the integrity of the original information. Finally, the two features are concatenated and fused in the final convolutional layer to generate comprehensive output features. Fig.10 Each group shuffle bottleneck module (GSBottleneck) also achieves feature fusion through a dual-path design: one path gradually strengthens the feature expression through two GSConv layers, and the other path directly transmits the input features; finally, the two paths of features are added to enhance the feature details and completeness.
[0070] The head network includes three detection heads, one bounding box selection module and one key point selection module. The first detection head forms the first input of the head network and is connected to the first output of the neck network; the second detection head forms the second input of the head network and is connected to the second output of the neck network; the third detection head forms the third input of the head network and is connected to the third output of the neck network; the bounding box outputs of the three detection heads are simultaneously connected to the input of the bounding box selection module, and the output of the bounding box selection module forms the bounding box output of the head network, i.e., the bounding box output of the human body bounding box and the skeleton key point detection model, and the key point outputs of the three detection heads are simultaneously connected to the input of the key point selection module, and the output of the key point selection module forms the key point output of the head network, i.e., the key point output of the human body bounding box and the skeleton key point detection model.
[0071] Each detection head consists of two parts: the first part is used to predict the bounding box of the underwater human body, which contains the bounding box coordinates and confidence of each detected human body; the second part is used to predict the key points of the underwater human body. Each detected target human body has a set of key points, and the output key point vector is in the form of [x i ,y i ,c i ], where x i ,y i is the coordinate of the i-th key point, c iis the corresponding confidence level. In this embodiment, there are 17 human key points, corresponding to 17 bone key points of the human body, namely the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle. The detection head uses a shared network structure to synchronously predict the human bounding box and bone key points. The bounding box is obtained by regressing the center point coordinates, width and height of the bounding box; the bone key points are obtained by predicting the normalized coordinates and confidence of each key point, and then denormalizing it in combination with the bounding box size, so as to accurately locate the key points on the original image. According to the preset bone connection key point index, the bounding box and bone structure are drawn and the results are output.
[0072] For the final result processing of the three detection heads: first calculate the intersection over union (IoU) of the bounding box output by each detection head, compare it with the boxes of other scales, retain the result of the box with the largest IoU, then check whether the key points of each box are completely located in the box, and select the detection result with more key points and more accurate key point position (higher confidence) as the final output.
[0073] (3) Human target tracking: Based on the human body bounding box in the detection results of each underwater image with a timestamp, a multi-target tracking algorithm is used to perform target allocation and trajectory tracking, and a tracking ID is assigned to each detected target; then, the skeletal key points in the detection results of each underwater image with a timestamp are associated with the tracking ID and organized into a target temporal key point sequence in the time order of the corresponding timestamps.
[0074] The multi-target tracking algorithm of the present invention is implemented by using the existing multi-target tracking algorithm ByteTrack. Fig.11 As shown, the specific process is as follows:
[0075] The detection results of underwater images of all timestamps are processed in order from the front to the back:
[0076] The following operations are performed on the detection results of the underwater image at the initial timestamp:
[0077] ① Extract bounding boxes and confidences from the detection results of the underwater image at the initial timestamp, and divide the bounding boxes into two categories by comparing them with the preset confidence threshold: i Greater than or equal to the confidence threshold T c (s i ≥T c ), it is defined as a high confidence bounding box; when the confidence of the bounding box s i Less than the confidence threshold T c (s i <T c ), it is defined as a low confidence bounding box;
[0078] ② Initialize each high-confidence bounding box in the underwater image at the initial timestamp as a tracking instance of each target trajectory, assign a unique tracking ID to each, and initialize the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the initial timestamp;
[0079] The following operations are performed on the detection results of the underwater images with the remaining time stamps after the initial time stamp:
[0080] ① The current Kalman filter predicts the position of each target track's tracking instance in the underwater image at the current timestamp based on the tracking instance state of each target track in the underwater image at the previous timestamp, and obtains the predicted bounding box of each target track;
[0081] ② Extract the bounding box and confidence from the detection result of the underwater image at the current timestamp, and divide the bounding box into two categories by comparing with the preset confidence threshold: i Greater than or equal to the confidence threshold T c (s i ≥T c ), it is defined as a high confidence bounding box; when the confidence of the bounding box s i Less than the confidence threshold T c (s i <T c ), it is defined as a low confidence bounding box;
[0082] ③ For each high-confidence bounding box in the underwater image of the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of each target trajectory of the underwater image of the current timestamp: if the high-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated high-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory; if the high-confidence bounding box is not successfully associated with any predicted bounding box and its confidence is higher than the preset tracking threshold, it is initialized as a new tracking instance of the target trajectory and is assigned a new tracking ID;
[0083] ④ For all low-confidence bounding boxes in the underwater image at the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of the target trajectory that has not been successfully associated with the underwater image at the current timestamp: if the low-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated low-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory;
[0084] ⑤ Update the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the current timestamp;
[0085] ⑥ If the predicted bounding box of the target track is not associated with a high confidence bounding box or a low confidence bounding box in a preset number of consecutive frames, the target track is removed from the tracking list.
[0086] The present invention not only performs the first association operation on the high-confidence bounding box, but also performs the second association operation on the low-confidence bounding box and the unsuccessfully matched trajectory bounding box, so as to enrich the information of the tracked trajectory and effectively reduce the missed detection rate. The unmatched low-confidence bounding box is discarded to maintain the accuracy and reliability of the trajectory data.
[0087] (4) Human behavior recognition: Each target time-series key point sequence is processed into a key point position stream and a motion information stream, and the key point position stream and the motion information stream are simultaneously fed into the behavior recognition model. The behavior recognition model outputs the classification result and obtains the final dangerous or normal behavior classification probability of the swimmer.
[0088] 1) Data stream division: Before the target time-series key point sequence is sent to TS-STG, it needs to be preprocessed and divided into two streams to make full use of the information: key point position stream (PTS) and motion information stream (MOT). The input of the PTS stream is the coordinates (x, y) of each key point at time t. t , and the input of the MOT stream is the motion change (x, y) of the key point between consecutive frames t+1 -(x,y) t , to capture the dynamic characteristics of the action.
[0089] 2) Behavior recognition: A behavior recognition model is used to determine the behavior type of the target time series key point sequence.
[0090] See also Fig.12The behavior recognition model (TS-STG) includes 2 batch normalization layers, 18 spatiotemporal graph convolution modules (ST-GCN), 2 global average pooling layers, 1 splicing layer, 1 fully connected layer, and 1 activation function layer. Among the 18 spatiotemporal graph convolution modules, 9 spatiotemporal graph convolution modules are connected in sequence to form the first spatiotemporal graph convolution group, and the other 9 spatiotemporal graph convolution modules are connected in sequence to form the second spatiotemporal graph convolution group; the input of the first batch normalization layer forms the key point position stream input of the behavior recognition model, and the input of the second batch normalization layer forms the motion information stream input of the behavior recognition model; the output of the first batch normalization layer is connected to the input of the first spatiotemporal graph convolution group, and the output of the second batch normalization layer is connected to the input of the second spatiotemporal graph convolution group; the output of the first spatiotemporal graph convolution group and the output of the second spatiotemporal graph convolution group are connected to the input of the splicing layer, the output of the splicing layer is connected to the input of the fully connected layer, the output of the fully connected layer is connected to the input of the activation function layer, and the output of the activation function layer forms the output of the behavior recognition model. Each spatiotemporal graph convolution module is Fig.13 shown.
[0091] The behavior recognition model uses a series of spatiotemporal graph convolution modules to extract spatial and temporal features, fuse the features and output the classification results to obtain the final dangerous / normal behavior classification probability. The specific process is as follows:
[0092] First, we use the spatiotemporal graph convolution module (such as Fig.12 The feature extraction is performed as shown in FIG.
[0093] In the spatial dimension, the graph convolution operation uses the adjacency matrix A k and the weight matrix W k Capturing the spatial association between key points, the formula is expressed as:
[0094]
[0095] Among them, X represents the input feature, X ′ Represents the output features.
[0096] In the time dimension, the one-dimensional convolution operation focuses on capturing the timing information in the sequence data while filtering out the noise to ensure that the continuity of the action is preserved.
[0097] The spatial features and temporal features are processed through nonlinear activation functions and deeply fused with the help of residual connections to update the feature map.
[0098] In addition, for each layer of the spatiotemporal graph convolution module, learnable parameters are introduced to optimize the adjacency matrix of the graph, which enhances the weights of key connections, allowing the model to focus more accurately on important structural information, thereby improving the recognition accuracy of action patterns.
[0099] The dual-stream design and feature extraction method for key point position stream and motion information stream provides rich spatiotemporal features for subsequent action recognition, greatly improving the recognition performance of the model.
[0100] Then, the output features of the key point position stream and the motion information stream are spliced through the splicing layer and integrated into the global action feature. This fusion feature generates the probability distribution of the action category by mapping the feature to the category space, and the activation function maps the output value to the range of [0,1]. The action classification result of the swimmer to be identified is output, and the action type is clearly divided into "dangerous" and "normal", providing a basis for real-time monitoring and early warning.
[0101] The present invention first uses an advanced human key point detection model to perform real-time analysis on the video stream and locate the key points of the human body. Subsequently, a multi-target tracking model is used to distinguish multiple swimmers and continuously track them, capturing the dynamic changes of each swimmer in the time series, thereby obtaining the motion trajectory of the human body. Finally, the extracted human key point features are deeply analyzed through the action recognition model to determine whether there are signs of drowning. The drowning detection scheme of the present invention can not only monitor and accurately identify drowning situations in real time, but also significantly improve the detection efficiency, accuracy and generalization performance of the model, providing more reliable technical support for swimming pool safety monitoring.
[0102] It should be noted that although the embodiments of the present invention described above are illustrative, they are not intended to limit the present invention, and therefore the present invention is not limited to the above specific embodiments. Without departing from the principles of the present invention, any other embodiments obtained by those skilled in the art under the guidance of the present invention are deemed to be within the protection of the present invention.
Claims
1. A method for detecting drowning in a swimming pool based on key points of human skeleton, characterized in that: The steps include: Step 1: Use the underwater monitoring equipment of the swimming pool to obtain underwater video, and extract a predetermined number of frame images from the underwater video at fixed time intervals, and pre-process these frame images to obtain underwater images, wherein each underwater image has a timestamp; Step 2: Each underwater image with a timestamp is sent to a human body bounding box and skeleton key point detection model, which outputs a detection result of each underwater image with a timestamp, wherein the detection result includes a human body bounding box and confidence, and a skeleton key point and the confidence of the corresponding point; Step 3: Based on the human body bounding box in the detection result of each underwater image with a timestamp, a multi-target tracking algorithm is used to perform target allocation and trajectory tracking, and a tracking ID is assigned to each detected target; Then, the skeleton key points in the detection results of each underwater image with a timestamp are associated with the tracking ID, and organized into a target time sequence key point sequence according to the time order of the corresponding timestamp; Step 4: Process each target time-series key point sequence into a key point position stream and a motion information stream, and simultaneously feed the key point position stream and the motion information stream into the behavior recognition model. The behavior recognition model outputs the classification result to obtain the swimmer's final dangerous or normal behavior classification probability.
2. A swimming pool drowning detection method based on human skeleton key points according to claim 1, characterized in that: The preprocessing of each image in step 1 includes resizing and filling processing, color and channel processing, and dimension and type adjustment processing.
3. A swimming pool drowning detection method based on human skeleton key points according to claim 1, characterized in that: In step 2, the human body bounding box and skeleton key point detection model consists of a backbone network, a neck network, and a head network; The backbone network includes 1 convolution layer, 4 lightweight convolution modules, 1 spatial pyramid pooling module, and 1 feature extraction module fused with mixed local channel attention; wherein 1 convolution layer, 4 lightweight convolution modules, 1 spatial pyramid pooling module and 1 feature extraction module fused with mixed local channel attention are connected in sequence; the input of the convolution layer forms the input of the backbone network, that is, the input of the human body bounding box and skeleton key point detection model; the output of the feature extraction module fused with mixed local channel attention forms the first output of the backbone network, the output of the third lightweight convolution module forms the second output of the backbone network, and the output of the second lightweight convolution module forms the third output of the backbone network; The above-mentioned neck network includes 4 splicing layers, 5 grouped spatial convolution modules, 4 cross-stage partial network modules, and 2 upsampling layers; wherein the input of the first grouped spatial convolution module forms the first input of the neck network, which is connected to the first output of the backbone network; an input of the first splicing layer forms the second input of the neck network, which is connected to the second output of the backbone network; an input of the second splicing layer forms the third input of the neck network, which is connected to the third output of the backbone network; the output of the first grouped spatial convolution module is connected to another input of the first splicing layer via the first upsampling layer, the output of the first splicing layer is connected to the input of the second grouped spatial convolution module via the first cross-stage partial network module, the output of the second grouped spatial convolution module is connected to another input of the second splicing layer via the second upsampling layer, and the output of the second splicing layer is connected to the second cross-stage via the third grouped spatial convolution module. The input of the partial network module, the output of the second cross-stage partial network module is connected to the input of the fourth grouped spatial convolution module; the output of the fourth grouped spatial convolution module and the output of the second grouped spatial convolution module are simultaneously connected to the input of the third grouped spatial convolution module, the output of the third grouped spatial convolution module is connected to the input of the third cross-stage partial network module, and the output of the third cross-stage partial network module is connected to the input of the fifth grouped spatial convolution module; the output of the fifth grouped spatial convolution module and the output of the first grouped spatial convolution module are simultaneously connected to the input of the fourth splicing layer, and the output of the fourth splicing layer is connected to the input of the fourth cross-stage partial network module; the output of the fourth cross-stage partial network module forms the first output of the neck network, the output of the third cross-stage partial network module forms the second output of the neck network, and the output of the second cross-stage partial network module forms the third output of the neck network; The head network includes three detection heads, one bounding box selection module and one key point selection module; the first detection head forms the first input end of the head network and is connected to the first output of the neck network; the second detection head forms the second input end of the head network and is connected to the second output of the neck network; the third detection head forms the third input end of the head network and is connected to the third output of the neck network; the bounding box outputs of the three detection heads are simultaneously connected to the input of the bounding box selection module, and the output of the bounding box selection module forms the bounding box output of the head network, that is, the bounding box output of the human body bounding box and the skeleton key point detection model, and the key point outputs of the three detection heads are simultaneously connected to the input of the key point selection module, and the output of the key point selection module forms the key point output of the head network, that is, the key point output of the human body bounding box and the skeleton key point detection model.
4. A swimming pool drowning detection method based on human skeleton key points according to claim 3, characterized in that: The feature extraction module integrating hybrid local channel attention includes 2 convolutional layers, N pyramid slice attention modules, and 1 concatenation layer, where N is a set value; Each pyramid slice attention module includes a hybrid local channel attention module, two convolutional layers, and two addition layers; the input of the hybrid local channel attention module forms the input of the pyramid slice attention module, the input and output of the fused hybrid local channel attention are simultaneously connected to the input of the first addition layer, the output of the first addition layer is connected to the input of the first convolutional layer, the output of the first convolutional layer is connected to the input of the second convolutional layer, the output of the second convolutional layer and the input of the first convolutional layer are simultaneously connected to the input of the second addition layer, and the output of the second addition layer forms the output of the pyramid slice attention module; The input of the first convolutional layer forms the input of the feature extraction module fused with hybrid local channel attention. The N pyramid slice attention modules are connected in sequence. The input of the first pyramid slice attention module is connected to the output of the first convolutional layer. The output of the last pyramid slice attention module and the output of the first convolutional layer are simultaneously connected to the input of the concatenation layer. The output of the concatenation layer is connected to the input of the second convolutional layer. The output of the second convolutional layer forms the output of the feature extraction module fused with hybrid local channel attention.
5. A swimming pool drowning detection method based on human skeleton key points according to claim 1, characterized in that: In step 2, there are 17 skeleton key points, namely nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle.
6. A swimming pool drowning detection method based on human skeleton key points according to claim 1, characterized in that: In step 3, the multi-target tracking algorithm is the ByteTrack multi-target tracking algorithm, namely: The detection results of underwater images of all timestamps are processed in order from the front to the back: The following operations are performed on the detection results of the underwater image at the initial timestamp: ① Extract bounding boxes and confidences from the detection results of the underwater image at the initial timestamp, and divide the bounding boxes into two categories by comparing them with the preset confidence threshold: when the confidence of the bounding box is greater than or equal to the confidence threshold, it is defined as a high-confidence bounding box; when the confidence of the bounding box is less than the confidence threshold, it is defined as a low-confidence bounding box; ② Initialize each high-confidence bounding box in the underwater image at the initial timestamp as a tracking instance of each target trajectory, assign a unique tracking ID to each, and initialize the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the initial timestamp; The following operations are performed on the detection results of the underwater images with the remaining time stamps after the initial time stamp: ① The current Kalman filter predicts the position of each target track's tracking instance in the underwater image at the current timestamp based on the tracking instance state of each target track in the underwater image at the previous timestamp, and obtains the predicted bounding box of each target track; ② Extract bounding boxes and confidences from the detection results of the underwater image at the current timestamp, and divide the bounding boxes into two categories by comparing them with the preset confidence threshold: when the confidence of the bounding box is greater than or equal to the confidence threshold, it is defined as a high-confidence bounding box; when the confidence of the bounding box is less than the confidence threshold, it is defined as a low-confidence bounding box; ③ For each high-confidence bounding box in the underwater image of the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of each target trajectory of the underwater image of the current timestamp: if the high-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated high-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory; if the high-confidence bounding box is not successfully associated with any predicted bounding box and its confidence is higher than the preset tracking threshold, it is initialized as a new tracking instance of the target trajectory and is assigned a new tracking ID; ④ For all low-confidence bounding boxes in the underwater image at the current timestamp, the Hungarian algorithm based on the intersection-over-union distance matrix is used to associate with the predicted bounding boxes of the target trajectory that has not been successfully associated with the underwater image at the current timestamp: if the low-confidence bounding box is successfully associated with a predicted bounding box, the successfully associated low-confidence bounding box is used as the tracking instance of the corresponding target trajectory and is assigned the tracking ID of the corresponding target trajectory; ⑤ Update the Kalman filter according to the position and size of the tracking instance of each target trajectory in the underwater image at the current timestamp; ⑥ If the predicted bounding box of the target track is not associated with a high confidence bounding box or a low confidence bounding box in a preset number of consecutive frames, the target track is removed from the tracking list.
7. A method for detecting drowning in a swimming pool based on key points of human skeleton according to claim 1, characterized in that: In step 4, the key point position stream is the coordinates (x, y) of each key point at time t t , the motion information flow is the motion change of the key point between consecutive frames, that is, the coordinates (x, y) of each key point at time t+1 t+1 Subtract the coordinates (x,y) of each key point at time t t .
8. The method for detecting drowning in a swimming pool based on key points of human skeleton according to claim 1, characterized in that: In step 4, the behavior recognition model includes 2 batch normalization layers, 18 spatiotemporal graph convolution modules, 2 global average pooling layers, 1 concatenation layer, 1 fully connected layer, and 1 activation function layer; Among the 18 spatiotemporal graph convolution modules, 9 spatiotemporal graph convolution modules are connected in sequence to form the first group of spatiotemporal graph convolution groups, and the other 9 spatiotemporal graph convolution modules are connected in sequence to form the second group of spatiotemporal graph convolution groups; the input of the first batch normalization layer forms the key point position stream input of the behavior recognition model, and the input of the second batch normalization layer forms the motion information stream input of the behavior recognition model; the output of the first batch normalization layer is connected to the input of the first group of spatiotemporal graph convolution groups, and the output of the second batch normalization layer is connected to the input of the second group of spatiotemporal graph convolution groups; the output of the first group of spatiotemporal graph convolution groups and the output of the second group of spatiotemporal graph convolution groups are connected to the input of the splicing layer, the output of the splicing layer is connected to the input of the fully connected layer, the output of the fully connected layer is connected to the input of the activation function layer, and the output of the activation function layer forms the output of the behavior recognition model.
Citation Information
Cited By
River drowning monitoring method based on lossless downsampling network and multi-modal semantic disambiguation
CN121725430A
A River Drowning Monitoring Method Based on Lossless Downsampling Network and Multimodal Semantic Disambiguation
CN121725430B