Swimming pool drowning alert method and system

CN122821598APending Publication Date: 2026-09-25ZHEJIANG HUANGLONG HULA NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611330449.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]本发明的目的在于克服现有技术中存在的上述缺陷,提供泳池溺水预警方法及系统,解决现有水下溺水识别方案存在的畸变矫正不足、识别精度低、误报漏报率高、实时性与隐私合规性无法兼顾的技术问题

Benefits of technology

[0035]本发明通过基于生成对抗网络(GAN)的水下光路折射逆映射矫正处理,从根源解决了水下图像非线性形变导致的骨骼关键点识别失效问题。同时采用自上而下(Top-Down)的两阶段卷积神经网络(CNN)架构,通过单阶段实时目标检测卷积神经网络(YOLOv8)精准定位人体目标,结合高分辨率保持型卷积神经网络(HRNet)提取标准化骨骼关键点,将水下场景的人体姿态识别准确率提升至98%以上,为后续分类判决提供了精准的基础数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821598A_ABST
    Figure CN122821598A_ABST
Patent Text Reader

Abstract

The application discloses a swimming pool drowning early warning method and system, and relates to the technical field of artificial intelligence and computer vision. The method first pre-processes the image collected by the underwater camera, then locates the human body through a single-stage real-time detection network, extracts seventeen skeletal key points through a high-resolution maintaining network, corrects the inverse mapping of underwater light refraction through a generative adversarial network to restore the real posture; meanwhile, the single-frame spatial features and continuous frame timing features are extracted, the classification and judgment are made according to the international lifesaving federation drowning posture rules, and the early warning is triggered when the drowning state lasts for more than a preset threshold. The system comprises image pre-processing, human body detection, key point extraction, distortion correction, feature extraction, classification judgment and early warning triggering modules, can be locally deployed, has high recognition accuracy, low false and missed report rate, short time delay and privacy protection, and is suitable for swimming pool safety monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and computer vision technology, specifically to a method and system for early warning of drowning in swimming pools. Background Technology

[0002] With the development of smart sports venues, swimming pool safety monitoring has become a core and essential aspect of venue operation. Currently, swimming pool drowning prevention and safety mainly rely on on-site lifeguards. However, due to limitations in human field of vision, fatigue and inattention during long hours of duty, and the distraction caused by monitoring multiple lanes simultaneously, blind spots and untimely responses are prone to occur, posing serious safety hazards.

[0003] Existing electronic drowning prevention systems are mainly divided into two categories: contact and non-contact. Contact systems require swimmers to wear waterproof wristbands and other equipment, which suffer from problems such as equipment malfunction, inconvenience in wearing, and failure due to missing wristbands. Non-contact systems are mostly based on the analysis of video streams captured by underwater cameras, which is the current mainstream research and development direction in the industry, but still have the following core shortcomings:

[0004] First, underwater imaging distortion leads to insufficient recognition accuracy. Due to the refraction of light at the interfaces between water and air, and between water and the camera glass, the human images captured by underwater cameras undergo nonlinear deformation. This causes standard human skeleton recognition algorithms to experience key point drift, inaccurate positioning, or even recognition failure in underwater scenes, directly resulting in a significant decrease in the accuracy of subsequent posture classification. Existing technologies, such as the invention patent with publication number CN116152928A, completely disregard the impact of underwater refraction distortion and directly perform human posture estimation on the original distorted image. Its key point positioning error can reach 15-20 pixels in underwater scenes, seriously affecting the accuracy of subsequent drowning determination.

[0005] Second, the high false positive and false negative rates fail to meet practical requirements. Existing solutions mostly rely on speckle detection and simple motion amplitude detection in single-frame images for judgment, without employing a standardized human skeleton keypoint detection system, thus failing to accurately capture the subtle joint features of drowning postures. The aforementioned comparison file uses a single-stage lightweight posture estimation model, extracting only the keypoint distance, angle, and velocity vectors from a single-frame image, and determining drowning by comparing cosine similarity with a custom set of 13 drowning states. This has significant drawbacks: first, it cannot distinguish between normal active diving breath-holding training, pool play, and real drowning struggle states, with a false positive rate exceeding 30%; second, the recognition rate for silent drowning postures conforming to the International Lifesaving Federation's definition is extremely low, posing a serious risk of false negatives; and third, the custom set of drowning states lacks scientific basis and cannot cover all real drowning scenarios.

[0006] Third, real-time performance and privacy compliance cannot be simultaneously achieved. Existing solutions mostly employ a model of uploading raw video streams to the cloud for processing. Video stream transmission introduces high latency, failing to meet the core requirement of millisecond-level response for drowning warnings. Furthermore, the raw video stream contains sensitive privacy information such as the swimmer's face and body, posing a serious risk of privacy breaches when stored and processed in the cloud, thus failing to comply with relevant personal information protection regulations. While the aforementioned comparative documents employ a cloud-edge collaborative architecture, they still require uploading some data to the cloud for model updates and do not clearly specify the processing method for the raw video, posing a potential privacy breach risk.

[0007] Therefore, there is an urgent need for a swimming pool drowning early warning technology that can solve the problem of underwater imaging distortion, significantly reduce the false alarm and false negative rates, and at the same time take into account real-time performance and privacy compliance. Summary of the Invention

[0008] The purpose of this invention is to overcome the above-mentioned defects in the prior art and provide a swimming pool drowning early warning method and system, which solves the technical problems of insufficient distortion correction, low recognition accuracy, high false alarm and false alarm rates, and inability to balance real-time performance and privacy compliance in existing underwater drowning identification schemes.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A method for early warning of drowning in swimming pools includes the following steps:

[0011] S1. Acquire real-time video stream images of the swimming pool captured by the underwater camera and complete basic preprocessing;

[0012] S2. Based on a single-stage real-time target detection convolutional neural network, locate human targets in the preprocessed image and output the corresponding human bounding boxes;

[0013] S3. Input the image region within the bounding box into a high-resolution preservative convolutional neural network to extract key point data of the human skeleton;

[0014] S4. Perform underwater optical path refraction inverse mapping correction processing on the preprocessed real-time video stream image based on generative adversarial network to restore the true proportions and posture of the human body.

[0015] S5. Based on the skeletal key point data of continuous multi-frame images, combined with the corrected real human posture, the spatial posture features of a single frame and the temporal motion features of continuous frames are extracted simultaneously.

[0016] S6. Based on the single-frame spatial posture features and continuous-frame temporal motion features, and combined with the preset posture classification rules that conform to the International Lifesaving League's definition of instinctive drowning postures, the human posture is classified and judged to distinguish between normal swimming state, active diving state and drowning struggle state.

[0017] S7. When it is determined that the person is in a state of drowning and struggling and the duration of this state exceeds a preset threshold, a corresponding warning instruction is triggered.

[0018] Preferably, the basic preprocessing in step S1 includes deblurring the real-time video stream image, underwater illumination correction, and foreground segmentation.

[0019] Preferably, in step S3, the human skeleton key point data is extracted, specifically seventeen key points of the standard human skeleton are extracted, including the key points of the nose, eyes and ears of the head, the key points of the shoulder, elbow and wrist of the upper limb, and the key points of the hip, knee and ankle of the lower limb, and the coordinates and confidence of each key point are output simultaneously.

[0020] Preferably, the underwater optical path refraction inverse mapping correction process in step S4 specifically adopts a generative adversarial network model that includes a generator and a discriminator. The generator outputs the corrected image, and the discriminator completes adversarial training optimization to achieve nonlinear distortion elimination and visual interference removal of the underwater image.

[0021] Preferably, the single-frame spatial pose features in step S5 include the angle between the line connecting the torso, shoulders, and hips and the water surface, the angle between the head and the torso, the elbow angle of the upper limbs, and the vertical height difference between key points of the shoulders and hips.

[0022] Preferably, the continuous frame temporal motion features in step S5 include the periodicity of key point displacement, the range of joint angle changes, and the displacement pattern of the human body's center of mass; the preset threshold in step S7 is set to three to eight seconds.

[0023] A swimming pool drowning early warning system includes the following modules connected in sequence for communication, used to implement the swimming pool drowning early warning method:

[0024] The image acquisition and preprocessing module is used to perform step S1, acquire the real-time video stream image of the swimming pool captured by the underwater camera, and complete basic preprocessing.

[0025] The human body detection module is used to perform step S2, which locates human targets in the preprocessed image based on a single-stage real-time target detection convolutional neural network and outputs the corresponding human body bounding box.

[0026] The key point extraction module is used to perform step S3, inputting the image region within the bounding box into a high-resolution preserving convolutional neural network to extract key point data of the human skeleton.

[0027] The distortion correction module is used to perform step S4, which performs underwater optical path refraction inverse mapping correction processing based on generative adversarial network on the preprocessed real-time video stream image to restore the true proportions and posture of the human body.

[0028] The feature extraction module is used to perform step S5, which extracts single-frame spatial pose features and continuous-frame temporal motion features based on the skeletal key point data of multiple consecutive frames of images and combined with the corrected real human posture.

[0029] The classification and decision module is used to execute step S6, and classify and decide the human posture based on the single-frame spatial posture features and the continuous frame temporal motion features, combined with the preset posture classification rules that conform to the International Lifesaving League's definition of instinctive drowning posture, and distinguish between normal swimming state, active diving state and drowning struggle state.

[0030] The warning triggering module is used to execute step S7. When it is determined that the person is in a state of drowning and struggling and the duration of the state exceeds a preset threshold, the corresponding warning instruction is triggered.

[0031] Preferably, the basic preprocessing performed by the image acquisition and preprocessing module includes deblurring of the real-time video stream image, underwater illumination correction, and foreground segmentation.

[0032] Preferably, the corresponding early warning instruction triggered by the early warning triggering module includes at least one of the following: sending a vibration reminder and alarm location coordinates to the smart terminal worn by the lifeguard, displaying alarm area information on the venue monitoring screen, and triggering an on-site sound and light alarm device.

[0033] An electronic device includes a processor, a memory, a communication interface, and a bus, wherein the processor, memory, and communication interface communicate with each other via the bus; the memory stores at least one executable instruction, which causes the processor to perform an operation corresponding to the swimming pool drowning warning method.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] This invention addresses the root cause of skeletal keypoint recognition failure due to nonlinear deformation of underwater images by employing underwater optical path refraction inverse mapping correction processing based on Generative Adversarial Networks (GANs). Simultaneously, it utilizes a top-down two-stage convolutional neural network (CNN) architecture, employing a single-stage real-time target detection convolutional neural network (YOLOv8) to accurately locate human targets, combined with a high-resolution preserving convolutional neural network (HRNet) to extract standardized skeletal keypoints. This improves the accuracy of human pose recognition in underwater scenes to over 98%, providing accurate foundational data for subsequent classification decisions.

[0036] This invention simultaneously extracts single-frame spatial posture features and continuous-frame temporal motion features, and constructs classification rules based on the International Lifesaving Federation's definition of instinctive drowning postures. It can accurately distinguish between three states: normal swimming, active diving, and drowning struggle, with a particularly high recognition rate for silent drowning postures. Testing shows that the false alarm rate of this invention is less than 2%, and the false alarm rate is less than 0.5%, fundamentally solving the problem of high false alarm and false alarm rates in existing solutions and significantly reducing unnecessary interference to lifeguards.

[0037] This invention can be deployed on local edge computing nodes within swimming pool venues. Video streams do not need to be uploaded to the cloud, and processing latency can be controlled within 50 milliseconds, meeting the real-time requirements of drowning warnings. Simultaneously, the system only extracts key skeletal data for analysis, without storing or transmitting raw facial and body images, thus mitigating the risk of swimmer privacy leaks at the source and fully complying with relevant personal information protection regulations. Attached Figure Description

[0038] Figure 1 This is a flowchart of the swimming pool drowning early warning method of the present invention.

[0039] Figure 2 This is a diagram of the swimming pool drowning early warning system architecture of the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0041] Example 1

[0042] This embodiment provides a swimming pool drowning early warning method, the complete implementation process of which is as follows: Figure 1 As shown, the specific steps are as follows:

[0043] Step S1: Acquire real-time video stream images of the swimming pool captured by the underwater camera and complete basic preprocessing.

[0044] First, real-time video stream images of the pool area are captured by wide-angle underwater cameras positioned on the sidewalls or bottom of the pool. In this embodiment, the underwater camera is an industrial-grade waterproof camera with 1080P resolution and a frame rate of 30 frames per second (fps), installed at a height of 1.5 meters below the water surface, with a horizontal viewing angle of 120 degrees, covering a 25-meter lane of a standard swimming pool. The camera is connected to a local edge computing server via Ethernet and uses the Real-time Streaming Protocol (RTSP) to transmit video stream data, with a transmission latency of less than 10 milliseconds.

[0045] The acquired real-time video stream images are subjected to basic preprocessing frame by frame, including deblurring, underwater illumination correction, and foreground segmentation.

[0046] The deblurring process employs an adaptive deblurring algorithm to eliminate image blurring caused by water flow and light scattering during underwater camera capture. The specific calculation formula is as follows:

[0047]

[0048] In the formula:

[0049] The output sharp image in pixel coordinates The grayscale value at that location; The input blurred image in pixel coordinates The grayscale value at that location; Blur kernel: Characterizes the degree and type of image blur; : Two-dimensional convolution operator; The square of the L2 norm is used to measure the pixel-level difference between the reconstructed and blurred images. Regularization parameter, used to balance the weights of data fidelity terms and regularization terms. In this embodiment... ; :image gradient, ; : The L1 norm of the image gradient, also known as the total variation regularization term, is used to preserve image edge information and suppress noise.

[0050] Underwater illumination correction processing balances image brightness in low-light and backlit underwater scenes by limiting contrast adaptive histogram equalization (CLAHE). The specific steps are as follows: the image is divided into 8×8 sub-blocks, histogram equalization is performed on each sub-block, then bilinear interpolation is used to eliminate boundary effects between sub-blocks, and finally, the maximum contrast ratio is limited to 4.0 to avoid excessive noise enhancement.

[0051] Foreground segmentation is performed using Gaussian Mixture Background Modeling (GMM) to separate the pool water background from the human foreground target. Specifically, 3-5 Gaussian distribution models are built for each pixel, and the model parameters are updated online to adapt to slow changes in the background (such as water surface ripples and lighting changes). When the grayscale value of a pixel matches all background models below a threshold, it is determined to be a foreground pixel. In this embodiment, the matching threshold is set to 2.5 times the standard deviation, and the background model update rate is set to 0.001.

[0052] Step S2: Locate human targets in the preprocessed image using a single-stage real-time object detection convolutional neural network (YOLOv8) and output the corresponding human bounding boxes.

[0053] After preprocessing, a single frame image is input into a pre-trained single-stage real-time object detection convolutional neural network (YOLOv8). The model outputs the bounding box coordinates and class confidence scores of all human targets in the image. Overlapping redundant bounding boxes are filtered using non-maximum suppression (NMS), and finally, the rectangular bounding box corresponding to each valid human target is output.

[0054] For underwater swimming pool scenarios, this embodiment fine-tunes and optimizes a single-stage real-time object detection convolutional neural network (YOLOv8) model. The fine-tuning dataset contains 5000 underwater human images, covering swimmer images at different water depths, under different lighting conditions, and with different postures, divided into training set:validation set:test set ratios of 8:1:1. During training, the initial learning rate is set to 0.001, a cosine annealing learning rate scheduling strategy is adopted, the batch size is set to 16, and the training epochs are set to 100. Testing shows that the fine-tuned model achieves a human target detection rate of 99.2% in underwater scenarios, with a mean average accuracy (mAP@0.5) of 98.7%.

[0055] The output format of the bounding box is

[0056] In the formula: : The pixel coordinates of the top-left corner of the bounding box; : The pixel coordinates of the bottom right corner of the bounding box.

[0057] Bounding boxes with a confidence level below 0.5 are filtered out to avoid false detections.

[0058] Step S3: Input the image region within the bounding box into a high-resolution preserving convolutional neural network (HRNet) to extract the human skeleton keypoint data.

[0059] After obtaining the bounding box corresponding to the human target, the image region within the bounding box is cropped, resized to 256×192 pixels, and input into a high-resolution preserving convolutional neural network (HRNet) to extract the human skeleton key point data.

[0060] This embodiment employs a high-resolution preserving convolutional neural network (HRNet-W32) as its backbone. This network maintains high-resolution feature map output throughout the entire processing, significantly improving the localization accuracy of key points on the human skeleton in underwater scenes through multi-scale feature fusion. The network output consists of heatmaps of 17 standard human skeleton key points, each measuring 64×48 pixels. The pixel coordinates of each key point are obtained by locating the maximum value in the heatmap, and the corresponding confidence score is also output.

[0061] The 17 key points of the standard human skeleton extracted include: No. 0 (nose), No. 1 (left eye), No. 2 (right eye), No. 3 (left ear), No. 4 (right ear), No. 5 (left shoulder), No. 6 (right shoulder), No. 7 (left elbow), No. 8 (right elbow), No. 9 (left wrist), No. 10 (right wrist), No. 11 (left hip), No. 12 (right hip), No. 13 (left knee), No. 14 (right knee), No. 15 (left ankle), and No. 16 (right ankle).

[0062] When filtering keypoint confidence, priority is given to ensuring the effective detection of upper limb keypoints. Specifically, for upper limb keypoints (numbers 5-10), the confidence threshold is set to 0.3; for head keypoints (numbers 0-4) and lower limb keypoints (numbers 11-16), the confidence threshold is set to 0.5. When the confidence of a keypoint is lower than the threshold, linear interpolation of keypoints from the preceding and following frames is used to complete the detection.

[0063] Step S4: Perform underwater optical path refraction inverse mapping correction processing on the preprocessed real-time video stream image based on generative adversarial network (GAN) to restore the true proportions and posture of the human body.

[0064] Simultaneously, underwater refraction distortion correction processing is performed on the preprocessed real-time video stream image to restore the true proportions and posture of the human body. This embodiment adopts an underwater optical path refraction inverse mapping model based on generative adversarial networks (GANs), which includes two core modules: a generator and a discriminator.

[0065] The generator employs an encoder-decoder architecture. The encoder consists of five convolutional layers to extract features from the underwater distorted image, while the decoder consists of five deconvolutional layers to map the features back to a clear, distortion-free image. The generator takes an underwater distorted image affected by refraction as input and outputs a corrected, distortion-free image.

[0066] The discriminator uses a PatchGAN architecture, consisting of six convolutional layers, to determine whether the input image is a real image of a human in mid-air or an image corrected by the generator. The discriminator outputs a 30×30 probability map, where each pixel represents the probability that the corresponding image patch is a real image.

[0067] The model's loss function consists of three parts: adversarial loss, content loss, and perceptual loss. The specific calculation formula is as follows:

[0068]

[0069] In the formula: The total loss value of the model; Adversarial loss measures the ability of a generator to deceive a discriminator; the formula is as follows: ,in For the discriminator to distinguish real images The output, For the generator to handle input noise The output; Content loss is calculated using mean squared error (MSE) to determine the pixel-level difference between the generated and real images. The calculation formula is as follows: ; Perceptual loss is calculated using the feature map differences extracted by a pre-trained Visual Geometric Group Network (VGG19). The calculation formula is as follows: ,in This is a feature map output by a certain layer of the Visual Geometry Group Network (VGG19). , , The weighting coefficients for each loss term are set to 0.001, 1.0, and 0.006 respectively in this embodiment.

[0070] This embodiment uses a dataset of 10,000 pairs of "underwater-air" paired images for adversarial training. The dataset includes underwater human images at different water depths (0.5 meters to 3 meters), under different lighting conditions (natural light and artificial light), and with different degrees of distortion, as well as standard human images in the air in the corresponding scenes. During training, the generator and discriminator are trained alternately, with an initial learning rate set to 0.0002, using the Adam optimizer, a batch size of 8, and 200 training epochs.

[0071] Tests have shown that the underwater refraction distortion correction model of this invention can reduce the key point positioning error of underwater images from 15-20 pixels to 2-3 pixels, and the human body proportion restoration error is less than 5%, effectively solving the problem of decreased recognition accuracy caused by underwater imaging distortion.

[0072] Step S5: Based on the skeletal keypoint data of multiple consecutive frames, combined with the corrected real human posture, simultaneously extract the spatial posture features of a single frame and the temporal motion features of consecutive frames.

[0073] After extracting skeletal key points and correcting image distortion, the spatial pose features of a single frame and the temporal motion features of consecutive frames are extracted based on the skeletal key point data of multiple consecutive frames. In this embodiment, 60 consecutive video images within 2-3 seconds are selected from the multiple consecutive frames, corresponding to the standard frame rate setting of 30 frames per second (fps) for a swimming pool monitoring camera.

[0074] Single-frame spatial pose features include the angle between the line connecting the torso, shoulders, and hips and the water surface; the angle between the head and torso; the elbow angle of the upper limbs; and the vertical height difference between key points on the shoulders and hips. The specific calculation formulas are as follows:

[0075] Angle between the line connecting the torso, shoulders, and hips and the water surface:

[0076]

[0077] In the formula:

[0078] : The angle between the line connecting the torso, shoulders, and hips and the horizontal water surface, in degrees;

[0079] Average coordinates of key points on the left and right shoulders. , ,in The coordinates of the key points on the left shoulder. Coordinates of the key points on the right shoulder;

[0080] : Average coordinates of key points on the left and right hips , ,in The coordinates of the key points on the left hip are: Coordinates of key points on the right hip;

[0081] Pi (π) has a value of approximately 3.1415926.

[0082] Angle between head and torso:

[0083]

[0084] In the formula:

[0085] The angle between the head and torso, measured in degrees; The vector from the head keypoint (nose, 0) to the neck keypoint (midpoint of left and right shoulders). ,in The coordinates of the key points of the nose; The vector from the key points of the neck to the midpoint of the torso (midpoint of the left and right hips). ; :vector and The dot product; , : are vectors and The length of the module.

[0086] Elbow angle of upper limb:

[0087]

[0088]

[0089] In the formula:

[0090] , : These are the left elbow angle and the right elbow angle, respectively, in degrees;

[0091] Regarding the left elbow angle, The vector from the left shoulder to the left elbow. Let be the vector from the left elbow to the left wrist, where Here are the coordinates of the key point on the left elbow. Coordinates of key points on the left wrist;

[0092] For the right elbow angle, Let be the vector from the right shoulder to the right elbow. Let be the vector from the right elbow to the right wrist, where Here are the coordinates of the key point on the right elbow. The coordinates are the key points on the right wrist.

[0093] Vertical height difference between key shoulder and hip points:

[0094]

[0095] In the formula:

[0096] The actual vertical height difference between the key points of the shoulder and hip, in meters;

[0097] : The conversion coefficient from pixel to actual distance, which in this embodiment is based on the camera calibration results. Meters per pixel.

[0098] The temporal motion characteristics of consecutive frames include the periodicity of keypoint displacements, the range of joint angle changes, and the displacement pattern of the human body's center of mass. The specific calculation methods are as follows:

[0099] Periodicity of keypoint displacement: Fourier transforms are performed on the x and y coordinates of each keypoint in 60 consecutive frames of images to obtain the displacement spectrum. If there is a significant peak in the spectrum, and the frequency corresponding to the peak is between 0.5 Hz and 2 Hz (corresponding to the stroke / kick frequency of normal swimming), it is determined to be periodic; otherwise, it is determined to be non-periodic.

[0100] Range of joint angle variation: The maximum, minimum, and standard deviation of each joint angle were statistically analyzed in 60 consecutive frames of images. If the standard deviation of the joint angle is less than 5 degrees, it is determined to be joint stiffness; if the standard deviation is greater than 30 degrees, it is determined to be irregular joint movement.

[0101] The displacement law of the human body's center of mass: The formula for calculating the coordinates of the human body's center of mass is:

[0102]

[0103] In the formula:

[0104] , : These are the x and y coordinates of the human body's center of mass, respectively.

[0105] Calculate the displacement velocity and direction of the centroid in 60 consecutive frames. If the average forward velocity of the centroid is greater than 0.1 m / s, it is considered normal forward movement; if the average forward velocity of the centroid is less than 0.05 m / s and the vertical displacement is greater than 0.2 m, it is considered drowning struggle.

[0106] Step S6: Based on the single-frame spatial posture features and continuous frame temporal motion features, and combined with the preset posture classification rules that conform to the International Lifesaving League's definition of instinctive drowning postures, classify and determine the human posture, distinguishing between normal swimming state, active diving state, and drowning struggle state.

[0107] After feature extraction, the current human posture is classified based on the extracted single-frame spatial pose features and continuous-frame temporal motion features, combined with preset pose classification rules. The preset pose classification rules strictly conform to the "Guidelines for Recognizing Instinctive Drowning Reactions" published by the International Lifesaving League (ILS), as follows:

[0108] Judgment rules for normal swimming condition (meeting all conditions):

[0109] 1. Angle between the line connecting the torso, shoulders, and hips and the water surface. When the temperature is below 30 degrees Celsius, the body should be in a basically horizontal position.

[0110] 2. Vertical height difference between key shoulder and hip points Less than 0.2 meters;

[0111] 3. Angle between head and torso Keep your head in a natural position if the angle is less than 15 degrees.

[0112] 4. Elbow angle of upper limb and The movements vary periodically between 90 and 180 degrees, with alternating movements of the left and right upper limbs.

[0113] 5. The knee angle of the lower limbs changes periodically between 120 and 180 degrees, with symmetrical movement of the left and right lower limbs;

[0114] 6. The displacement of all key skeletal points exhibits a clear periodicity, with the stroke / leg kick cycle ranging from 0.5 seconds to 2 seconds;

[0115] 7. The center of mass of the human body tends to move forward at a constant speed, with an average forward speed greater than 0.1 m / s.

[0116] Decision rules for active diving (meeting all conditions):

[0117] 1. Angle between the line connecting the torso, shoulders, and hips and the water surface. Less than 30 degrees, the body is horizontal or tilted;

[0118] 2. The limbs move with small and gentle amplitude, and the range of joint angle changes is less than 20 degrees;

[0119] 3. Key point displacements are smooth, with no sudden or irregular movements;

[0120] 4. The head remains submerged for an extended period, with irregular surfacing movements;

[0121] 5. The center of gravity of the human body shows a slow downward or horizontal trend, without significant up-and-down fluctuations.

[0122] Judgment rules for drowning and struggling states (all conditions must be met):

[0123] 1. Angle between the line connecting the torso, shoulders, and hips and the water surface. Greater than 60 degrees, with the torso nearly perpendicular to the water surface;

[0124] 2. Vertical height difference between key shoulder and hip points Greater than 0.5 meters;

[0125] 3. The head is tilted back excessively, and the angle between the head and the torso is too large. At a temperature greater than 30 degrees, the key points of the nose / mouth barely touch the water surface, and there are no regular breathing movements;

[0126] 4. The upper limb is rigidly extended laterally in a horizontal position, with the elbow angle... and The wrist remains in a rigid state of 160-180 degrees for a long time, without periodic flexion and extension, the wrist key point is higher than the elbow and shoulder key points, and the left and right upper limbs are symmetrical without alternating movements.

[0127] 5. Irregular leg kicking movements of the lower limbs, with the knee angle maintained at an excessive flexion state below 90 degrees or a completely rigid state above 170 degrees for a long period of time, and asymmetry between the left and right lower limbs;

[0128] 6. The displacement of key points within consecutive frames is not periodic or regular, and there is no effective stroke / leg kick cycle within 10 seconds; 7. The center of gravity of the human body does not move forward, but only floats slightly in the vertical direction or even sinks continuously, with an average forward speed of less than 0.05 m / s.

[0129] This embodiment uses a weighted voting method for the final judgment, with the weight of each condition set according to its importance in determining drowning. For example, the weight of upper limb stiffness and lack of periodic movement is set to 2.0, while the weight of other conditions is set to 1.0. When the total score is greater than or equal to 8 points, the corresponding state is determined.

[0130] Step S7: When a drowning struggle is detected and the duration of this state exceeds a preset threshold, a corresponding warning instruction is triggered.

[0131] After completing the posture classification judgment, the drowning struggle state is continuously tracked. When the duration of continuous drowning struggle exceeds a preset threshold, a corresponding warning command is triggered. The preset threshold is set to 3 to 8 seconds, and in this embodiment, it is preferably set to 5 seconds.

[0132] The warning instructions include:

[0133] 1. Send vibration alerts and alarm location coordinates to the smart bracelets worn by lifeguards, with the location coordinates accurate to the specific lane and area;

[0134] 2. Display the real-time image of the alarm area and the key points of the skeleton on the venue's monitoring screen, and simultaneously pop up an alarm prompt box;

[0135] 3. Trigger the venue's audio-visual alarm device, emitting a high-decibel alarm sound and flashing red lights;

[0136] 4. Record the alarm time, alarm location, and corresponding video clips, and store them on the local server for easy tracing later.

[0137] If the human body returns to a normal swimming or active diving state before the warning is triggered, the warning determination will be cancelled and real-time monitoring will continue.

[0138] Example 2

[0139] This embodiment provides a swimming pool drowning early warning system to implement the swimming pool drowning early warning method described in Embodiment 1. The system architecture is as follows: Figure 2 As shown, it specifically includes:

[0140] The image acquisition and preprocessing module is used to perform step S1, acquiring real-time video stream images of the swimming pool captured by the underwater camera and completing basic preprocessing. This module interfaces with multiple underwater cameras in the swimming pool via Real-time Streaming Protocol (RTSP) or Open Network Video Interface Protocol (ONVIF), supporting simultaneous processing of 16 channels of 1080P 30 frames per second (fps) video streams.

[0141] The human detection module, communicating with the image acquisition and preprocessing module, is used to execute step S2. It locates human targets in the preprocessed image based on a single-stage real-time object detection convolutional neural network (YOLOv8) and outputs the bounding box corresponding to each valid human target. This module uses a hardware acceleration unit for inference, achieving a processing latency of less than 10 milliseconds per frame.

[0142] A key point extraction module, in communication connection with the human body detection module, is configured to perform step S3: input the image area within the bounding box into a high-resolution preservation convolutional neural network (HRNet), extract and obtain human skeleton key point data, and synchronously output the coordinates and confidence of each key point. This module also uses a hardware acceleration unit for inference, and the key point extraction delay of a single human target is less than 5 milliseconds.

[0143] A distortion correction module, in communication connection with the image acquisition and preprocessing module, is configured to perform step S4: perform inverse mapping correction processing based on a generative adversarial network (GAN) for underwater optical path refraction on the preprocessed real-time video stream image, and restore the real proportion and posture of the human body. This module can automatically adjust correction parameters according to the installation position and angle of different cameras, and adapt to different underwater scenarios.

[0144] A feature extraction module, in communication connection with the key point extraction module and the distortion correction module respectively, is configured to perform step S5: based on the skeleton key point data of consecutive multi-frame images, combine the corrected real human posture, and extract single-frame spatial posture features and continuous-frame temporal motion features at the same time. This module adopts a parallel computing architecture and can process feature extraction of multiple human targets simultaneously.

[0145] A classification and judgment module, in communication connection with the feature extraction module, is configured to perform step S6: according to the single-frame spatial posture features and continuous-frame temporal motion features, combined with preset posture classification rules conforming to the definition of instinct drowning posture by the International Life Saving Federation, perform classification and judgment on human postures, and distinguish between normal swimming state, active diving state and drowning struggle state. This module has a built-in classification rule library that meets the standards of the International Life Saving Federation, and supports users to fine-tune according to actual scenarios.

[0146] An early warning triggering module, in communication connection with the classification and judgment module, is configured to perform step S7: when the state is determined as drowning struggle and the duration of this state exceeds a preset threshold, trigger a corresponding early warning instruction. This module supports docking with various peripheral devices such as smart bracelets, monitoring large screens, sound and light alarm devices, and realizes multi-terminal linked early warning.

[0147] Embodiment 3

[0148] This embodiment provides an electronic device for running the swimming pool drowning early warning method described in Embodiment 1, and the hardware structure of the electronic device is as follows:

[0149] The electronic device comprises a processor, a memory, a communication interface and a bus, wherein the processor, the memory and the communication interface communicate with each other through the bus.

[0150] The bus can be an Industry Standard Architecture (ISA) bus, a PCI bus, or an Extended Industry Standard Architecture (EISA) bus, etc., and can be divided into address bus, data bus, control bus, etc.

[0151] The memory stores at least one executable instruction that causes the processor to perform the operation corresponding to the swimming pool drowning warning method described in Embodiment 1. The memory can be high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. In this embodiment, the memory, with a capacity of 1 terabyte (TB), is deployed in an edge computing server located locally at the swimming pool venue and is used to store program code, model parameters, and alarm records.

[0152] The communication interface is used to enable communication between electronic devices and peripheral devices such as underwater cameras, smart bracelets, and stadium screens. It supports multiple communication protocols such as Ethernet, WiFi, 4G / 5G, and Bluetooth.

[0153] The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. In this embodiment, the processor is an edge computing processor equipped with an artificial intelligence (AI) acceleration unit, with a computing power of 32 trillion operations per second (TOPS), capable of simultaneously processing real-time inference of 16 video streams, meeting the system's performance requirements.

[0154] This electronic device adopts a fully local deployment mode, where all video stream data and processing are completed locally without uploading to the cloud, ensuring both real-time performance and protecting user privacy.

[0155] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application.

Claims

1. A method for early warning of drowning in swimming pools, characterized in that, Includes the following steps: S1. Acquire real-time video stream images of the swimming pool captured by the underwater camera and complete basic preprocessing; S2. Based on a single-stage real-time target detection convolutional neural network, locate human targets in the preprocessed image and output the corresponding human bounding boxes; S3. Input the image region within the bounding box into a high-resolution preservative convolutional neural network to extract key point data of the human skeleton; S4. Perform underwater optical path refraction inverse mapping correction processing on the preprocessed real-time video stream image based on generative adversarial network to restore the true proportions and posture of the human body. S5. Based on the skeletal key point data of continuous multi-frame images, combined with the corrected real human posture, the spatial posture features of a single frame and the temporal motion features of continuous frames are extracted simultaneously. S6. Based on the single-frame spatial posture features and continuous-frame temporal motion features, and combined with the preset posture classification rules that conform to the International Lifesaving League's definition of instinctive drowning postures, the human posture is classified and judged to distinguish between normal swimming state, active diving state and drowning struggle state. S7. When it is determined that the person is in a state of drowning and struggling and the duration of this state exceeds a preset threshold, a corresponding warning instruction is triggered.

2. The swimming pool drowning early warning method according to claim 1, characterized in that, The basic preprocessing in step S1 includes deblurring the real-time video stream image, underwater illumination correction, and foreground segmentation.

3. The swimming pool drowning early warning method according to claim 1, characterized in that, In step S3, the key point data of the human skeleton is extracted. Specifically, seventeen key points of the standard human skeleton are extracted, including the key points of the nose, eyes and ears of the head, the key points of the shoulder, elbow and wrist of the upper limb, and the key points of the hip, knee and ankle of the lower limb. The coordinates and confidence of each key point are output simultaneously.

4. The swimming pool drowning early warning method according to claim 1, characterized in that, The underwater optical path refraction inverse mapping correction process in step S4 specifically employs a generative adversarial network model that includes a generator and a discriminator. The generator outputs the corrected image, and the discriminator completes adversarial training optimization to achieve nonlinear distortion elimination and visual interference removal of the underwater image.

5. The swimming pool drowning early warning method according to claim 1, characterized in that, The single-frame spatial pose features in step S5 include the angle between the line connecting the torso, shoulders, and hips and the water surface, the angle between the head and the torso, the elbow angle of the upper limbs, and the vertical height difference between the key points of the shoulders and hips.

6. The swimming pool drowning early warning method according to claim 1, characterized in that, The continuous frame temporal motion features in step S5 include the periodicity of key point displacement, the range of joint angle changes, and the displacement pattern of the human body's center of mass; the preset threshold in step S7 is set to three to eight seconds.

7. A swimming pool drowning early warning system, characterized in that, The method includes the following modules, which are sequentially connected in communication, for implementing the swimming pool drowning early warning method according to any one of claims 1 to 6: The image acquisition and preprocessing module is used to perform step S1, acquire the real-time video stream image of the swimming pool captured by the underwater camera, and complete basic preprocessing. The human body detection module is used to perform step S2, which locates human targets in the preprocessed image based on a single-stage real-time target detection convolutional neural network and outputs the corresponding human body bounding box. The key point extraction module is used to perform step S3, inputting the image region within the bounding box into a high-resolution preserving convolutional neural network to extract key point data of the human skeleton. The distortion correction module is used to perform step S4, which performs underwater optical path refraction inverse mapping correction processing based on generative adversarial network on the preprocessed real-time video stream image to restore the true proportions and posture of the human body. The feature extraction module is used to perform step S5, which extracts single-frame spatial pose features and continuous-frame temporal motion features based on the skeletal key point data of multiple consecutive frames of images and combined with the corrected real human posture. The classification and decision module is used to execute step S6, and classify and decide the human posture based on the single-frame spatial posture features and the continuous frame temporal motion features, combined with the preset posture classification rules that conform to the International Lifesaving League's definition of instinctive drowning posture, and distinguish between normal swimming state, active diving state and drowning struggle state. The warning triggering module is used to execute step S7. When it is determined that the person is in a state of drowning and struggling and the duration of the state exceeds a preset threshold, the corresponding warning instruction is triggered.

8. The swimming pool drowning early warning system according to claim 7, characterized in that, The basic preprocessing performed by the image acquisition and preprocessing module includes deblurring of real-time video stream images, underwater illumination correction, and foreground segmentation.

9. The swimming pool drowning early warning system according to claim 7, characterized in that, The corresponding early warning instructions triggered by the early warning triggering module include at least one of the following: sending a vibration reminder and alarm location coordinates to the smart terminal worn by the lifeguard, displaying alarm area information on the venue's monitoring screen, and triggering an on-site audible and visual alarm device.

10. An electronic device, characterized in that, It includes a processor, a memory, a communication interface, and a bus, wherein the processor, memory, and communication interface communicate with each other via the bus; the memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the swimming pool drowning warning method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Anti-drowning early warning method and system based on lightweight human body posture estimation model

    CN116152928A