Drowning detection method and device and storage medium

By using multi-view video data processing and cross-frame tracking technology, we have achieved comprehensive, blind-spot-free drowning detection in swimming venues and accurate drowning monitoring in multi-person scenarios. This solves the problems of reliance on manpower, high cost, and poor environmental adaptability in existing technologies, and improves the accuracy of drowning detection and the level of safety management.

CN120997765APending Publication Date: 2025-11-21SHENZHEN HEISHIBAN TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511092918.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing drowning detection technologies in swimming venues suffer from problems such as insufficient reliance on manpower, high costs, poor environmental adaptability, and low accuracy, making it difficult to achieve accurate drowning monitoring in all areas and in multi-person scenarios without blind spots.

Method used

The method employs multi-view video data acquisition, preprocessing, target detection model recognition, multi-view fusion, and cross-frame tracking. After acquiring and preprocessing multi-view video data, target detection is performed to generate target detection results from each viewpoint. Multi-view fusion and cross-frame tracking are then performed in a virtual projection pool overlooking the viewpoint to assess the risk of drowning.

Benefits of technology

It enables continuous monitoring of the entire swimming pool area without blind spots, improves detection accuracy and environmental adaptability, solves the problem of accurate identification and trajectory recording of individuals in multi-person scenarios, reduces reliance on scarce real drowning data, reduces false alarms and missed alarms, and improves the safety management level of swimming venues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997765A_ABST
    Figure CN120997765A_ABST
Patent Text Reader

Abstract

The invention discloses a drowning detection method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining multi-view video data corresponding to a target swimming pool area; preprocessing the multi-view video data, and detecting the preprocessed video data of each view through a target detection model to generate a target detection result of each view; mapping the target detection result of each view angle to a bird's-eye view virtual projection swimming pool, and carrying out multi-view fusion to obtain a fused target detection result; tracking the target object in the continuous multi-frame video frame image based on the fusion target detection result; and based on the tracking result, determining whether the target object has a drowning risk. In the embodiment of the invention, by fusing the multi-view detection results and combining AI intelligent analysis and dynamic tracking collaborative design, the persistence, environmental adaptability, multi-person scene adaptability and judgment reliability of drowning prevention monitoring of the swimming pool are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision and artificial intelligence technology, and in particular to a drowning detection method, device and storage medium. Background Technology

[0002] Swimming is a popular fitness and leisure activity, frequently conducted in various swimming venues. However, monitoring the safety of the pool environment remains a challenge. Currently, drowning prevention monitoring in swimming venues mainly relies on the following methods, but all have significant limitations:

[0003] Traditional manual monitoring is the mainstream practice in the industry, relying on lifeguards' real-time observation and monitoring of live feeds. The fundamental flaw in this approach lies in the inherent limitations of human attention: firstly, swimming pools typically encompass multiple lanes and complex zones such as shallow and deep water, making it difficult for lifeguards to maintain continuous focus on every corner of such a large area; secondly, prolonged monitoring can lead to visual fatigue, and in crowded situations (such as peak hours), it's easy to miss subtle abnormalities in the early stages of drowning (such as struggling limbs or loss of balance), creating safety hazards. With technological advancements, AI-assisted monitoring solutions are gradually being applied. One type uses infrared cameras to detect swimmers by identifying their infrared signatures to determine whether they are above or below water, thus indirectly inferring the risk of drowning. However, this solution faces dual bottlenecks in terms of equipment cost and technology compatibility: the deployment and maintenance costs of infrared cameras are high, making it difficult to popularize them in small and medium-sized swimming venues; more importantly, infrared features can only reflect the macroscopic state of "above / below water" and cannot distinguish individual identities. In scenarios where multiple people are swimming at the same time, it is impossible to continuously track specific individuals, and misjudgments are easily caused by overlapping positions of people, making it difficult to meet actual monitoring needs.

[0004] Another type of AI-assisted solution analyzes swimmers' status through conventional monitoring images, relying on computer vision technology to detect features such as posture and movement trajectory to determine drowning risk. While this approach is theoretically suitable for the needs of swimming pool scenarios, it still suffers from significant shortcomings in environmental adaptability and data dependence. Firstly, the lighting conditions in swimming pool environments are complex (such as strong sunlight causing water surface reflection on sunny days, dim lighting at night, and underwater light refraction), making existing image detection models susceptible to lighting interference, leading to a decrease in posture recognition accuracy. Secondly, the collection of real drowning data faces ethical and practical obstacles, making it difficult to accumulate enough samples to train models. This results in a lack of accurate data support for drowning judgment rules (such as the threshold for static duration and abnormal posture standards), leading to high false alarm and false negative rates, and limiting the effectiveness of practical applications.

[0005] Therefore, given the shortcomings of existing solutions in terms of manpower reliance, cost control, environmental adaptability, and judgment accuracy, improving the safety management level of swimming venues has become an urgent problem to be solved. Summary of the Invention

[0006] Therefore, it is necessary to provide a drowning detection method, device, and storage medium to address the above-mentioned technical problems and solve at least one of the problems existing in the prior art.

[0007] In a first aspect, a drowning detection method is provided, characterized in that the method includes:

[0008] Acquire multi-view video data corresponding to the target pool area;

[0009] The multi-view video data is preprocessed, and the preprocessed video data from each viewpoint is detected by a target detection model to generate target detection results for each viewpoint.

[0010] The target detection results from each perspective are mapped onto the virtual projection pool overhead, and multi-view fusion is performed to obtain the fused target detection result;

[0011] Based on the fused target detection results, the target object is tracked in multiple consecutive video frame images;

[0012] Based on the tracking results, it is determined whether the target object is at risk of drowning.

[0013] Secondly, a drowning detection device is provided, comprising:

[0014] The multi-view video data acquisition unit is used to acquire multi-view video data corresponding to the target pool area;

[0015] The multi-view target detection result generation unit is used to preprocess the multi-view video data and detect the preprocessed video data from each view using a target detection model to generate target detection results from each view.

[0016] The target detection result generation unit is used to map the target detection results from various perspectives onto the virtual projection pool from above, and perform multi-view fusion to obtain the fused target detection result;

[0017] A cross-frame tracking unit is used to track a target object in multiple consecutive video frames based on the fused target detection results.

[0018] A drowning detection unit is used to determine whether the target object is at risk of drowning based on the tracking results.

[0019] Thirdly, a readable storage medium is provided that stores computer-readable instructions, which, when executed by a processor, implement the steps of the drowning detection method described above.

[0020] The aforementioned drowning detection method, device, and storage medium, implemented by the method, includes: acquiring multi-view video data corresponding to the target pool area; preprocessing the multi-view video data and detecting the preprocessed video data from each view using a target detection model to generate target detection results for each view; mapping the target detection results from each view onto a virtual projection pool from above and performing multi-view fusion to obtain a fused target detection result; tracking the target object in multiple consecutive video frames based on the fused target detection result; and determining whether the target object poses a drowning risk based on the tracking result. In this embodiment, by fusing multi-view detection results and combining AI intelligent analysis with dynamic tracking collaborative design, the limitations of manual monitoring attention are overcome, achieving continuous monitoring of the entire pool area without blind spots; preprocessing and multi-view fusion adapt to complex lighting environments, improving detection accuracy; leveraging the fusion result and cross-frame tracking, accurate individual identification and trajectory recording are achieved in densely populated scenes, solving the problem of tracking multiple people; and drowning risk is judged based on the dynamic features of movement and state obtained from tracking, reducing reliance on scarce real drowning data and minimizing false alarms and missed alarms. It comprehensively improves the continuity, environmental adaptability, multi-person scenario adaptability, and judgment reliability of drowning prevention monitoring in swimming venues, effectively enhancing the safety management level of swimming venues. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of an application environment for a drowning detection method according to one embodiment of this application. Figure 1 ;

[0023] Figure 2 This is a schematic diagram of an application environment for a drowning detection method according to one embodiment of this application. Figure 2 ;

[0024] Figure 3 This is a flowchart illustrating a drowning detection method according to one embodiment of this application;

[0025] Figure 4 This is a timing diagram of a drowning detection method in one embodiment of this application;

[0026] Figure 5 This is a schematic diagram of a drowning detection system according to one embodiment of this application;

[0027] Figure 6This is a scene demonstration diagram of the perspective transformation method in one embodiment of this application;

[0028] Figure 7 This is a schematic diagram of images taken from four perspectives in one embodiment of this application;

[0029] Figure 8 This is an embodiment of the present application. Figure 7 A schematic diagram showing the target detection and tracking results projected onto a virtual projection pool from four different perspectives.

[0030] Figure 9 This is a schematic diagram of a drowning detection device according to one embodiment of this application;

[0031] Figure 10 This is a schematic diagram of a computer device according to one embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0033] The drowning detection method provided in this embodiment can be applied to, for example, Figure 1 , Figure 2In its application environment, the architecture includes a data input layer, an intelligent analysis layer, a decision logic layer, and an output response layer. The data input layer comprises multiple cameras deployed at different locations within the pool, such as the poolside and from aerial views, to collect multi-view video data, covering the entire pool area, including lanes, deep water, shallow water, and poolside corners. The intelligent analysis layer includes an AI algorithm module and a time-series analysis module. AI algorithms perform target detection and status recognition on the multi-view video data collected by the cameras, such as first identifying swimmers, distinguishing between the human body and the background, and then determining the swimmer's status label, such as "above water" (e.g., floating, swimming) or "under water" (e.g., diving, drowning). The time-series analysis module continuously tracks multiple frames of data, recording the swimmer's "underwater duration" and "displacement changes" (e.g., if a person's underwater stay exceeds a threshold or their displacement is abnormally small, a risk warning will be triggered). The decision logic layer is configured with judgment rules to set risk assessment rules based on the AI-recognized "status labels" and "time-series data." For example: if the underwater duration > X seconds + displacement distance < Y meters, it is judged as "suspected drowning" (the person is still / moving slightly underwater and may have lost the ability to move independently). The output response layer is used to issue drowning alarms. For example, it can trigger a buzzer on the lifeguard's wristband / wristband to notify the lifeguard on site immediately and accurately locate the risk location (to avoid lifeguards being unable to keep an eye on the area or being distracted). It can also simultaneously trigger an audible and visual alarm in the area to remind people in the vicinity to assist, or link with the broadcast system to intervene (such as "There is a suspected drowning person on the east side of the deep water area, please send lifeguards to the scene").

[0034] In one embodiment, such as Figure 3 , Figure 4 As shown, a drowning detection method is provided, including the following steps:

[0035] In step S110, multi-view video data corresponding to the target pool area is acquired;

[0036] Optionally, multiple cameras can be deployed within the pool area, at different locations such as the poolside and from elevated views, to collect multi-angle video data, covering the entire pool area, including lanes, deep end, shallow end, and poolside corners. The system can connect to each camera via its network IP address through the corresponding SDK open platform to obtain real-time video stream data.

[0037] The number of cameras can be 2, 4, 5, or other specific numbers, which can be set according to the size of the pool or the layout of the environment. This application does not limit the number of cameras.

[0038] In addition, after acquiring multi-view video data, the video data of each view can be calibrated. For example, each vertex of the pool and the corresponding ID identifier of each vertex can be calibrated. The ID identifier is a unique identifier for each vertex, ensuring that the same pool vertex has the same ID in video footage from different viewpoints.

[0039] It should be noted that calibration can be performed manually or using a calibration object. For example, a chessboard of known size can be used as a calibration object. By taking multiple images of the chessboard at different angles and positions, and utilizing the correspondence between the world coordinates and image coordinates of the chessboard corner points, a system of equations can be established based on the pinhole imaging model and perspective transformation principles. These equations relate to the camera's intrinsic parameters (including focal length, principal point coordinates, distortion coefficients, etc.) and extrinsic parameters (rotation matrix and translation vector, used to describe the relationship between the camera coordinate system and the world coordinate system). Then, optimization algorithms such as the least squares method can be used to solve these parameters, thereby obtaining calibration information and providing a unified benchmark for subsequent multi-view coordinate mapping.

[0040] In step S120, the multi-view video data is preprocessed, and the preprocessed video data from each view is detected by a target detection model to generate target detection results for each view.

[0041] Optionally, after acquiring multi-view video data, due to the heterogeneity of multi-view videos (differences in resolution, frame rate, and lighting between different cameras), frame extraction can be performed to reduce the video frame rate and decrease redundant computation. This unifies the image resolution (resize) to adapt to the model's input requirements. Simultaneously, to address the specific interferences of the pool scene (water surface reflection, monochromatic color, and varying lighting), data augmentation can be performed. Specifically, SAM (Segment Anything Model) segmentation can be used to segment the foreground and background of the image, highlighting key areas such as people and pool walls while suppressing background interference; color images can be converted to grayscale (single-channel) to reduce color interference; overall image brightness can be increased to address the problem of excessive darkness in underwater areas; and data augmentation techniques (such as HSV adjustment, MixUp blending, and CutMix cropping and stitching) can be used to increase the diversity of training data. This ensures that the preprocessed data meets the model's input requirements and reduces the impact of environmental interference on detection.

[0042] Then, the preprocessed multi-view video frame images are input into the pre-trained target detection model. The model's feature learning ability for pool scenes (such as human contour and motion posture recognition) is used to identify and locate target objects in the videos from each viewpoint, i.e. swimmers. The model outputs structured detection results that include the target object's location (such as bounding box coordinates) and state classification (normal swimming / suspected drowning, etc.), providing basic data support for subsequent multi-view fusion, target tracking and drowning judgment.

[0043] It should be noted that the object detection model can be the YOLOv11 model, or other models capable of object detection, such as CNN and RetinaNet. Taking the YOLOv11 model as an example, its training process is as follows: First, a large amount of swimming pool video is collected, covering diverse scenes, such as different lighting (sunny daylight, cloudy day, nighttime lighting), different time periods (morning, noon, evening), different population densities (peak / off-peak), and different water surface states (calm, rippling, reflective), ensuring that the data reflects the complexity of the actual swimming pool environment. Frames are then extracted at a certain frequency (the frame extraction frequency can be determined according to the video frame rate, avoiding excessively small differences between adjacent frames (reducing data redundancy) while preserving the continuity of motion (such as keyframes of swimming movements). After frame extraction, all people in the extracted video frame image data are manually labeled (as ground truth), and the detection box for each person is drawn and categorized. Then, the person's labeling and classification are treated as a whole, and the entire person's detection box is drawn. At the same time, it is determined whether the person's mouth and nose are above or below the water, and the whole thing is categorized and labeled.

[0044] Then, the large amount of labeled data can be divided into training and testing sets (approximately 9:1). The training set data undergoes scene-specific data augmentation (such as random cropping to simulate occlusion, water ripple overlay, and lighting perturbation) to enhance the model's robustness to environmental disturbances. This data is then fed into the YOLOv11 model for multiple rounds of learning and training, and the training loss curve is observed. After a preset number of training epochs, such as 250, if the curve stabilizes, it indicates good convergence. At this point, the trained YOLOv11 model can be tested and validated on the testing set data. The actual results calculated by the model are compared with the actual labeled results (GT), and various accuracy metrics (including recall, precision, and accuracy) are calculated. When the accuracy metrics of the model's calculated results and the actual labeled results reach the expected values, training is complete. If the accuracy metrics do not meet the requirements, the reasons are analyzed, the training method is adjusted, or more data is added, and the above steps are repeated until the accuracy metrics reach the expected values, resulting in a successfully trained object detection model.

[0045] In step S130, the target detection results from each viewpoint are mapped onto the virtual projection pool from above, and multi-view fusion is performed to obtain the fused target detection result.

[0046] It should be noted that due to the different installation angles and heights of various cameras, directly stitching multiple videos together can lead to misalignment of the same person's position in different viewpoints. Therefore, by using a global coordinate system corresponding to the virtual projected pool from above, the detection results from each viewpoint (such as the coordinates of a person in a side view) can be converted into a unified global coordinate system, aligning the target positions from different viewpoints (i.e., the coordinates of the same target object coincide or are adjacent after mapping the detection results from different viewpoints). Furthermore, multi-view fusion can prevent the detection of people whose side view is obscured by the poolside from being supplemented by the detection results from the top view (solving "missed detections"); and if a person is mistakenly identified as "swimming normally" from one viewpoint but is detected as "suspected drowning" from other viewpoints, it can be corrected through weighted fusion voting, thereby improving detection accuracy.

[0047] Specifically, the process begins by using the physical boundaries of the target pool (such as the four corner vertices) as a reference. Boundary feature points are collected from various viewpoints, and the homography matrix formula is substituted to solve for the homography matrix, thus completing multi-view spatial calibration and establishing coordinate mapping rules between video data from each viewpoint and the virtual projection pool from an overhead view. Next, the obtained homography matrix is ​​used to transform the target detection results (personnel detection box coordinates, etc.) from different viewpoints into the virtual pool coordinate system, eliminating coordinate distortion caused by viewpoint differences. This normalizes the multi-view detection data to a unified virtual overhead view, providing accurate and consistent structured data for subsequent cross-view target tracking and drowning risk analysis. This achieves the integration and optimization of the entire pool detection information, compensating for issues such as occlusion and missed detections, and misjudgments that easily occur in single-view detection, thereby improving overall detection accuracy and reliability.

[0048] In step S140, based on the fused target detection results, the target object is tracked in multiple consecutive video frame images;

[0049] Optionally, the fused target detection results may include the normalized coordinates of each target object, such as the swimmer's status classification (e.g., above water, underwater, suspected drowning, static floating, etc.) in the virtual projection pool from above, the multi-view confidence (e.g., 0.9 from above, 0.6 from the side), the overall confidence after fusion (e.g., 0.8 after weighted average), and quality features (e.g., occlusion ratio, sharpness, etc.). After obtaining the fused target detection results, a preset tracking algorithm, such as OC-SORT, DeepSORT, or TransTrack, can be selected according to real-time requirements to perform cross-frame dynamic tracking of each target object in the fused target detection results. The tracking results are displayed on a large screen for management personnel to view, promptly detect the presence of drowning or other dangerous events, and take timely action.

[0050] In addition, during tracking, the center point position of the detection box of the target object in multiple frames of images [(x1+x2) / 2,(y1+y2) / 2] can be recorded as the continuous motion trajectory of the target object, where (x1,y1) is the coordinate of the upper left corner of the detection box and (x2,y2) is the coordinate of the lower right corner.

[0051] In step S150, based on the tracking results, it is determined whether the target object is at risk of drowning.

[0052] The tracking results can include motion trajectories, motion state sequences, and temporal feature data. Motion trajectories can include a continuous sequence of coordinates of the target object within the virtual projection pool, such as [(x1,y1),(x2,y2),(x3,y3),...], with each frame's coordinates associated with a unique ID for the target object. Motion state sequences refer to the sequence of state classification labels for the target object across consecutive frames, such as [above water, underwater, mouth and nose underwater, suspected drowning,...]. Temporal feature data refers to temporal features strongly correlated with the pool scene, such as the duration the target object remains underwater and the intervals between the mouth and nose emerging from the water.

[0053] Optionally, the risk analysis window is initialized, such as the hyperparameter window width W (e.g., corresponding to approximately 1-3 seconds of actual time, calculated based on the frame rate). For each tracked target object, the detection data of its past W frames is continuously extracted to calculate multiple secondary data such as displacement distance, movement speed, movement direction, and state value change trend. Then, these secondary data are combined and analyzed from the dimensions of motion and state. For example, when the displacement distance is continuously extremely small, the movement speed is abnormally slow, the movement direction is disordered and irregular, and the state value changes rapidly from normal swimming to suspected drowning or static floating, it is determined that the person is drowning or at risk of drowning, thereby achieving accurate identification of the drowning state of people in the pool.

[0054] It should be noted that the window width of the risk analysis window can be dynamically changed according to the risk level (for example, 10 frames when normal and 5 frames when suspected of being abnormal), balancing detection accuracy and real-time performance.

[0055] When drowning is detected, alarm information can be output, such as through sound and light alarms, to remind on-site staff, lifeguards, or people in the vicinity to carry out rescue in a timely manner. Alternatively, alarm information can be sent to users wearing smart bracelets or carrying smart terminals, such as through vibration, ringing, or buzzing, to alert them so that drowning rescue can be carried out in a timely manner.

[0056] like Figure 5As shown, this drowning detection method can be applied to a drowning detection system, which may include a front-end module, an algorithm module, a back-end module, and an interaction module. The front-end module includes an initialization interface and a large screen. Users can enter their IP address on the initialization interface to bind the camera and complete the connection between the system and the hardware. The large screen is used to display the tracking results. The algorithm module includes a detection algorithm and a tracking algorithm. The detection algorithm is used to detect target objects in multi-view video data collected by the cameras and output the detection results. The tracking algorithm is used to perform cross-frame tracking based on the multi-view fused detection results, generating tracking results for display on the large screen. The back-end module is used to associate camera IDs with detection results to ensure that multi-camera data is not confused; to perform noise reduction and alignment on the detection and tracking results; to integrate the detection results from multiple cameras, output the fused results, and determine whether drowning has occurred based on the tracking results, and output drowning alarm information. The interaction module includes multiple cameras for collecting multi-view video and multiple smart bracelets for receiving alarm information generated by the back-end and providing timely drowning rescue.

[0057] This application provides a drowning detection method, comprising: acquiring multi-view video data corresponding to a target pool area; preprocessing the multi-view video data and detecting the preprocessed video data from each view using a target detection model to generate target detection results for each view; mapping the target detection results from each view onto a virtual projection pool from above and performing multi-view fusion to obtain a fused target detection result; tracking the target object in multiple consecutive video frames based on the fused target detection result; and determining whether the target object poses a drowning risk based on the tracking result. In this application embodiment, by fusing multi-view detection results and combining AI intelligent analysis with dynamic tracking collaborative design, the method overcomes the attention limitations of manual monitoring, achieving continuous monitoring of the entire pool area without blind spots; by adapting to complex lighting environments through preprocessing and multi-view fusion, the detection accuracy is improved; by leveraging the fusion result and cross-frame tracking, accurate individual identification and trajectory recording are achieved in densely populated scenes, solving the problem of tracking multiple people; and the drowning risk is judged based on the dynamic features of movement and state obtained from tracking, reducing reliance on scarce real drowning data and minimizing false alarms and missed alarms. It comprehensively improves the continuity, environmental adaptability, multi-person scenario adaptability, and judgment reliability of drowning prevention monitoring in swimming venues, effectively enhancing the safety management level of swimming venues.

[0058] In one embodiment of this application, mapping the target detection results from each viewpoint onto the overhead virtual projection pool includes:

[0059] Based on the physical boundary of the target pool area, multi-view spatial calibration is performed on the video data from each perspective to establish a coordinate mapping relationship between the video data from each perspective and the virtual projection pool from above.

[0060] Based on the coordinate mapping relationship, the target detection results from each viewpoint are projected onto the virtual projection pool from above, and the target detection results from each viewpoint after projection are fused to obtain the fused target detection result.

[0061] Optionally, after acquiring multi-view video data, the physical boundary feature points of the target swimming pool (such as the four corner vertices) can be used as anchor points in the video data through manual calibration or automatic system calibration. That is, the pixel coordinates of the "corner vertices" in each view video are marked (e.g., the coordinates of corner A in the side view are (x1, y1), and the coordinates of corner A in the top view are (x1′, y1′)). Since the real-world coordinates of these physical boundary points are known (e.g., the actual coordinates of corner A are (X, Y, Z)), they can be used as calibration reference systems.

[0062] Because homography exists between images of the same plane (the pool surface can be approximated as a plane) when a camera captures images of the pool from different positions, perspective transformation can be used to represent the conversion relationship between them. For example... Figure 6 As shown, it is recognized that a planar object appears from two different perspectives O. R O L Images taken at the time can be converted to each other using the homography matrix formula.

[0063] Based on this, the pixel coordinates of the pool corner vertices obtained from different viewpoints captured by different cameras can be substituted into the homography matrix formula, as shown below:

[0064]

[0065] Yes, by using the four corner points of the pool from each labeled viewpoint as matching points, then for each set of matching points (x... i y i ) and (x' i y' i ), Then the following equation holds:

[0066]

[0067] Homography matrix H 3×3Although it contains 9 unknowns, due to the characteristic that "overall scaling does not change the transformation effect" (i.e., scale equivalence), only 8 independent parameters (8 degrees of freedom) are actually needed to determine it. Solving for these 8 degrees of freedom requires exactly 4 sets of non-collinear matching points. This is because each set of matching points provides 2 independent equations, and the 4 sets, totaling 8 equations, can be solved to obtain a unique solution. In the pool scene, selecting the 4 corner points of the pool from each viewpoint as matching points is extremely reasonable: they are not only strictly on the same plane (the pool plane), meeting the application premise of "objects on the same plane" for the homography matrix, but also naturally arranged in a rectangular shape (non-collinear), directly satisfying the solution condition of "4 sets of non-collinear matching points." Furthermore, the physical positions of the corner points are clear, accurately establishing the relationship between the image coordinates and the actual plane of the pool. Finally, through this scene-adaptive calibration method, a homography matrix that can be used for subsequent coordinate mapping is efficiently obtained, namely, a 3×3 matrix H. 3×3 .

[0068] When the homography matrix H is solved 3×3 Then, the target object detection boxes obtained from multiple different perspectives, such as four target detections (each box is defined by the top left vertex (x1, y1) and the bottom right vertex (x2, y2)), can be substituted into the mapping formula of the homography matrix (homogeneous coordinate transformation) to obtain their corresponding coordinates in a virtual two-dimensional plane (overlooking the virtual projected swimming pool). This virtual plane is not arbitrarily set; its size ratio strictly matches the actual physical size of the swimming pool (e.g., 1 pixel in the virtual plane corresponds to 0.1 meters in the real swimming pool). Therefore, the mapped detection boxes can not only achieve multi-view coordinate alignment in the virtual plane (e.g., the bounding boxes of the same target appearing in the side view and top view overlap after mapping), but also directly reflect the actual position of the target object in the real swimming pool (e.g., the coordinates of a target object in the virtual plane correspond to the center of the deep water area of ​​the real swimming pool).

[0069] It should be noted that after detecting the same scene from n perspectives, each target object can have a detection result in its corresponding perspective. After homography matrix projection, the same person will have n projections of detection results in the planar rectangle overlooking the virtual projection pool. Therefore, it is necessary to fuse the n projected detection results to determine that the n detection results are for the same target object, and to distinguish adjacent and close detection boxes as different target objects. Then, a preset matching algorithm, such as the Hungarian matching algorithm, intersection-union matching algorithm, or deep learning matching (such as the Siamese network), can be used to calculate the matching degree of the detection results (such as the detection boxes of the target object) in different perspectives (which can be measured by indicators such as positional overlap and size similarity). Then, the n matching results with the highest matching degree (usually the detection results of the same target object from different perspectives) are selected from the matching results and fused into a single detection result. The final projected position of the person is obtained by averaging the coordinates of these n detection results (or by weighted averaging according to the confidence of each perspective), thereby reducing the single-view detection error and improving the positional accuracy. This process essentially involves using algorithms to solve the correspondence between multi-view detection results and the same target, and then integrating the coordinates through statistical methods to achieve precise aggregation of target locations. For example... Figure 7 The image shows pool area images captured by cameras from four different perspectives. The detection results of the target objects in the images are projected onto a planar rectangle overlooking the virtual projected pool using a homography matrix, forming an image like... Figure 8 The target detection (Det) shown on the left illustrates this. The selected areas are labeled "under" or "above," and the accompanying numerical value (e.g., "under 0.89") represents the detection confidence score, indicating the degree to which the algorithm classifies the area as an "underwater / above-water target." The closer the value is to 1, the higher the confidence score. Its function is to identify potential "underwater human" or "above-water human" areas within a single image. Furthermore, Figure 8 The Track on the right side represents the target tracking result, indicating the tracking progress. Figure 7 Tracking results of the same target in multi-view images. By using different identifiers (such as sequence number and color), the position and "above water / underwater" state changes of the same target in consecutive video frames can be tracked, and the target's motion trajectory and behavior patterns can be analyzed.

[0070] In one embodiment of this application, fusing the target detection results from each projected viewpoint to obtain the fused target detection result includes:

[0071] Based on the spatial attributes of the video data from each viewpoint, determine the weight corresponding to each viewpoint;

[0072] Based on the weights, the target detection results from each perspective are weighted and fused to obtain the fused target detection result.

[0073] Optionally, when observing the same target from multiple perspectives (such as different cameras or shooting angles), the target detection results from each perspective may differ (e.g., positional deviation, different confidence levels). By fusing these results, the limitations of a single perspective (e.g., occlusion, angular deviation) can be offset, improving the accuracy of target detection. Specifically, the spatial attributes of the video data from each perspective, such as position (the shooting angle relative to the target's orientation (e.g., front, side, back)) and distance (the distance between the camera and the target), can be combined to determine the weight of each perspective: for example, a frontal perspective can capture the entire target more clearly (e.g., a face, complete limb shape), so it has a higher weight; side / back perspectives may have lower weights due to occlusion or incomplete features. Perspectives at a moderate distance (appropriate target size, clear details) have higher weights; perspectives that are too close (target exceeds the frame) or too far (target is blurry) have lower weights. Finally, based on the determined weights, the target detection results from each perspective, such as classification labels, confidence levels, and positional information, are weighted and fused to obtain the fused target detection result.

[0074] In one embodiment of this application, the target detection includes location information, confidence level, and classification label. The weighted fusion of the target detection results from each viewpoint based on the weights to obtain the fused target detection result includes:

[0075] The positional information of the same target object in each viewpoint is weighted according to the corresponding weights and the average value is calculated to obtain the fused coordinates;

[0076] The classification labels and confidence scores of the same target object from different perspectives are voted on according to their weights. The classification label with the highest weight is selected as the fusion classification label, and the confidence score with the highest weight is selected as the fusion confidence score.

[0077] Optionally, the location of the target object is calculated by weighted average of the coordinates of each viewpoint, and the state of the target object (such as normal swimming or suspected drowning) is determined by voting based on the state labels of each viewpoint according to weight. This results in a fusion target detection result that covers all areas of the pool (including occluded areas), highlights the advantages of clear viewpoint detection, and reduces interference from blurry viewpoints, providing high-quality data support for subsequent target tracking and drowning judgment.

[0078] Specifically, taking a target object in a swimming pool as an example, assume its coordinates are detected as (x1, y1) from a side view, with a weight of 0.3; and its coordinates are detected as (x2, y2) from a top view, with a weight of 0.7. The fused coordinates are calculated as follows: fused x-coordinate = x1 × 0.3 + x2 × 0.7, fused y-coordinate = y1 × 0.3 + y2 × 0.7. Because the top view has higher clarity (higher weight), its coordinates account for a larger proportion in the fused result, effectively reducing the positional deviation caused by occlusion and reflection from the side view (for example, a side view might misjudge a person as being "to the left of the lane line," but a top view clearly places them "to the right of the lane line," resulting in fused coordinates that are closer to the top view result). Furthermore, the weighted average is not a simple summation, but rather a balancing of errors from different perspectives through weights: if a certain perspective occasionally experiences extreme deviations (such as a 1-meter positional shift due to splashing water from the side view), because of its lower weight (0.3), the impact on the fused result is controlled within 0.3 meters, ensuring coordinate stability.

[0079] When different perspectives classify the same target object's state differently (e.g., a side view detects "normal swimming" with a weight of 0.3, while a top view detects "suspected drowning" with a weight of 0.7), the final label must be determined by the sum of the weights: the sum of the weights for "normal swimming" is 0.3, and the sum of the weights for "suspected drowning" is 0.7, therefore "suspected drowning" is selected as the fusion classification label. The classification result from a clear perspective (top view "suspected drowning") has a higher weight and a greater impact on the final label, avoiding misjudgment interference from low-resolution perspectives (e.g., a side view might be misjudged as "normal" due to glare, but its low weight in the voting process prevents it from dominating the result).

[0080] Similarly, for the classification confidence of the same target object in each perspective, a weighted vote is performed based on the weights corresponding to each perspective. The confidence of the target classification label under each perspective is multiplied by the weight of the corresponding perspective, and then the results of all perspectives are summed. Finally, the confidence value with the highest sum of weights is selected as the fused confidence value.

[0081] In one embodiment of this application, the step of detecting preprocessed video data from various perspectives using a target detection model to generate target detection results for each perspective includes:

[0082] In the multi-view video frame images, all target objects corresponding to each view image are detected, and an object detection box corresponding to each target object is generated;

[0083] The classification label and confidence level of the target object selected in each object detection box are identified, and the classification label includes at least the motion state category of the target object.

[0084] Optionally, for multi-view video frame images (such as side view, top view, oblique view, etc.), object detection models such as YOLOv11 and Faster R-CNN can be used to process the image from each viewpoint separately, identify and box all target objects (mainly people in the pool), and generate corresponding object detection boxes. The detection boxes mark the position of the person in the current viewpoint image in coordinate form (such as the upper left corner (x1, y1), the lower right corner (x2, y2)) to ensure that no person is missed (even partially occluded persons should be detected as much as possible through model optimization). While selecting target objects, image features can be extracted, such as the target object's limb posture, motion trajectory segments, and the relative position of the body to the water surface, to identify the motion state of the target object within each detection box and generate classification labels. These classification labels may include: a primary label for the overall state (normal swimming, suspected drowning, static floating), and a secondary label for the mouth and nose state (above / below water), and ensure that the label logic is consistent (such as "suspected drowning" is preferentially associated with "mouth and nose underwater and continuously").

[0085] Additionally, it can output a confidence score, which can be a value between 0 and 1. The closer the confidence score is to 1, the higher the accuracy of the classification result; conversely, the closer the confidence score is to 0, the lower the accuracy of the classification result.

[0086] In one embodiment of this application, tracking the target object in multiple consecutive video frames based on the fused target detection result includes:

[0087] Based on the fused target detection results, target object information is obtained, including the position information of the target object in the overhead virtual projection pool;

[0088] Based on the location information, the target object is tracked across multiple consecutive video frames.

[0089] Optionally, through multi-view fusion, the fused target detection results are unified into the coordinate system of the "overlooking virtual projection pool." Regardless of how the target object moves or is occluded in the original multi-view video frames, its position in the virtual pool remains physically consistent (e.g., the virtual coordinates will change accordingly when swimming from the east side to the west side of the pool). Position information and state labels can be extracted from the fused target detection results. Position information refers to the normalized coordinates of the target object in the overlooking virtual projection pool; state labels refer to the target object's state in the overlooking virtual projection pool, such as normal swimming. Then, using preset tracking algorithms, such as OC-SORT, DeepSORT, and TransTrack, cross-frame dynamic tracking of each target object in the fused target detection results can be performed.

[0090] Taking the OC-SORT algorithm as an example, first determine the center point (x0, y0) of the detection box of the target object in the current frame (F0), and predict the position (x', y') of the next frame (F1) based on the uniform speed principle; then, with the help of the Kalman filter, match the actual detection center point (x1, y1) of F1 with the predicted position to eliminate normal jitter interference. If the match is successful, it is determined to be the same target object. If occlusion occurs (not detected in frame t+1), the ID is temporarily stored based on the historical trajectory, and the association is restored when the target reappears.

[0091] Meanwhile, the OC-SORT algorithm optimizes tracking performance and solves problems such as dense crowds and easy occlusion in swimming pool scenes through three innovative methods: OOOS (Observation-Centric Online Smoothing), OCM (Observation-Centric Momentum), and OCR (Observation-Centric Recovery). It provides continuous and stable data on the movement trajectory and state sequence of people for drowning detection, and achieves accurate positioning of the same target object across frames, distinguishing between continuously moving people and objects in different positions.

[0092] It should be noted that the OC-SORT algorithm can be used to assign a unique ID to each target object included in the fused target detection results, so that the detection box of the same target object in consecutive frames is identified as the same target, solving the problem of single-frame detection crossing continuous tracking, and providing motion trajectory and state sequence data for drowning judgment.

[0093] In this embodiment of the application, the tracking result includes the movement trajectory, posture change sequence, and dwell time of the target object. Determining whether the target object is at risk of drowning based on the tracking result includes:

[0094] Continuously acquire the tracking results of the target object within a preset risk analysis window;

[0095] Based on the tracking results, the motion and state parameters of the target object within the preset risk analysis window are calculated. The motion and state parameters include at least one of displacement distance, movement speed, movement direction, and state value change trend.

[0096] Based on the motion and state parameters, it is determined whether the target object is at risk of drowning.

[0097] Optionally, since drowning is sudden (usually within a few seconds to tens of seconds), a risk analysis window can be set, i.e., a preset hyperparameter frame window width W (e.g., the most recent 10 frames, corresponding to a time slice of about 3 seconds) as the analysis unit. From the tracking results, complete data of the target object within the preset hyperparameter window width is extracted, such as motion trajectory segments [(x1,y1),(x2,y2),...,(x10,y10)], reflecting the positional changes of the target object; attitude change sequences, such as status label sequences ["normal swimming", "normal swimming", "static floating", "static floating"], recording attitude transition trends; and time series sequences, which can reflect the number of consecutive stay frames in the same area, etc.

[0098] Based on the tracking results using the preset hyperparameter frame window width W, key parameters reflecting drowning characteristics are extracted, such as displacement distance, movement speed, movement direction, and state value change trends, among other secondary data. Then, based on the calculated secondary data, combined with state changes and motion characteristics, it is determined whether a person is drowning or at risk of drowning.

[0099] For example, 5 frames per second (5fps) can be captured, for a total of 25 frames (5s) in a window. (Both the window size and the frame rate can be adjusted according to actual needs). The risk of drowning can be determined by the displacement trend and the duration of time above and below water. The displacement trend can be used as the main criterion; if there is a displacement of a certain length and with a regular direction, drowning is ruled out (because drowning victims usually cannot make regular displacements on their own and mostly make small, struggling movements).

[0100] Displacement trend = Sum of the magnitudes of vectors between each frame (Σ|vi|) / Magnitude of vectors from the first and last frames (|v0-vn|);

[0101] Where Σ|vi| represents the total distance moved between every two frames (reflecting all the small movements in the intermediate process), and |v0-vn| represents the straight-line distance from the first frame to the last frame (reflecting the final displacement result).

[0102] It should be noted that the closer the value is to 1, the clearer the movement trend. The larger the value, the less the total displacement, but there are a lot of small movements in between (such as struggling in place, which may be a characteristic of drowning).

[0103] For cases where the swimmer turns back within 5 seconds, additional judgment conditions can be added:

[0104] Take the middle frame of the window (e.g., frame 13), calculate the vector v1 from the first frame to the middle frame and the vector v2 from the middle frame to the last frame, and use (|v0-vn|) / v1 and (|v0-vn|) / v2 to determine if there is a backtracking. If these two values ​​are closer to 2, it means that there is no backtracking and the flow is continuous.

[0105] The duration of time spent above and below water can be determined by observing the changes in the "above water" (head exposed) and "under water" (head submerged) states in each frame. This indicates whether the person is maintaining a normal breathing rhythm (not drowning). Within a 25-frame window, the above-water and underwater states should change at regular intervals in each frame. Theoretically, when the frame rate (fps) is high enough and the window is large enough, the number of above-water frames should be approximately equal to the number of underwater frames, indicating regular breathing and swimming. Alternatively, if the number of above-water frames is much greater than the number of underwater frames, it indicates floating or swimming on their back.

[0106] In one embodiment of this application, determining whether the target object is at risk of drowning based on the motion and state parameters includes:

[0107] Based on the motion and state parameters, and in accordance with preset soft rules and / or hard rules, it is determined whether the target object is at risk of drowning.

[0108] The soft rule refers to threshold judgment on multiple indicators among the target object’s underwater time percentage, movement speed and acceleration, trajectory jitter, continuous underwater time, trajectory deviation, trajectory repetition frequency, directional inconsistency and path curvature. When the number of indicators that meet the soft rule reaches a preset number, it is determined that there is a risk of drowning.

[0109] The hard rule refers to the determination of drowning risk when a target object experiences a sudden descent.

[0110] Specifically, a combination of soft and hard rules can be used to accurately identify drowning risks. Soft rules are designed for progressive drowning situations. Due to the gradual abnormal characteristics of drowning (such as the gradual emergence of "slow movement and erratic trajectory" as one struggles from normal swimming), soft rules capture the "accumulation of risk before drowning" by quantifying physical parameters such as the person's movement trajectory and underwater state. Hard rules are designed for sudden drowning situations. When a sudden abnormality occurs during drowning (such as a person suddenly sinking or disappearing due to drowning), a high-priority warning is directly triggered. Drowning is determined when ≥4 soft rules are met or when a hard rule is triggered, covering both slow drowning (such as gradual sinking due to exhaustion) and sudden drowning (such as sudden descent due to a sudden illness).

[0111] The soft rules can include eight items, specifically: most of the time underwater, low speed and low acceleration, path jitter intensity, continuous underwater time, trajectory deviation, trajectory repetition frequency, directional inconsistency, and path curvature ratio. If any four or more of these items meet the corresponding preset conditions, it indicates a risk of drowning.

[0112] "Most of the time underwater" refers to normal swimming, where a person needs to frequently surface for air. In the event of drowning, due to exhaustion and panicked struggle, a person lacks the strength to lift their head above water, leaving their mouth and nose submerged for an extended period. The frame ratio of the head underwater within the sliding window can be obtained using the following formula:

[0113]

[0114] Where W represents the total number of frames within the sliding window, h i This represents the head state in the i-th frame. If the proportion of frames in the sliding window where the head is underwater is higher than 70%, that is, R... sub A value ≤0.7 indicates that the subject has been submerged for an extended period and is unable to breathe normally, posing a risk of drowning.

[0115] Low speed + low acceleration refers to the fact that after drowning, a person lacks the strength to paddle, and their speed drops drastically to a very low level. At the same time, because they cannot actively exert force to change their state of motion, their acceleration also approaches zero. This is different from being at rest during normal rest, where one actively stops and can accelerate at any time if they want to move. Drowning, however, means that one wants to move but cannot. Therefore, it can be judged by average speed and average acceleration.

[0116] The average velocity can be calculated using the following formula:

[0117]

[0118] Where W represents the total number of frames within the sliding window, ||p i -p i-1 || represents the Euclidean distance between the i-th frame and the (i-1)-th frame, reflecting the displacement between two adjacent frames, and Δt represents the time interval between two adjacent frames.

[0119] The average acceleration can be calculated using the following formula:

[0120]

[0121] Among them, ||p i -p i-1 ||-||p i-1 -p i-2 || represents the difference between two adjacent displacements, and Δt represents the time interval between the two displacements.

[0122] If both the average velocity and average acceleration are low, such as V avg <0.2and A avg <0.1 indicates that the subject is stationary or has limited movement in the water, and is likely to lose self-control.

[0123] The intensity of path jitter refers to the violent shaking of the movement trajectory caused by loss of body control during drowning struggles. It is determined by calculating the change in direction, which can be obtained using the following formula:

[0124]

[0125] Where W represents the total number of frames within the sliding window, θ i θ represents the orientation angle of the i-th frame. i-1 This represents the direction angle of the (i-1)th frame.

[0126] θ i It can be calculated using the following formula:

[0127] θ i =atan2(y i -y i-1 x i -x i-1 );

[0128] Among them, y i This represents the position coordinates of the target object in the vertical direction (e.g., the y-axis of the image coordinate system) in the i-th frame. i-1 The x-coordinate represents the vertical position coordinate of the target object in the (i-1)th frame. i This represents the horizontal coordinates (e.g., the x-axis of the image coordinate system) of the target object in the i-th frame. i-1 represents the horizontal position coordinates of the subject in the (i-1)th frame, and atan2 represents the tangent function.

[0129] If J > 0.4, it indicates severe path jitter, suggesting that the main body's direction is changing repeatedly and it may be struggling in the water.

[0130] Continuous underwater time refers to the duration during which a drowning person is unable to lift their head, resulting in their mouth and nose being covered by water for an extended period. Monitoring the duration of continuous underwater submersion beyond a preset time indicates that the person may no longer be able to breathe actively, posing a risk of drowning. Longer continuous underwater time can be calculated using the following formula:

[0131] L ma x = max{l:h} k =0,h k-1 =0,...,h k-l+1 =0}·Δt;

[0132] Where l is used to count the number of consecutive underwater frames, h k h k-1 h k-l+1 : These represent the state of the head being underwater at frame k, frame k-1, and frame k-l+1, respectively, and Δt represents the time interval corresponding to each frame.

[0133] If L max ≥2.0 indicates that the person remained underwater for more than 2 seconds in multiple consecutive frames, indicating that they were unable to breathe independently.

[0134] Trajectory deviation refers to the phenomenon where, after drowning, a person loses directional control and drifts in unusual directions such as pool walls or corners, causing their trajectory to deviate from the normal path. This can be determined by the deviation variance of the target object's centroid position, which can be calculated using the following formula:

[0135]

[0136] Where W represents the total number of frames within the sliding window. The centroid position is represented by p, which corresponds to the position of all frames within the sliding window. i The average value, p i This represents the position vector of the target object in the i-th frame.

[0137] The deviation variance can be calculated using the following formula:

[0138]

[0139] in, D represents the Euclidean distance between the centroid position and the position vector. If D≤0.1, it means that the target object is shaking or staying in one area for a long time, indicating that it is in a struggling and floating state.

[0140] The frequency of trajectory repetition refers to the repeated overlapping of trajectories caused by a person struggling while drowning, lacking the strength to move and only able to sway back and forth within a small area. A hash table can be used to statistically analyze trajectory projections, identifying continuous trajectory points p. i =(x i y i Discretize the data into a two-dimensional raster space (e.g., a 0.1m × 0.1m cell) and count how many cells were accessed more than twice. The trajectory repetition frequency can be calculated using the following formula:

[0141]

[0142] Where K represents the total number of grid cells in the two-dimensional grid space (or the total number of grid cells corresponding to the grid range to be counted), n k This indicates the number of times the k-th grid cell is visited by the trajectory point.

[0143] If C≥3, it means that at least 3 grids are visited repeatedly, reflecting that there is a lot of repetition in the main trajectory, and it is likely in a state of "shaking / staying in one area, struggling and floating" (normal swimming trajectory is highly directional and has few grids visited repeatedly; the trajectory is more chaotic and repetitive when drowning and struggling).

[0144] Inconsistent direction refers to a significant deviation between the current velocity direction and the overall trajectory direction. Consistency can be measured by the average directional angle.

[0145] Total direction vector (window start and end points):

[0146] Δp=p W -p1;

[0147] Where W represents the total number of frames within the sliding window, p W p1 represents the position vector of the target object in the last frame of the sliding window, and p2 represents the position vector of the target object in the first frame of the sliding window.

[0148] The angle deviation between the velocity vector direction and the total direction in each frame:

[0149]

[0150] in, Let ||Δp|| represent the velocity vector of the target object in the i-th frame, and ||Δp|| represent the magnitude of the total direction vector. Represents the velocity vector The length of the module.

[0151] Average directional deviation angle:

[0152]

[0153] like, This indicates that the current velocity direction deviates significantly from the overall trajectory direction, consistent with the chaotic movement direction characteristic of a drowning person struggling; during normal swimming, this value is even smaller.

[0154] The path curvature ratio refers to the fact that a normal swimming path is usually close to a straight line, while the path during drowning struggles becomes more complex, exhibiting spiral, circular, or zigzag patterns. The degree of curvature is measured by "actual path length / length of the line connecting the start and end points." The actual path length can be expressed by the following formula:

[0155]

[0156] Where W represents the total number of frames within the sliding window, p i p represents the position vector of the target object in the i-th frame. i-1 This represents the position vector of the target object in the (i-1)th frame.

[0157] The straight-line distance (between the first and last points) can be expressed by the following formula:

[0158] L straight =||p W -p1||;

[0159] Where, p Wp1 represents the position vector of the target object in the last frame of the sliding window, and p2 represents the position vector of the target object in the first frame of the sliding window.

[0160] The curvature ratio can be expressed by the following formula:

[0161]

[0162] If T > 2.5, the main body's movement path is curved and complex, which is consistent with the characteristics of a complex pattern such as spiral, circle, and reversal when struggling in drowning. Based on this, it can be judged that the body may be in a state of drowning.

[0163] One of the hard rules is a sudden descent event. For example, when a person is swimming normally, their body floats near the surface with minimal vertical displacement. However, in the event of drowning (e.g., a sudden heart attack or leg cramps), buoyancy is suddenly lost, causing rapid sinking and a drastic change in vertical displacement, or the person disappears entirely from the monitoring screen. This "sudden descent" is an emergency warning signal for drowning. Once detected, drowning is immediately diagnosed without waiting for soft rules to accumulate, as sudden drowning develops rapidly and requires immediate alert. Specifically, the first half of the underwater distance can be calculated, and the second half can be used for judgment. The first half of the underwater distance can be calculated using the following formula:

[0164]

[0165] The underwater proportion of the second half can be calculated using the following formula:

[0166]

[0167] Where W represents the total number of frames within the sliding window, h i This indicates the state of the head underwater in the i-th frame.

[0168] like That is, the head spends most of its time above water (underwater percentage <20%) in the first half and most of its time underwater (underwater percentage ≥80%) in the second half, with at least one continuous underwater state for the head. max The statement indicates that the first half of the incident occurred while the person was mostly above water, and then suddenly dropped to the bottom, remaining submerged for more than one second without surfacing. This aligns with the characteristics of drowning, where the head suddenly sinks and the person is unable to surface on their own. Therefore, it is classified as a "drop event," indicating a risk of drowning.

[0169] In this embodiment, by fusing multi-view detection results and combining AI intelligent analysis with dynamic tracking collaborative design, the limitations of attention in manual monitoring are overcome, achieving continuous monitoring of the entire swimming pool area without blind spots. Preprocessing and multi-view fusion are adapted to complex lighting environments, improving detection accuracy. Through fusion results and cross-frame tracking, accurate individual identification and trajectory recording are achieved in densely populated scenarios, solving the problem of tracking multiple people. Drowning risk is assessed based on the dynamic features of movement and status obtained from tracking, reducing reliance on scarce real drowning data and minimizing false alarms and missed alarms. This comprehensively improves the continuity, environmental adaptability, multi-person scenario adaptability, and judgment reliability of drowning prevention monitoring in swimming venues, effectively enhancing the safety management level of swimming venues.

[0170] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0171] In one embodiment, a drowning detection device is provided, which corresponds one-to-one with the drowning detection methods described in the above embodiments. For example... Figure 9 As shown, the drowning detection device includes a multi-view video data acquisition unit 10, a multi-view target detection result generation unit 20, a fused target detection result generation unit 30, a cross-frame tracking unit 40, and a drowning detection unit 50. Detailed descriptions of each functional module are as follows:

[0172] The multi-view video data acquisition unit 10 is used to acquire multi-view video data corresponding to the target pool area;

[0173] The multi-view target detection result generation unit 20 is used to preprocess the multi-view video data and detect the preprocessed video data from each view using a target detection model to generate target detection results from each view.

[0174] The target detection result generation unit 30 is used to map the target detection results from each viewpoint onto the virtual projection pool from above, and perform multi-view fusion to obtain the fused target detection result.

[0175] The cross-frame tracking unit 40 is used to track the target object in multiple consecutive video frame images based on the fused target detection results;

[0176] The drowning detection unit 50 is used to determine whether the target object is at risk of drowning based on the tracking results.

[0177] In one embodiment of this application, the target detection result generation unit 30 is further configured to:

[0178] Based on the physical boundary of the target pool area, multi-view spatial calibration is performed on the video data from each perspective to establish a coordinate mapping relationship between the video data from each perspective and the virtual projection pool from above.

[0179] Based on the coordinate mapping relationship, the target detection results from each viewpoint are projected onto the virtual projection pool from above, and the target detection results from each viewpoint after projection are fused to obtain the fused target detection result.

[0180] In one embodiment of this application, the multi-view video data includes a top-down view and a side-view view. The target detection result generation unit 30 is further used for:

[0181] Based on the spatial attributes of the video data from each viewpoint, determine the weight corresponding to each viewpoint;

[0182] Based on the weights, the target detection results from each perspective are weighted and fused to obtain the fused target detection result.

[0183] In one embodiment of this application, the target detection includes location information, confidence level, and classification label. The target detection result generation unit 30 is further configured to:

[0184] The positional information of the same target object in each viewpoint is weighted according to the corresponding weights and the average value is calculated to obtain the fused coordinates;

[0185] The classification labels and confidence scores of the same target object from different perspectives are voted on according to their weights. The classification label with the highest weight is selected as the fusion classification label, and the confidence score with the highest weight is selected as the fusion confidence score.

[0186] In one embodiment of this application, the cross-frame tracking unit 40 is further configured to:

[0187] In the multi-view video frame images, all target objects corresponding to each view image are detected, and an object detection box corresponding to each target object is generated;

[0188] The classification label and confidence level of the target object selected in each object detection box are identified, and the classification label includes at least the motion state category of the target object.

[0189] In one embodiment of this application, the multi-view target detection result generation unit 20 is further configured to:

[0190] Based on the fused target detection results, target object information is obtained, including the position information of the target object in the overhead virtual projection pool;

[0191] Based on the location information, the target object is tracked across multiple consecutive video frames.

[0192] In one embodiment of this application, the drowning detection unit 50 is further configured to:

[0193] Continuously acquire the tracking results of the target object within a preset risk analysis window;

[0194] Based on the tracking results, the motion and state parameters of the target object within the preset risk analysis window are calculated. The motion and state parameters include at least one of displacement distance, movement speed, movement direction, and state value change trend.

[0195] Based on the motion and state parameters, it is determined whether the target object is at risk of drowning.

[0196] In one embodiment of this application, the drowning detection unit 50 is further configured to:

[0197] Based on the motion and state parameters, and in accordance with preset soft rules and / or hard rules, it is determined whether the target object is at risk of drowning.

[0198] The soft rule refers to threshold judgment on multiple indicators among the target object’s underwater time percentage, movement speed and acceleration, trajectory jitter, continuous underwater time, trajectory deviation, trajectory repetition frequency, directional inconsistency and path curvature. When the number of indicators that meet the soft rule reaches a preset number, it is determined that there is a risk of drowning.

[0199] The hard rule refers to the determination of drowning risk when a target object experiences a sudden descent.

[0200] In this embodiment, by fusing multi-view detection results and combining AI intelligent analysis with dynamic tracking collaborative design, the limitations of attention in manual monitoring are overcome, achieving continuous monitoring of the entire swimming pool area without blind spots. Preprocessing and multi-view fusion are adapted to complex lighting environments, improving detection accuracy. Through fusion results and cross-frame tracking, accurate individual identification and trajectory recording are achieved in densely populated scenarios, solving the problem of tracking multiple people. Drowning risk is assessed based on the dynamic features of movement and status obtained from tracking, reducing reliance on scarce real drowning data and minimizing false alarms and missed alarms. This comprehensively improves the continuity, environmental adaptability, multi-person scenario adaptability, and judgment reliability of drowning prevention monitoring in swimming venues, effectively enhancing the safety management level of swimming venues.

[0201] Specific limitations regarding drowning detection devices can be found in the limitations of drowning detection methods described above, and will not be repeated here. Each module in the aforementioned drowning detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0202] In one embodiment, a computer device is provided, which may be a terminal device, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium storing computer-readable instructions. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer-readable instructions implement a drowning detection method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0203] In this application embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the drowning detection method described above.

[0204] In one embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, they implement the steps of the drowning detection method described above.

[0205] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0206] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0207] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for detecting drowning, characterized in that, The method includes: Acquire multi-view video data corresponding to the target pool area; The multi-view video data is preprocessed, and the preprocessed video data from each viewpoint is detected by a target detection model to generate target detection results for each viewpoint. The target detection results from each perspective are mapped onto the virtual projection pool overhead, and multi-view fusion is performed to obtain the fused target detection result; Based on the fused target detection results, the target object is tracked in multiple consecutive video frame images; Based on the tracking results, it is determined whether the target object is at risk of drowning.

2. The drowning detection method as described in claim 1, characterized in that, The process of mapping the target detection results from each viewpoint onto the virtual projection pool from an overhead view includes: Based on the physical boundary of the target pool area, multi-view spatial calibration is performed on the video data from each perspective to establish a coordinate mapping relationship between the video data from each perspective and the virtual projection pool from above. Based on the coordinate mapping relationship, the target detection results from each viewpoint are projected onto the virtual projection pool from above, and the target detection results from each viewpoint after projection are fused to obtain the fused target detection result.

3. The drowning detection method as described in claim 2, characterized in that, The process of fusing the target detection results from each projected viewpoint to obtain the fused target detection result includes: Based on the spatial attributes of the video data from each viewpoint, determine the weight corresponding to each viewpoint; Based on the weights, the target detection results from each perspective are weighted and fused to obtain the fused target detection result.

4. The drowning detection method as described in claim 3, characterized in that, The target detection includes location information, confidence level, and classification label. Based on the weights, the target detection results from each viewpoint are weighted and fused to obtain the fused target detection result, including: The positional information of the same target object in each viewpoint is weighted according to the corresponding weights and the average value is calculated to obtain the fused coordinates; The classification labels and confidence scores of the same target object from different perspectives are voted on according to their weights. The classification label with the highest weight is selected as the fusion classification label, and the confidence score with the highest weight is selected as the fusion confidence score.

5. The drowning detection method as described in claim 1, characterized in that, The step of detecting targets in preprocessed video data from various perspectives using a target detection model to generate target detection results for each perspective includes: In the multi-view video frame images, all target objects corresponding to each view image are detected, and an object detection box corresponding to each target object is generated; The classification label and confidence level of the target object selected in each object detection box are identified, and the classification label includes at least the motion state category of the target object.

6. The drowning detection method as described in claim 1, characterized in that, The step of tracking the target object in multiple consecutive video frames based on the fused target detection results includes: Based on the fused target detection results, target object information is obtained, including the position information of the target object in the overhead virtual projection pool; Based on the location information, the target object is tracked across multiple consecutive video frames.

7. The drowning detection method according to any one of claims 1-6, characterized in that, The tracking results include the target object's movement trajectory, posture change sequence, and dwell time. Determining whether the target object is at risk of drowning based on the tracking results includes: Continuously acquire the tracking results of the target object within a preset risk analysis window; Based on the tracking results, the motion and state parameters of the target object within the preset risk analysis window are calculated. The motion and state parameters include at least one of displacement distance, movement speed, movement direction, and state value change trend. Based on the motion and state parameters, it is determined whether the target object is at risk of drowning.

8. The drowning detection method as described in claim 7, characterized in that, The determination of whether the target object is at risk of drowning based on the motion and state parameters includes: Based on the motion and state parameters, and in accordance with preset soft rules and / or hard rules, it is determined whether the target object is at risk of drowning. The soft rule refers to threshold judgment on multiple indicators among the target object’s underwater time percentage, movement speed and acceleration, trajectory jitter, continuous underwater time, trajectory deviation, trajectory repetition frequency, directional inconsistency and path curvature. When the number of indicators that meet the soft rule reaches a preset number, it is determined that there is a risk of drowning. The hard rule refers to the determination of drowning risk when a target object experiences a sudden descent.

9. A drowning detection device, characterized in that, The device includes: The multi-view video data acquisition unit is used to acquire multi-view video data corresponding to the target pool area; The multi-view target detection result generation unit is used to preprocess the multi-view video data and detect the preprocessed video data from each view using a target detection model to generate target detection results from each view. The target detection result generation unit is used to map the target detection results from various perspectives onto the virtual projection pool from above, and perform multi-view fusion to obtain the fused target detection result; A cross-frame tracking unit is used to track a target object in multiple consecutive video frames based on the fused target detection results. A drowning detection unit is used to determine whether the target object is at risk of drowning based on the tracking results.

10. A readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by a processor, they implement the steps of the drowning detection method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Visual monitoring method and system for operation state of pier anti-collision facility

    CN121935632A

  • Multi-vision lifesaving system and lifesaving method thereof

    CN121990138A