Multi-mode fusion swimming pool drowning detection system and method based on improved YOLOv5
Through multimodal data acquisition and improvement of YOLOv5 detection model, combined with IoT linked rescue, the accuracy, timeliness and rescue linkage problems of swimming pool drowning detection are solved, and efficient drowning risk identification and rapid rescue are achieved.
Patent Information
- Application Number
- CN202510453922.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing swimming pool drowning detection technology has shortcomings in accuracy, timeliness and rescue linkage. Traditional manual monitoring is prone to missed inspections. A single visual model has low detection accuracy in complex water surface environments. A single modal data cannot accurately distinguish between normal postures and drowning postures, and rescue methods are highly lagging.
The multimodal data acquisition module is used to combine the improved YOLOv5 detection model, and integrate a wide-angle camera, infrared thermal imager, ToF camera and attitude capture unit. The multimodal features are fused through the CBAM attention mechanism and the BRA module, and the human body movement is analyzed in combination with the LSTM timing model. The red light positioning and autonomous navigation life-saving robot are triggered through the Internet of Things linkage rescue module.
It improves the accuracy and timeliness of drowning detection, achieves seamless connection from detection to rescue, significantly reduces the harm of swimming pool drowning accidents, and ensures the safety of drowning people's lives.
Smart Images

Figure CN120388420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent security and computer vision, and more particularly to a multi-modal fusion swimming pool drowning detection system and method based on improved YOLOv5. Background Art
[0002] In modern society, as an important place for people's daily leisure, entertainment and fitness, the safety guarantee of swimming pools is of vital importance. However, there are many problems to be solved urgently in traditional pool monitoring methods.
[0003] Limitations of manual monitoring: Traditional pool monitoring relies heavily on manual observation. It is difficult for people to maintain a high level of concentration for a long time, resulting in frequent missed detections. A large number of drowning risks cannot be discovered in time. There is an obvious delay in manual response from detecting drowning signs to taking effective measures, which is crucial in drowning rescue.
[0004] Defects of single vision models: Single vision models represented by the widely used YOLO series perform poorly in complex water surface scenes. The fluctuations of the water surface, the reflection and refraction of light often interfere with the model's recognition of targets. Strong direct light or backlight environments have a great impact on the detection accuracy. Moreover, for small targets, such as the limbs of a drowning person partially exposed above the water surface, the detection accuracy is low, which greatly affects the timely judgment of drowning incidents.
[0005] Insufficiency of single-modal data: Most existing drowning detection systems rely only on single-modal data, such as RGB images. This approach ignores the importance of multi-dimensional information such as human body postures and infrared thermal imaging. Due to the lack of comprehensive analysis of multi-dimensional information, the system cannot accurately distinguish normal swimming postures from drowning postures, resulting in a large number of false alarms and bringing great trouble to pool management.
[0006] Lag of rescue means: Most current drowning detection systems can only trigger alarms in terms of functions and lack an effective linkage mechanism with automatic rescue equipment. Even when the alarm goes off, it still takes some time for the rescue personnel to arrive at the scene and carry out the rescue. During this period, the life safety of the drowning person is faced with a great threat.
[0007] In summary, the existing pool drowning detection technologies have serious deficiencies in terms of accuracy, timeliness and rescue linkage. Based on this, a multi-modal fusion swimming pool drowning detection system and method based on an improved YOLOv5 are proposed to solve these problems. Summary of the Invention
[0008] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-modal fusion swimming pool drowning detection system and method based on improved YOLOv5 to solve the problems existing in the above-mentioned background art.
[0009] The present invention provides the following technical solutions: A multi-modal fusion swimming pool drowning detection system and method based on improved YOLOv5, including:
[0010] A multi-modal data acquisition module for collecting data of multiple modalities, including a visual sensor, a depth sensor, and a pose capture unit; the visual sensor deploys a wide-angle camera and an infrared thermal imager at the same time to obtain the RGB image and the thermal map of the pool area respectively; the depth sensor uses a ToF camera, and by measuring the time of flight t of light, according to the formula the distance information d between the human body and the water surface is obtained, where c is the speed of light; the pose capture unit combines the MediaPipe model to extract the human body bone key point information in real time;
[0011] An improved YOLOv5 detection model, whose backbone network replaces the original CSP module with a lightweight GSConv module, and integrates the CBAM attention mechanism before the SPPF layer; at the same time, there is a BRA module for multi-modal feature fusion, and a pose analysis sub-network added after the detection head, which analyzes human actions through the LSTM time series model; in the CBAM attention mechanism, the channel attention module performs global average pooling and global max pooling on the input feature map F to obtain and After being processed by a multi-layer perceptron (MLP), the formula is used to generate the channel attention map M c (F), where σ is the sigmoid activation function, which is used to enhance the attention to important features in the channel dimension;
[0012] An Internet of Things linked rescue module with positioning and alarm functions. When drowning is detected, it triggers a red light laser device to mark the position and sends the coordinates to the administrator terminal through the NB-IoT module; it also includes a surface floating robot equipped with a GPS and an electromagnetic release life buoy, which is used to automatically navigate to the target point to release the rescue equipment.
[0013] Further, in the improved YOLOv5 detection model, the BRA module and the CBAM attention mechanism work together to improve the multi-modal feature fusion efficiency, enhance the extraction of small target features and the feature expression after the fusion of different modal data; when the BRA module fuses multi-modal data, taking the RGB image feature F RGB , the thermal map feature F thermal , and the depth information feature F depth as an example, through a specific fusion formula F fused =aF RGB +βF thermal +γF depth an enhanced feature map F fused, improve the extraction of small target features and the feature expression after the fusion of different modality data. α, β, and γ are fusion weights, which are obtained through training and learning.
[0014] Furthermore, the red light positioning device in the Internet of Things linkage rescue module can accurately mark the drowning position and provide intuitive visual guidance for rescue personnel; the autonomous navigation rescue robot quickly and accurately drives to the drowning point and releases rescue equipment according to the received coordinates.
[0015] Furthermore, the visual sensor in the multi-modal data acquisition module, in low light environments, fuses and analyzes the thermal map obtained by the infrared thermal imager and the RGB image obtained by the wide-angle camera, effectively improving the accuracy of target detection and making up for the deficiencies of single visual data under complex lighting conditions.
[0016] Furthermore, when the improved YOLOv5 detection model processes small targets, it focuses on small target features through the CBAM attention mechanism, and combines the BRA module to fuse multi-modal data such as RGB images, thermal maps, and depth information, significantly improving the detection accuracy of small targets.
[0017] Furthermore, a multi-modal fusion swimming pool drowning detection method based on the improved YOLOv5 is provided, including the following steps:
[0018] Data preprocessing and enhancement step: Use guided filtering and MSRCR algorithms to process underwater images, enhance their contrast and solve the problem of color distortion; perform non-uniformity correction on infrared images to eliminate thermal noise interference. In guided filtering, for the input image I and the guidance image p, the output filtered image q satisfies the formula where ω i is the window centered on pixel i, and a j and b j are coefficients calculated from the pixels within the window;
[0019] Multi-modal target detection step: Input multi-source data into the improved YOLOv5 model, fuse features through the BRA module, and output the human detection box and the preliminary drowning probability; the pose analysis sub-network calculates the movement trajectory of the skeletal key points. When it is detected that the head is below the water surface and the limbs are not paddling for 5 consecutive seconds, a secondary alarm is triggered; comprehensively considering the visual detection results, abnormal infrared body temperature, and depth data, generate the final drowning determination signal; in the pose analysis sub-network, use the LSTM model to process the sequence of skeletal key points S = [s1, s2, …, s n of consecutive frames. The output h t of the LSTM unit is calculated through the formula h t = LSTM(h t-1 , x t ) where xt is the input at the current moment, h t-1 is the hidden state at the previous moment, used to determine whether the human body movement conforms to the drowning characteristics;
[0020] Hierarchical response steps: When the system detects that a person is approaching the deep water area, it enters the early warning state and issues a voice reminder through the speaker; when the drowning alarm is triggered, it enters the alarm state, activates the sound and light alarm device, activates the red light positioning function, and sends out rescue equipment.
[0021] Furthermore, in the multi-modal target detection step, the pose analysis sub-network accurately judges whether the human body movement conforms to the drowning characteristics by continuously monitoring the movement trajectories of the bone key points, provides a more reliable basis for drowning detection, and further confirms the drowning risk.
[0022] Furthermore, in the hierarchical response step, the voice reminder in the early warning state can prevent the drowning risk in advance; the sound and light alarm, red light positioning and rescue equipment dispatch in the alarm state form an all-round and efficient rescue response mechanism. Technical effects and advantages of the present invention:
[0023] The present invention combines visual, depth and pose information through the multi-modal data acquisition module, effectively overcoming the disadvantages of traditional detection means relying on single data; the wide-angle camera and the infrared thermal imager cooperate to enable the system to accurately capture the target under various lighting conditions; the depth data provided by the ToF camera helps to judge the drowning risk; the MediaPipe model tracks the human bone key points in real time, laying a foundation for identifying abnormal postures;
[0024] In the improved YOLOv5 detection model, the GSConv module significantly reduces the calculation amount and improves the operation efficiency; the CBAM attention mechanism combined with the BRA module not only strengthens the extraction of small target features, but also significantly improves the multi-modal feature fusion effect; the pose analysis sub-network uses the LSTM time series model to deeply analyze the human body movement in continuous frames, accurately identify drowning actions, and greatly improves the accuracy and reliability of detection;
[0025] The Internet of Things linkage rescue module realizes seamless docking from detection to rescue; once drowning is detected, the red light laser device quickly marks the position, the NB-IoT module timely sends the coordinates to the administrator terminal, and the surface floating robot equipped with GPS can quickly drive to the drowning point and release rescue equipment, greatly shortening the rescue response time and providing a strong guarantee for the life safety of the drowning person;
[0026] The data preprocessing and enhancement steps adopt the guided filter and MSRCR algorithms to improve the quality of underwater images. At the same time, the infrared images are corrected to ensure the reliability of multi-source data. The multi-modal object detection process and the hierarchical response mechanism cooperate with each other, which not only improves the accuracy of drowning determination but also enables corresponding measures to be taken in a timely manner according to different situations to prevent risks in advance or carry out rescue operations efficiently. In summary, the present invention has made remarkable breakthroughs in terms of accuracy, timeliness, and rescue linkage, and can effectively reduce the harm of pool drowning accidents. Brief Description of the Drawings
[0027] Figure 1 It is the architecture diagram of the multi-modal fusion swimming pool drowning detection system based on the improved YOLOv5 in the present invention;
[0028] Figure 2 It is the flowchart of the multi-modal fusion swimming pool drowning detection method based on the improved YOLOv5 in the present invention. Detailed Embodiment
[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0030] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish one element from another.
[0031] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0032] Embodiment:
[0033] The present invention provides a multi-modal fusion swimming pool drowning detection system and method based on the improved YOLOv5, including:
[0034] A multi-modal data acquisition module for collecting data of multiple modalities, including a visual sensor, a depth sensor, and a pose capture unit; the visual sensor simultaneously deploys a wide-angle camera and an infrared thermal imager to respectively obtain the RGB image and the thermal map of the pool area; the depth sensor adopts a ToF camera, and by measuring the time of flight t of light, according to the formula Obtain the distance information d between the human body and the water surface, where c is the speed of light; the pose capture unit combines the MediaPipe model to extract the key point information of the human body skeleton in real time;
[0035] In actual deployment, the installation positions and angles of the wide-angle camera and the infrared thermal imager can be flexibly adjusted according to the size and shape of the swimming pool to ensure maximum coverage of the swimming pool area and reduce monitoring blind spots; for example, for a rectangular swimming pool, the camera can be installed in the middle of the long side of the swimming pool and tilted slightly downward so that the shooting range can cover the entire water surface of the swimming pool; at the same time, calibrate the ToF camera regularly to ensure the accuracy of the measured distance, and multiple ToF cameras can also be added to further improve the accuracy of the depth information through triangulation; for the pose capture unit, as the MediaPipe model is continuously updated, new versions can be introduced in a timely manner to improve the accuracy and speed of key point extraction of the skeleton, and its ability to capture the poses of multiple people simultaneously can also be utilized to achieve comprehensive monitoring of all the people in the swimming pool;
[0036] The improved YOLOv5 detection model replaces the original CSP module with a lightweight GSConv module in its backbone network and integrates the CBAM attention mechanism before the SPPF layer; at the same time, it is equipped with a BRA module for multi-modal feature fusion and a pose analysis sub-network added after the detection head, and this sub-network analyzes human actions through the LSTM time series model; in the CBAM attention mechanism, the channel attention module obtains and After being processed by a multi-layer perceptron (MLP), the formula is used to generate the channel attention map M c (F), where σ is the sigmoid activation function, which is used to enhance the attention to important features in the channel dimension;
[0037] In the model training stage, more swimming pool data in different scenarios can be collected, including different weather conditions, different time periods, different swimming pool layouts, etc., to enrich the training set and enable the model to better adapt to various complex situations; at the same time, transfer learning technology can be adopted to utilize the model parameters pre-trained in other related fields (such as human behavior recognition, security monitoring, etc.) to accelerate the convergence speed of this model and improve the training efficiency; for the CBAM and BRA modules, their parameter optimization strategies can be further studied, and by adjusting the weight coefficients, the model can achieve better effects in small target detection and multi-modal feature fusion; other advanced attention mechanisms or feature fusion methods can also be tried to be introduced into the model, and comparative experiments can be carried out with the existing modules to explore a more efficient model structure;
[0038] The Internet of Things (IoT) linkage rescue module has positioning and alarm functions. When drowning is detected, it triggers a red light laser device to mark the location and sends the coordinates to the administrator terminal through the NB-IoT module. It also includes a surface floating robot equipped with GPS and an electromagnetic release lifebuoy, which is used to automatically navigate to the target point to release the rescue equipment.
[0039] To improve the marking effect of the red light laser device, a high-brightness and high-directivity laser emitter can be used, combined with a flashing mode, making it more easily detectable by rescue personnel in a complex pool environment. For the NB-IoT module, its communication protocol can be optimized to increase the stability and anti-interference ability of data transmission, ensuring that the drowning coordinates can be sent to the administrator terminal in a timely and accurate manner. In terms of the surface floating robot, in addition to being equipped with GPS and an electromagnetic release lifebuoy, it can also be equipped with water quality detection sensors, life detectors and other devices, enabling it to obtain more information during the rescue process, such as the water quality situation around the drowning person and whether there are other potential dangers. At the same time, by optimizing the power system and navigation algorithm of the robot, its driving speed and obstacle avoidance ability on the water surface can be improved, ensuring that it can reach the drowning point quickly and safely.
[0040] In the improved YOLOv5 detection model, the BRA module and the CBAM attention mechanism work together to improve the efficiency of multi-modal feature fusion, enhance the extraction of small target features and the feature expression after the fusion of different modal data. When the BRA module fuses multi-modal data, taking the RGB image feature F RGB , the heatmap feature F thermal , and the depth information feature F depth as examples, through a specific fusion formula F fused = αF RGB + βF thermal + γF depth an enhanced feature map F fused is generated to improve the extraction of small target features and the feature expression after the fusion of different modal data. α, β, and γ are fusion weights obtained through training and learning. In practical applications, the fusion weights can be dynamically adjusted according to the characteristics and requirements of different pools. For example, for indoor pools, due to relatively stable lighting conditions, the weight of the infrared heatmap feature can be appropriately reduced; while for outdoor pools, in strong light or low light environments, the weights of the heatmap and depth information features can be increased. In addition, the reinforcement learning algorithm can be used to enable the model to automatically optimize the fusion weights during operation according to the detection results, continuously improving the effect of multi-modal feature fusion. It is also possible to study how to integrate more modal data (such as sound data, detecting the cries of drowning people) into the model to further enrich the feature information and improve the detection accuracy.
[0041] The red light positioning device in the Internet of Things linked rescue module can accurately mark the drowning position, providing intuitive visual guidance for rescue personnel; the autonomous navigation rescue robot quickly and accurately drives towards the drowning point and releases rescue equipment according to the received coordinates; in addition to the red light positioning device, other auxiliary positioning methods can be added, such as setting multiple Bluetooth positioning beacons around the pool. When drowning is detected, the drowning position can be further accurately located through Bluetooth signals, improving the accuracy of positioning. For the autonomous navigation rescue robot, a linkage mechanism can be established with other surrounding rescue equipment (such as lifeboats and rescue ropes on the shore). When the robot reaches the drowning point, if other auxiliary rescue equipment is needed, a request signal can be automatically sent to achieve the coordinated operation of multiple rescue equipment. At the same time, the rescue robot is equipped with a remote control function. When encountering complex situations (such as robot failures and special water area environments), the administrator can operate the robot through remote control to ensure the smooth progress of the rescue operation.
[0042] The visual sensor in the multi-modal data acquisition module, in low-light environments, uses the thermal map obtained by the infrared thermal imager and the RGB image obtained by the wide-angle camera for fusion analysis, effectively improving the accuracy of target detection and making up for the deficiencies of single visual data under complex lighting conditions; develop an adaptive image fusion algorithm to automatically adjust the fusion ratio of the RGB image and the thermal map according to the change of light intensity. In extremely dark environments, the weight of the thermal map can be increased to highlight the heat information of the human body; in the case of certain light but interference such as reflection, the fusion ratio can be appropriately adjusted to balance the advantages of the two images. Image enhancement technology can also be combined to perform secondary processing on the fused image to further improve the clarity and contrast of the image. For example, the super-resolution reconstruction technology of deep learning is used to improve the resolution of the fused image, enabling the detection model to better identify the target.
[0043] When the improved YOLOv5 detection model processes small targets, it focuses on the features of small targets through the CBAM attention mechanism, combines the BRA module to fuse multi-modal data such as RGB images, thermal maps, and depth information, significantly improving the detection accuracy of small targets; introduce a dedicated module for small target detection, such as adding an attention pyramid module (APM) to the model to further enhance the feature extraction ability for small targets. APM can highlight the features of small targets by performing attention weighting on feature maps of different scales, work together with the CBAM and BRA modules, and improve the overall small target detection performance. At the same time, use data augmentation technology to oversample small target samples in the training set, increasing the diversity of small targets and enabling the model to better learn the features of small targets. The feature differences of small targets in different modal data can also be studied, and the feature fusion strategy can be optimized accordingly to improve the detection accuracy of small targets.
[0044] Provided is a multi-modal fusion swimming pool drowning detection method based on improved YOLOv5, including the following steps:
[0045] Data preprocessing and enhancement step: Use guided filtering and MSRCR algorithm to process underwater images, enhance their contrast and solve the problem of color distortion; perform non-uniformity correction on infrared images to eliminate thermal noise interference. In guided filtering, for the input image I and the guidance image p, the output filtered image q satisfies the formula where ω i is the window centered on pixel i, and a j and b j are coefficients calculated from the pixels within the window;
[0046] In the data preprocessing stage, deep learning denoising methods can be combined, such as denoising autoencoders (DAE) based on convolutional neural networks, to jointly denoise underwater images and infrared images. DAE can learn the noise patterns in the images, more effectively remove noise, and at the same time retain the detailed information of the images. In addition, for the MSRCR algorithm, its parameter selection strategy can be optimized to automatically adjust parameters according to different image characteristics to achieve better image enhancement effects. The detection and processing function for occlusions in the images can also be added. When an object occlusion is detected, image inpainting technology is used to reconstruct the occluded part to improve the integrity of the images
[0047] Multi-modal object detection step: Input multi-source data into the improved YOLOv5 model, fuse features through the BRA module, and output the human detection box and the preliminary drowning probability; the pose analysis sub-network calculates the movement trajectories of skeletal key points. When it is detected that the head is below the water surface and the limbs do not move for 5 consecutive seconds, a secondary alarm is triggered; combining the visual detection results, abnormal infrared body temperature, and depth data, generate the final drowning determination signal; in the pose analysis sub-network, use the LSTM model to process the sequence of skeletal key points of consecutive frames S = [s1, s2, …, s n , and the output h t of the LSTM unit is calculated through the formula h t = LSTM(h t-1 , x t ), where x t is the input at the current moment, and h t-1 is the hidden state at the previous moment, so as to judge whether the human action conforms to the drowning characteristics;
[0048] In order to improve the accuracy of the posture analysis sub-network, the semantic understanding of human body movements can be increased. For example, combined with natural language processing technology, the motion information of key skeletal points can be converted into semantic descriptions, such as "swimming action" and "struggling action", and then the semantic analysis model can be used to understand and judge these descriptions to more accurately identify drowning actions. At the same time, in multimodal target detection, a multi-model fusion method can be introduced to fuse the improved YOLOv5 model with other advanced target detection models (such as Faster R-CNN, Mask R-CNN, etc.), and the detection results of multiple models can be integrated to improve the reliability of detection. It is also possible to study how to use time series analysis technology to conduct in-depth mining of multimodal data in continuous frames, discover potential signs of drowning risk, and issue early warnings;
[0049] Gradual response steps: When the system detects a person approaching deep water, it enters the early warning state and issues a voice reminder through the speaker; when the drowning alarm is triggered, it enters the alarm state, starts the sound and light alarm device, activates the red light positioning function, and dispatches life-saving equipment;
[0050] In the early warning state, in addition to voice reminders, a mobile app can send reminder messages to people near deep water to ensure the effectiveness of the reminder. In the alarm state, a linkage mechanism can be established with surrounding medical institutions. When a drowning alarm is triggered, the system automatically sends drowning incident information to nearby hospitals, including the drowning location and the general condition of the drowning person, allowing hospitals to prepare for rescue in advance and improve rescue efficiency. Intelligent rescue equipment storage points can also be set up around the swimming pool, using IoT technology to automatically unlock and access equipment, making it easier for rescue personnel to obtain the necessary rescue equipment immediately.
[0051] In the multimodal target detection step, the posture analysis subnetwork continuously monitors the motion trajectory of the skeleton key points to accurately determine whether the human body movement meets the characteristics of drowning, providing a more reliable basis for drowning detection and further confirming the drowning risk; using the clustering algorithm in machine learning, a large amount of normal swimming and drowning skeleton key point motion trajectory data is clustered and analyzed to find the cluster center and distribution characteristics of different movement patterns. In actual detection, by calculating the distance between the current skeleton key point motion trajectory and the cluster center, it is more accurate to judge whether the human body movement belongs to the normal or drowning category. At the same time, combined with the principles of biomechanics, the movement of the human body in water is modeled and analyzed to predict the movement trend of the human body under different movements and detect possible drowning risks in advance. It is also possible to increase the monitoring of human physiological indicators (such as heart rate, respiratory rate, etc.), obtain these data through wearable devices, and combine them with the posture analysis results to improve the accuracy of drowning detection.
[0052] In the hierarchical response step, voice reminders in the early warning state can prevent drowning risks in advance; the sound and light alarms, red light positioning, and the dispatch of rescue equipment in the alarm state form an all-round and efficient rescue response mechanism; in the early warning state, the frequency and content of voice reminders can be adjusted according to the distance and speed of personnel approaching the deep water area. For example, when personnel approach the deep water area quickly, increase the reminder frequency and issue more urgent warning messages. For the alarm state, virtual reality (VR) or augmented reality (AR) technology can be used to provide more intuitive on-site information for rescue personnel. Through AR glasses, rescue personnel can see the real-time position of the drowning person, the surrounding environment information, and the optimal rescue path planning. At the same time, establish a rescue effect evaluation mechanism. After the rescue is completed, evaluate the entire rescue process, summarize experience and lessons, and continuously optimize the rescue response mechanism.
[0053] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0054] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
[0055] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-modal fusion swimming pool drowning detection system based on improved YOLOv5, characterized in that: Including: The multi-modal data acquisition module is used to acquire data of multiple modalities, including a visual sensor, a depth sensor, and an attitude capture unit; the visual sensor deploys a wide-angle camera and an infrared thermal imager at the same time to respectively obtain the RGB image and the thermal map of the pool area; the depth sensor uses a ToF camera, and by measuring the time of flight t of light, according to the formula obtain the distance information d between the human body and the water surface, where c is the speed of light; The pose capture unit combines with the MediaPipe model to extract the information of human skeletal key points in real time; An improved YOLOv5 detection model, whose backbone network replaces the original CSP module with a lightweight GSConv module and integrates the CBAM attention mechanism before the SPPF layer; at the same time, it is equipped with a BRA module for multi-modal feature fusion and a pose analysis sub-network added after the detection head. This sub-network analyzes human actions through an LSTM time-series model; in the CBAM attention mechanism, the channel attention module obtains and through global average pooling and global max pooling for the input feature map F. After being processed by a multi-layer perceptron (MLP), the formula is used to generate the channel attention map M c (F), where σ is the sigmoid activation function, which is used to enhance the attention to important features in the channel dimension; The Internet of Things (IoT) linked rescue module has positioning and alarm functions. When drowning is detected, it triggers a red light laser device to mark the position and sends the coordinates to the administrator terminal through the NB-IoT module. It also includes a surface floating robot equipped with GPS and an electromagnetic release life buoy, which is used to automatically navigate to the target point to release the rescue equipment.
2. The multi-modal fusion swimming pool drowning detection system based on the improved YOLOv5 according to claim 1, wherein: In the improved YOLOv5 detection model, the BRA module works in collaboration with the CBAM attention mechanism to improve the efficiency of multi-modal feature fusion, enhance the extraction of small target features, and the feature expression after the fusion of different modal data. When the BRA module fuses multi-modal data, taking the RGB image feature F RGB , the heatmap feature F thermal , and the depth information feature F depth as examples, through a specific fusion formula F fused = αF RGB + βF thermal + γF depth to generate the enhanced feature map F fused , which improves the extraction of small target features and the feature expression after the fusion of different modal data. α, β, and γ are fusion weights and are obtained through training and learning.
3. The multi-modal fusion swimming pool drowning detection system based on the improved YOLOv5 according to claim 1, characterized in that: The red light positioning device in the IoT linked rescue module can accurately mark the drowning position, providing an intuitive visual guidance for rescue personnel. The autonomous navigation rescue robot quickly and accurately drives towards the drowning point and releases the rescue equipment according to the received coordinates.
4. The multi-modal fusion swimming pool drowning detection system based on the improved YOLOv5 according to claim 1, characterized in that: The visual sensor in the multi-modal data acquisition module, in low light environments, uses the thermal map obtained by the infrared thermal imager and the RGB image obtained by the wide-angle camera for fusion analysis, effectively improving the accuracy of target detection and making up for the deficiencies of single visual data under complex lighting conditions.
5. The multi-modal fusion swimming pool drowning detection system based on the improved YOLOv5 according to claim 1, characterized in that: When dealing with small targets, the improved YOLOv5 detection model focuses on the features of small targets through the CBAM attention mechanism, and combines the BRA module to fuse multi-modal data such as RGB images, thermal maps and depth information, significantly improving the detection accuracy of small targets.
6. A multi-modal fusion swimming pool drowning detection method based on improved YOLOv5, characterized in that, Including the following steps: Data preprocessing and enhancement steps: The guided filter and MSRCR algorithm are used to process underwater images, enhance their contrast and solve the problem of color distortion; non-uniformity correction is performed on infrared images to eliminate thermal noise interference. In the guided filter, for the input image I and the guidance image p, the output filtered image q satisfies the formula where ω i is the window centered on pixel i, and a j and b j are coefficients calculated from the pixels within the window; Steps of multi-modal object detection: Input multi-source data into the improved YOLOv5 model, fuse features through the BRA module, and output the human detection box and the preliminary drowning probability; The pose analysis sub-network calculates the movement trajectory of the skeletal key points. When it is detected that the head is below the water surface and the limbs do not move for 5 consecutive seconds, a secondary alarm is triggered; Synthesize the visual detection results, abnormal infrared body temperature, and depth data to generate the final drowning determination signal; In the pose analysis sub-network, use the LSTM model to process the sequence of skeletal key points S = [s1, s2, …, s n of consecutive frames. The output h t of the LSTM unit is calculated through the formula h t = LSTM(h t-1 , x t ), where x t is the input at the current moment, and h t-1 is the hidden state at the previous moment, so as to judge whether the human body movement conforms to the drowning characteristics; Hierarchical response step: When the system detects that a person is approaching the deep water area, it enters the early warning state and issues a voice reminder through the speaker; when the drowning alarm is triggered, it enters the alarm state, activates the sound and light alarm device, activates the red light positioning function, and sends out the rescue equipment.
7. The multi-modal fusion swimming pool drowning detection method based on the improved YOLOv5 according to claim 6, characterized in that: In the multi-modal target detection step, the pose analysis sub-network accurately judges whether the human body movement conforms to the drowning characteristics by continuously monitoring the movement trajectory of the skeletal key points, providing a more reliable basis for drowning detection and further confirming the drowning risk.
8. The multi-modal fusion swimming pool drowning detection method based on improved YOLOv5 according to claim 6, characterized in that: In the hierarchical response step, the voice reminder in the early warning state can prevent the drowning risk in advance; the sound and light alarm, red light positioning and the dispatch of rescue equipment in the alarm state form an all-round and efficient rescue response mechanism.