A method and system for drowning detection combining polarized multi-modal

By employing a polarization multimodal detection method, utilizing a polarization camera and a deep learning network, the problem of low accuracy in drowning detection under complex environments such as water surface reflection is solved, achieving high-precision drowning identification.

CN121482554BActive Publication Date: 2026-05-08XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-01-12
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing drowning detection methods have low accuracy in complex water conditions, such as water surface reflections, and are difficult to adapt to complex underwater environments, such as bathing beaches.

Method used

A polarization multimodal detection method is adopted, which acquires polarization degree map and intensity map through polarization camera, obtains data using Stokes vector method, and performs feature fusion by combining pre-trained polarization fusion network. A hierarchical convolutional network with Focus, CSP and SPPF structure is used for feature extraction and fusion, and a multi-scale prediction mechanism is used for drowning behavior detection.

Benefits of technology

It significantly improves the accuracy and robustness of drowning detection, effectively eliminates water surface reflection and environmental noise interference, preserves swimmer morphology and posture information, adapts to complex water surface environments, and improves the integrity and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482554B_ABST
    Figure CN121482554B_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods and systems for detecting drowning in combination with polarization multimodal, belong to the technical field of detecting drowning, to solve the problem of low accuracy caused by water surface reflection, complex water environment affecting existing method.The method comprises the following steps: collecting the polarization degree map and intensity map of the swimmer; input the two maps into the pre-trained polarization fusion network, and obtain the fusion image through feature extraction, fusion and reconstruction; use the simulated drowning scene video / image shot by the polarization camera to train the target detection network combined with the polarization fusion feature; input the fusion image into the trained detection network to determine whether drowning.The application eliminates water surface reflection, improves image information integrity and detection accuracy, adapts to multiple scenarios, and is suitable for intelligent drowning detection in places such as swimming pools and beaches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of drowning detection technology, specifically relating to a drowning detection method and system that combines polarization multimode. Background Technology

[0002] Drowning detection refers to the use of technology to identify drowning behavior or risks in water so that timely rescue measures can be taken. With the rapid development of computer vision and artificial intelligence technologies, deep learning-based drowning detection systems have emerged. These systems typically analyze images or video data from surveillance cameras to identify abnormal drowning behavior in real time. Because AI-based drowning detection systems offer advantages over traditional methods, such as high real-time performance, wide coverage, and reduced fatigue, they are widely used in swimming pools and other locations for intelligent drowning detection and alarm systems.

[0003] Current drowning detection methods based on video processing primarily involve detecting swimmers, tracking them, and analyzing their behavior to determine if they are drowning. Early representative machine vision drowning detection systems used infrared and (Red, Green, Blue, RGB) cameras positioned above and on the pool walls to track swimmers' movements in real time. The system analyzed the length of time swimmers remained underwater to determine the likelihood of a drowning accident and issued timely alerts. Existing methods utilize mean background modeling for underwater human detection, combining wavelet thresholding and Retinex algorithms to enhance the image and reduce underwater noise. Another method converts RGB images to HueSaturation Value (HSV) color space and combines a prior thresholding mechanism with contour detection to track swimmers. The system triggers an alarm when it detects a swimmer's contour disappearing from the water surface for a period of time. However, this method involves manually calculating a threshold each time, which is complex, and relying solely on contours leads to poor accuracy. Furthermore, complex water conditions significantly impact detection accuracy. A camera-based system is proposed for detecting drowning events in swimming pools at the earliest possible stage. The system consists of two main parts: a vision module and an event reasoning module. The vision module uses a model-based approach to represent and distinguish between the background pool area and the foreground swimmer. The event reasoning module is built on a finite state machine, which integrates multiple reasoning rules formulated based on the general movement characteristics of drowning victims. A sequential change detection algorithm is used to quickly detect potential drowning events. The system has been applied to several simulated drowning video clips with good results. A real-time vision system for outdoor swimming pools is also proposed. This system proposes a set of methods including background subtraction, denoising, data fusion, and speckle segmentation, taking into account the characteristics of the water background and crowded pool scenes. In the drowning event detection step, visual indicators of distress and drowning are incorporated into a set of foreground descriptors. A module incorporating data fusion and Hidden Markov Modeling is designed to learn the unique features of different swimming behaviors. Another drowning early prediction technique is based on a new equation-based prediction technique for early near-drowning events (NEPTUNE). The formulas and rules used by NEPTUNE are capable of detecting near drowning using video sequences of at least 1 second but no more than 5 seconds, with very few false alarms. NEPTUNE's backbone consists of a mixture of statistical image processing to merge images from the video sequence, followed by K-means clustering to extract segments from the merged images, and finally, a return to statistical image processing to derive variables for each segment, which will be used to determine drowning.Another method for extracting human body regions from videos based on drowning posture features is proposed, and a drowning detection algorithm based on convolutional autoencoders and low-level feature differences is designed accordingly. The convolutional autoencoder reconstructs and models the features of a normal swimmer. During detection, the swimmer's features are reconstructed using the trained encoder, and the error between the input and reconstructed features is used to determine whether the swimmer is drowning. A drowning detection system suitable for common swimming pools is designed based on image processing technology. The target detection module uses a mask-based convolutional neural network (Mask RCNN) algorithm, trained with labeled data to detect swimming, standing in water, and drowning behaviors. Another swimming pool drowning detection algorithm is based on human posture estimation. This method uses the OpenPose human posture estimation model to label human keypoints in the swimmer's image, constructing a set of keypoint distance vectors. Drowning behavior is determined by comparing the similarity between the swimmer's keypoint distance vectors and those of a drowning state. However, real-world scenarios are complex, and obtaining complete keypoint information is difficult, resulting in poor applicability of this method.

[0004] In summary, existing drowning detection methods have several shortcomings. For example, they require underwater cameras, which are often unsuitable for murky underwater environments or beach settings. In more complex scenarios, such as beaches, the water surface is affected by natural factors like waves, tides, and wind direction, resulting in significant fluctuations and complex water movements. Furthermore, low underwater visibility, coupled with surface reflections and refractions, presents challenges for which effective solutions are currently lacking. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a drowning detection method and system that combines polarization multimode to address the shortcomings of the prior art, thereby solving the technical problem that the accuracy of drowning detection is low due to complex water surface conditions such as water surface reflection.

[0006] The present invention adopts the following technical solution:

[0007] A drowning detection method combining polarization multimode includes the following steps:

[0008] S1. Collect polarization and intensity maps of the swimming scene, which are obtained based on the Stokes vector method;

[0009] S2. Input the polarization degree map and intensity map obtained in step S1 into the pre-trained polarization fusion network for feature fusion to obtain the fused image;

[0010] S3. Collect and label videos and images of simulated drowning scenes taken by polarization cameras, and use them as training data to train a target detection network that adapts to polarization fusion features, and obtain a trained target detection network.

[0011] The target detection network includes:

[0012] The input module is used to perform preprocessing operations on the polarization fusion image, including adaptive image scaling and adaptive anchor frame configuration, to maintain the target proportion and adapt to multi-scale detection.

[0013] The feature extraction module employs a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used to enhance the extraction of human limb and floating object contour features, the CSP structure is used to suppress water surface background interference, and the SPPF structure is used to capture multi-scale contextual information.

[0014] The feature fusion module adopts a bidirectional fusion structure that combines a feature pyramid network and a path aggregation network. The feature pyramid network is used to pass the deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features. The path aggregation network is used to pass the shallow spatial features from bottom to top and fuse them with the deep semantic features again.

[0015] The detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from the feature maps of different scales output by the feature fusion module. The categories include normal swimming and drowning state. The final detection result is output through non-maximum suppression.

[0016] S4. The fused image obtained in step S2 is fed into the target detection network trained in step S3 to detect drowning behavior and determine whether the person is in a drowning state.

[0017] Preferably, in step S1, the difference between the right-handed and left-handed circular polarization intensities is set to 0, and the simplified form of the Stokes vector method is as follows:

[0018]

[0019] in, For intensity maps, The difference in polarization intensity between the 0° and 90° directions. The difference in polarization intensity is between the 45° and 135° directions. For Stokes vectors;

[0020] The polarization degree diagram is a parameter that measures the degree of linear polarization of light waves, and its value ranges from 0 to 1.

[0021] Preferably, in step S2, the pre-training process of the polarization fusion network includes:

[0022] A publicly available polarization image dataset is used as the basic training data, and the basic training data is divided into a training set and a test set.

[0023] The polarization and intensity maps in the training set are split into non-overlapping image blocks, and image enhancement operations are performed on the image blocks.

[0024] The polarization fusion network, which includes a feature extraction module, a feature fusion module, and a feature reconstruction module, is trained using the training set after image enhancement. Under the constraint of a preset loss function, the polarization degree map and intensity map in the test set are input into the polarization fusion network during the training phase to obtain the fusion result corresponding to the test set. The pre-training is completed when the fusion result meets the preset accuracy requirements.

[0025] Preferably, the feature extraction module adopts a dual-stream encoder structure to perform targeted feature extraction on the intensity map and polarization map respectively, specifically as follows:

[0026] For the intensity map: first, a 3×3 standard convolution is used to capture shallow edge features while preserving brightness and texture details; then, a 5×5 convolution is used to progressively capture deep features.

[0027] For the polarization map: First, local polarization features are extracted through 3×3 depth-separable convolution to reduce parameter redundancy and focus on subtle polarization differences related to material properties; then, 7×7 convolution is used to progressively capture deep features.

[0028] The dual-stream encoder introduces a dense connection mechanism on both sides of the branch, which directly transmits the features of each layer to all subsequent layers through skip connections.

[0029] Preferably, the feature fusion module adopts a combined structure of multi-scale processing and polarization-guided attention, specifically including:

[0030] The features output by the feature extraction module are convolved by 3×3 and then processed at the original scale, 2x scale, and 4x scale respectively. After performing upsampling, convolution, and downsampling operations on the features at each scale, they are fused together.

[0031] The feature maps corresponding to the intensity map and the polarization degree map are input into the global average pooling layer and the global max pooling layer, respectively, to obtain the global statistical features of the two.

[0032] Introducing global statistical features of polarization degree maps into intensity maps The attention generation branch uses a learnable weight matrix. Achieve channel-level fusion of cross-modal features to generate polarization-guided channel attention weight vectors. ;

[0033] Gradient features of polarization direction changes are extracted by convolutional layers to generate a spatial weight mask for the polarization degree map. The spatial weight mask Multiplying the features by the channel weighting achieves adaptive fusion of spatial dimensions.

[0034] Preferably, the feature reconstruction module achieves feature reconstruction through three layers of convolution:

[0035] The first 3×3 convolution layer is used to integrate the correlation of cross-scale features and filter redundant information.

[0036] The second 3×3 convolution layer is used to enhance local details and global structure;

[0037] The third 1×1 convolutional layer is used to compress the channel dimension, mapping high-dimensional features to feature dimensions that match the input image, and generating the fused image.

[0038] The preset loss function Multi-scale average structural similarity loss gradient loss With strength loss The weighted sum.

[0039] Preferably, in step S3, the input module specifically comprises:

[0040] The sample distribution is expanded through diversified strategies; the original proportions of the swimmer targets are maintained, and subtle boundary features and posture information are preserved; the anchor box size is dynamically adjusted by statistically annotating the bounding boxes of the samples based on the true size distribution of the swimmer targets in the training data.

[0041] Preferably, in step S3, the feature extraction module specifically comprises:

[0042] Focus structure: The pixels of the input image are divided into four spatial locations: top left, top right, bottom left, and bottom right. After increasing the number of channels, convolution processing is performed to preserve the human body posture and floating object contour features.

[0043] CSP structure: The feature map is divided into two parts. One part performs a regular convolution operation, and the other part is fused with the first part of the feature map after passing through the residual path, thus filtering out redundant information of the water surface background.

[0044] SPPF structure: Extracts multi-scale contextual information by combining max pooling operations of different scales.

[0045] Preferably, in step S3, the training process of the target detection network uses the CIoU loss function to optimize the target localization accuracy. for:

[0046]

[0047] in, The crossover ratio loss function value, Let Euclidean distance be the center point of the target box and the predicted box. The distance is the diagonal of the target bounding box. It is a parameter that measures the aspect ratio.

[0048] Secondly, embodiments of the present invention provide a drowning detection system combining polarization multimode, comprising:

[0049] Acquisition module: used to acquire polarization and intensity maps in swimming scenarios, which are obtained based on the Stokes vector method;

[0050] Network module: Used to store the pre-trained polarization fusion network, receive the polarization degree map and intensity map output by the acquisition module, and perform feature fusion through the polarization fusion network to generate a fused image;

[0051] Training module: Used to collect and label videos and images of simulated drowning scenes captured by a polarization camera, and use the videos and images as training data to train a target detection network adapted to polarization fusion features, thereby obtaining a trained target detection network;

[0052] The target detection network includes:

[0053] The input module is used to perform preprocessing operations on the polarization fusion image, including adaptive image scaling and adaptive anchor frame configuration, to maintain the target proportion and adapt to multi-scale detection.

[0054] The feature extraction module employs a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used to enhance the extraction of human limb and floating object contour features, the CSP structure is used to suppress water surface background interference, and the SPPF structure is used to capture multi-scale contextual information.

[0055] The feature fusion module adopts a bidirectional fusion structure that combines a feature pyramid network and a path aggregation network. The feature pyramid network is used to pass the deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features. The path aggregation network is used to pass the shallow spatial features from bottom to top and fuse them with the deep semantic features again.

[0056] The detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from the feature maps of different scales output by the feature fusion module. The categories include normal swimming and drowning state. The final detection result is output through non-maximum suppression.

[0057] Drowning Module: Inputs the fused image into the target detection network to detect drowning behavior and determine whether the person is drowning.

[0058] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the above-described method for detecting drowning using combined polarization multimodal methods.

[0059] Fourthly, embodiments of the present invention provide a computer-readable storage medium including a computer program, which, when executed by a processor, implements the steps of the above-described drowning detection method incorporating polarization multimodalities.

[0060] Fifthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the above-described method for detecting drowning using combined polarization multimodal methods.

[0061] In a sixth aspect, embodiments of the present invention provide an electronic device including a computer program, which, when executed by the electronic device, implements the steps of the above-described drowning detection method combining polarization multimodalities.

[0062] Compared with the prior art, the present invention has at least the following beneficial effects:

[0063] A multimodal polarization-based drowning detection method is proposed. This method utilizes a polarization camera to acquire polarization degree maps and intensity maps, fully leveraging the complementary physical characteristics of these two image types. The intensity map preserves scene brightness and texture details, while the polarization degree map effectively mitigates water surface reflections and ripple refraction. When acquiring data using the Stokes vector method, the calculation process is simplified by ignoring the minimal circularly polarized light component, improving processing efficiency while ensuring data validity. This provides low-noise, high-purity input data for subsequent feature fusion, avoiding the environmental interference issues inherent in traditional single-modal data. Deep fusion is then performed using a pre-trained polarization fusion network, effectively eliminating… While mitigating water surface reflections and environmental noise interference, this method preserves key details such as swimmer morphology and limb posture to the greatest extent possible, significantly improving the integrity and reliability of image information. This provides high-quality feature support for the accurate identification of drowning behavior and solves the problems of image information distortion and low target-background differentiation in existing methods. The target detection network is specifically adapted to polarization fusion features and integrates the core structures of Focus, CSP, and SPPF. The Focus structure improves feature extraction efficiency, the CSP structure filters redundant information from the water surface background, and the SPPF structure enhances the capture of multi-scale contextual information. The three work together to balance feature extraction efficiency and multi-scale detection capabilities. Simultaneously, a multi-scale prediction mechanism is employed to output detection results from feature maps of different sizes, enabling accurate identification of small targets at long distances and stable capture of large targets at close range. A non-maximum suppression strategy is used to filter candidate boxes, retaining the prediction results with the highest confidence, effectively reducing redundant annotations and misjudgments, and achieving high-precision drowning detection. A pre-trained polarization fusion network collaborates with the target detection network to form an end-to-end detection solution covering the entire process from data acquisition and feature fusion to behavior detection and result output, ensuring technical integrity. Adaptable to various complex scenarios such as swimming pools and bathing beaches, it comprehensively addresses the problems of low detection accuracy and poor scene adaptability in existing methods, while also providing core support for the detailed optimization of dependent claims, forming the foundation for achieving high-precision drowning detection.

[0064] Furthermore, the Stokes vector integrates the total light intensity with the polarization intensities at 0° / 90° and 45° / 135°, providing accurate data for polarization degree map calculation. The polarization degree map quantifies the degree of linear polarization, specifically eliminating water surface reflections and suppressing their interference, thus solving the problem of target-background confusion caused by reflections in existing images. This provides high-quality, low-interference foundational data for subsequent fusion networks, ensuring that the fused image accurately reflects the swimmer's morphology and posture, and avoiding the impact of poor raw data quality on detection accuracy.

[0065] Furthermore, by dividing the training and test sets, performing image enhancement, and conducting training and testing validation, the network's generalization ability is ensured. A publicly available polarized image dataset is used to guarantee the reliability of the training data. Data is split into non-overlapping image patches and enhancement is performed to expand sample diversity and avoid network overfitting. Test set validation is introduced during training; the training effect is judged by the fusion results on the test set, avoiding the memory problem caused by relying solely on the training set. This ensures that the pre-trained fusion network can still stably output high-quality fused images on newly acquired data, solving the problems of non-standard training and poor generalization ability in existing networks, providing reliable image input for subsequent drowning detection, and ensuring the stability of the detection process.

[0066] Furthermore, a dual-stream encoder structure is adopted to specifically extract intensity and polarization maps, adapting to the different attributes of the two types of data: the intensity map first extracts basic features of shallow edges and brightness texture through 3×3 standard convolution, and then progressively captures deep structure and detail features at different scales through a dual-path parallel structure composed of 5×5 and 7×7 standard convolutions. The two features are refined by 3×3 convolution and then fused, and a dense connection mechanism is used to skip connections between shallow and deep features to avoid information attenuation. Similarly, the polarization map first extracts basic polarization information through 3×3 standard convolution, and then uses a dual-path parallel structure composed of 5×5 and 7×7 depth-separable convolutions to accurately capture polarization details and local polarization differences while reducing parameter redundancy. The two features are refined by 3×3 convolution and then fused, and the features of each layer are transmitted through dense connections. Finally, a multi-scale feature representation with strong targeting and complete information for the two types of images is formed, providing a high-quality feature foundation for subsequent feature fusion.

[0067] Furthermore, a combined structure of multi-scale processing and polarization-guided attention is employed. Multi-scale processing covers local details, intermediate structure, and global context, avoiding feature loss at a single scale. Polarization-guided attention introduces global statistical features of the polarization degree map into the intensity map attention branch, combined with a spatial weight mask, to achieve cross-modal channel and spatial fusion, highlighting the swimmer target area and suppressing background interference such as water ripples. By balancing multi-scale information and polarization characteristics, this approach addresses the problems of existing fusion methods that simply superimpose data and have low target-background differentiation. The fused image retains details while highlighting the target, providing highly discriminative features for the detection network and reducing misjudgments caused by background interference.

[0068] Furthermore, feature optimization and dimensionality matching are achieved through three convolutional layers: the first layer integrates cross-scale correlations and filters redundant information; the second layer enhances the consistency between local details and global structure; and the third layer compresses channel dimensions to match the input image, ensuring the fused image is compatible with subsequent detection networks. The preset loss function is a weighted sum of multi-scale average structural similarity loss, gradient loss, and intensity loss, constraining the fusion result from multiple dimensions to avoid structural distortion, edge blurring, or brightness distortion in the fused image. This addresses the problem of existing reconstructions easily losing key information from the source image, ensuring that the fused image accurately conveys complementary information from the intensity map and polarization map, providing high-quality input for the detection network.

[0069] Furthermore, image enhancement expands the sample distribution through diversified strategies, improving the model's robustness to complex scenes; adaptive scaling maintains the swimmer's original proportions, avoiding the loss of pose information caused by forced scaling; and the anchor boxes are dynamically adjusted based on the target's true size in the training data, adapting to swimmers of different distances and body types. Optimizing input quality through data preprocessing addresses the problem of existing input processing being simple and prone to detection errors due to target deformation or scale mismatch, providing high-quality data for subsequent feature extraction and detection, and improving the adaptability of the detection network.

[0070] Furthermore, Focus is split into four pixel spatial locations, and convolution is performed after increasing the number of channels. This reduces computational cost while preserving human limb posture and floating object contour features. CSP processes the feature map in two paths to filter redundant information from the water surface background. SPPF extracts contextual information through multi-scale max pooling, covering features of targets of different sizes. The feature fusion module adopts FPN+PAN collaboration: FPN transmits semantic information from top to bottom, and PAN transmits spatial features from bottom to top. Bidirectional fusion takes into account both deep semantics and shallow details. This solves the problems of incomplete feature extraction and information loss caused by unidirectional fusion in existing systems, ensuring that the network can stably identify swimmer targets in complex backgrounds.

[0071] Furthermore, the training employs the CIoU loss function, adding constraints on the Euclidean distance, diagonal distance, and aspect ratio between the center points of the target box and the predicted box to optimize target localization accuracy and reduce errors caused by positional offsets or shape mismatches. This addresses the existing problems of easily missing small targets and inaccurate localization, improving the accuracy and precision of drowning target recognition, reducing false positives caused by localization deviations, and ensuring rapid identification of drowning behavior.

[0072] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0073] In summary, this invention systematically solves the problem of low accuracy in drowning detection in complex water environments by fusing polarization multimodal information and using a specially optimized deep learning network architecture, resulting in significant improvements in feature enhancement, model robustness, and detection accuracy.

[0074] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0075] Figure 1 This is a flowchart of the drowning detection method combining polarization images according to the present invention;

[0076] Figure 2 This is a diagram of the polarization fusion network structure proposed in this invention;

[0077] Figure 3 This is a network structure diagram of the detection model network of the present invention;

[0078] Figure 4 This is a flowchart of the drowning detection process of the present invention;

[0079] Figure 5 A schematic diagram of a computer device provided in an embodiment of the present invention;

[0080] Figure 6 This is a block diagram of a chip provided according to an embodiment of the present invention.

[0081] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation

[0082] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0083] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0084] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0085] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.

[0086] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0087] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0088] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0089] This invention provides a drowning detection method combining polarization multimodal analysis. It simultaneously acquires polarization degree and intensity maps of a scene using a polarization camera, and constructs a specially designed polarization fusion network to fuse these two features, generating an enhanced fused image. A trained target detection network is then used to analyze the fused image, achieving automatic identification of drowning behavior. This method effectively overcomes the problems of information loss and high false detection rates in traditional visual detection under complex scenarios such as water surface reflection and wave interference. By introducing and deeply fusion polarization information, it significantly improves the integrity and representational ability of image information, thereby greatly enhancing the accuracy and robustness of drowning detection. It is applicable to various water safety monitoring scenarios such as swimming pools and beaches.

[0090] Please see Figure 1 and Figure 4 The present invention discloses a drowning detection method combining polarization multimode, comprising the following steps:

[0091] S1. Polarization maps and intensity maps of swimmers are acquired using a polarization camera and used as input data for subsequent fusion and detection model networks. The polarization camera can simultaneously acquire image information at different polarization angles, thus reducing sea surface reflection to some extent and providing a multimodal input basis for subsequent feature fusion and drowning state recognition.

[0092] The principles for obtaining polarization and intensity images are as follows:

[0093] Sunlight, initially lacking a specific direction of vibration, is chaotic. However, when these rays encounter a smooth surface, the reflected light begins to exhibit an ordered vibration pattern—this is known as polarized light. The degree of this transformation is closely related to the angle at which the light strikes the smooth surface. Therefore, whether living or non-living, anything with a smooth surface can reflect polarized light, such as water, leaves, skin, or scales. Thus, the polarization state of the reflected light can be used to determine the properties of a substance. The polarization state of light describes the direction and mode of vibration of the electric field, primarily including linear polarization, circular polarization, and elliptical polarization. The electric field of linearly polarized light vibrates in a fixed direction, circularly polarized light rotates along a circular trajectory, while elliptically polarized light vibrates in two orthogonal directions, forming an elliptical trajectory.

[0094] The commonly used method for representing polarized light is the Stokes vector method, which is represented as follows:

[0095]

[0096] in, , , and The images are polarization images at four angles of 0°, 45°, 90°, and 135°. Represents total intensity, indicating an intensity image; This represents the difference in polarization intensity between the 0° and 90° directions; This represents the difference in polarization intensity between the 45° and 135° directions; It represents the difference in intensity between right-handed and left-handed circular polarization.

[0097] Because circularly polarized light is relatively rare in nature, it is generally believed that... S 3 equals 0. Therefore, the Stokes vector is often simplified to the following form:

[0098]

[0099] The degree of linear polarization (DoLP) is a measure of the degree of linear polarization of light waves, reflecting the relative intensity of the linearly polarized components in the light wave. The value of DoLP ranges from 0 to 1, where 0 represents completely unpolarized light and 1 represents completely linearly polarized light. Its calculation formula is as follows:

[0100]

[0101] A polarization camera can be used to obtain linear polarization maps and intensity maps, i.e. and .but While containing only total intensity information, it lacks detailed information about the polarization state of light waves, making analysis of specific scenes incomplete. DoLP, on the other hand, reflects the degree of polarization of light waves, and on water surfaces or other reflective surfaces, DoLP can effectively eliminate unwanted reflections. Therefore, combining DoLP with... Fusion can provide more comprehensive information, better distinguish polarized objects from the background, and improve the accuracy of target recognition and tracking.

[0102] S2. Using the trained fusion network, the polarization map and intensity map obtained in step S1 are fused together.

[0103] Before using this method for drowning detection, the polarization fusion network must first be trained. Training data can utilize publicly available polarization image datasets, such as the Polarization Image Dataset (PIF). To improve the model's generalization ability, the dataset is divided into training and testing sets in an 8:2 ratio. During training, the polarization and intensity images are first split into smaller, non-overlapping blocks to more efficiently process local features. Subsequently, various image enhancement operations are performed on the dataset, including random flipping, rotation, illumination changes, and noise perturbation, to expand sample diversity and improve model robustness. The fusion network obtains a fusion result containing complementary information from the source images through feature extraction, feature fusion, and feature reconstruction, constrained by a loss function.

[0104] Please see Figure 2 The polarization fusion network used to fuse polarization multimodalities includes a feature extraction module, a feature fusion module, and a feature reconstruction module; specifically as follows:

[0105] This invention employs an innovative feature extraction method. The feature extraction module uses a dual-stream encoder to extract features from the source code. Targeted extraction of DoLP features: For First, 3×3 standard convolutions are used to extract basic features of shallow edges and brightness textures. Then, a dual-path parallel structure consisting of 5×5 and 7×7 standard convolutions is used to progressively capture deep structural and detail features at different scales. The two sets of features are refined by 3×3 convolutions and then fused. A dense connection mechanism is used to skip connections between shallow and deep features to avoid information attenuation. For DoLP, similarly, 3×3 standard convolutions are used to extract basic polarization information. Then, a dual-path parallel structure consisting of 5×5 and 7×7 depth-separable convolutions is used to accurately capture polarization details and local polarization differences while reducing parameter redundancy. The two sets of features are refined by 3×3 convolutions and then fused. Dense connections are used to transfer features from each layer, ultimately forming a multi-scale feature representation that is highly targeted and complete for both types of images, providing a high-quality feature foundation for subsequent feature fusion.

[0106] The feature fusion module adopts a structural design that combines multi-scale and polarization-guided attention, aiming to fully acquire and integrate multi-level feature information of the image from different spatial levels, so as to take into account both global structure and local details during the fusion process.

[0107] In the implementation phase, the features extracted by the encoder first enter the multi-scale feature extraction module. This module maps the features in parallel to 2x and 4x downsampling scales for processing. Through the multi-scale feature representation mechanism, the network can simultaneously capture local details, intermediate-scale structure, and global contextual information. Specifically, the 2x downsampling scale focuses on semantic features within an intermediate range, while the 4x downsampling scale effectively perceives the overall structural information of the image by expanding the receptive field. After upsampling, convolutional transformation, and downsampling operations to achieve dimensionality and semantic alignment, the features at each scale are fused. This multi-scale fusion strategy not only enhances the network's ability to perceive and integrate features at different scales but also significantly improves the model's robustness to scale changes and geometric deformations.

[0108] Next, the intensity feature map and polarization feature map are input into the max pooling and global average pooling layers, respectively, to obtain their global statistical features. Then, the statistics of the polarization features are introduced into the attention generation branch of the intensity feature, and an adaptively adjustable weight matrix is ​​learned through convolutional layers. W c This generates polarization-guided channel attention weight vectors, thereby achieving channel-level fusion of cross-modal features:

[0109]

[0110] Where σ() is the Sigmoid activation function.

[0111] Then, gradient features of polarization direction changes are extracted through convolutional layers to generate a spatial weight mask for the polarization degree map DoLP. The spatial weight mask Multiplying with channel-weighted features achieves adaptive spatial dimension fusion:

[0112] in, Indicates channel weighting, This indicates a space-weighted operation.

[0113] After multi-scale fusion, the features need to be refined and reconstructed through three layers of convolution to complete the final image fusion.

[0114] Specifically, the fused features are first progressively optimized through two 3×3 convolutional layers:

[0115] The first 3×3 convolution layer focuses on integrating the correlation of cross-scale features and filtering redundant information;

[0116] The second 3×3 convolution further enhances the consistency between local details and global structure;

[0117] Finally, a 1×1 convolutional layer is used to compress the channel dimension, mapping the high-dimensional features to feature dimensions that match the input image, ultimately generating reconstructed features and completing the entire image fusion process.

[0118] In the fusion network, in order to preserve as much structural and texture information as possible from the source images, this invention combines multi-scale average structural similarity loss. gradient loss and intensity loss The network is constrained as a loss function; where... By calculating the structural similarity between the source image and the fused image at multiple scales and then weighting the average, the ability of the fusion result to retain the structural information of the source image is carefully constrained. The multi-scale design can take into account the structural details at different resolutions, and the weighting mechanism can highlight the structural importance of key scales according to task requirements, thereby reducing structural distortion during the fusion process. L g It is used to constrain the fusion result to retain the edge sharpness and texture level of the source image, avoid detail blurring or edge loss caused by fusion, and enhance the visual sharpness of the image; L i The main focus is on the difference in grayscale values ​​of image pixels. By measuring the deviation in pixel intensity between the source image and the fused image, it ensures that the overall brightness, contrast and other intensity characteristics of the fused image are consistent with the source image, preventing global overbrightness, underbrightness or local intensity distortion, and maintaining the basic visual consistency of the image.

[0119] The loss function is defined as follows:

[0120]

[0121]

[0122]

[0123]

[0124] in, α , β and λ The parameter represents the weights that control the loss function. L ssim This represents a function for calculating structural similarity. , and These represent the two source images and the fused image, respectively. This represents the function for calculating the gradient.

[0125] S3. Collect and label relevant videos and images of simulated drowning taken by a polarization camera, and use them to train the detection network;

[0126] A target detection network combining polarization fusion features is employed. This network uses video and images of simulated drowning scenarios captured by a polarization camera as training data to ensure that the model training process fully reflects the polarization feature distribution in real drowning scenarios. The detection network mainly consists of an input module, a feature extraction module, a feature fusion module, and a detection output module. Its overall process is as follows:

[0127] S301, the input module is mainly responsible for preprocessing training and detection images to improve the model's generalization ability and detection accuracy; specifically, the input module includes three sub-processes: image enhancement, adaptive image scaling, and adaptive anchor box configuration.

[0128] Image augmentation aims to improve the generalization ability of a model, especially in drowning detection scenarios where swimmers are often surrounded by water reflections, ripples, and complex backgrounds, and the target size is relatively small with significant shape variations. By introducing diverse image augmentation strategies during the training phase, the sample distribution can be effectively expanded, allowing the model to better adapt to different lighting conditions, water textures, and pose changes, thereby improving the robustness of small target recognition.

[0129] Image adaptive scaling is used to maintain the original scale of the target during the input preprocessing stage, avoiding deformation or information loss caused by forced scaling, thereby preserving the swimmer's subtle boundary features and posture information, which is particularly important for distinguishing between normal swimming and abnormal drowning conditions.

[0130] The adaptive anchor box mechanism dynamically adjusts the anchor box size based on the actual size distribution of swimmer targets in the dataset during model training. By statistically analyzing the bounding boxes of labeled samples, it generates anchor box configurations that better match the actual target scale, which can significantly improve the model's detection accuracy for swimmers of different distances and body sizes, especially for capturing small targets at long distances.

[0131] S302. The feature extraction module is used to extract multi-level, multi-scale feature representations from the input image to capture the morphological and motion change features of the target. This module can adopt a hierarchical convolutional neural network structure, extracting spatial and semantic features through stacked convolutional layers, activation functions, and residual connections, providing a high-quality feature foundation for subsequent feature fusion and detection. In an optional implementation, the feature extraction module may include a focus structure (Focus), a cross-stage partial structure (CSP), and a space pyramid pooling-fast structure (SPPF), each playing a different role in the feature extraction process.

[0132] The Focus structure is primarily used to quickly capture spatial details of targets on the water surface. It increases the number of channels by a factor of four by splitting the pixels of the input image into four spatial locations: top left, top right, bottom left, and bottom right, before performing convolution. This design can fully preserve key features such as human limb posture and floating object contours while reducing computational load, providing fundamental feature support for subsequent determination of drowning actions (such as struggling, floating on one's back, etc.).

[0133] The CSP architecture optimizes the feature extraction process through cross-stage partial connections. This architecture divides the feature map into two parts: one part undergoes regular convolution, and the other part is fused with the first part via a residual path. This design filters out redundant information from the water surface background, enhances feature representation, and reduces computational complexity, thereby improving model inference speed while maintaining detection accuracy.

[0134] SPPF, a spatial pyramid pooling architecture, efficiently extracts multi-scale contextual information from images by combining max pooling operations at different scales. This design enables the network to capture feature details across scales, thereby enhancing its performance in recognizing targets of varying sizes. Compared to traditional spatial pyramid pooling architectures, SPPF further reduces computational load while maintaining multi-scale information extraction capabilities, thus improving overall inference efficiency.

[0135] Through the above multi-structure collaboration, the feature extraction module can take into account both local details and global semantic features, enabling the network to stably identify human morphological features even in complex water surface reflection and disturbance backgrounds.

[0136] S303, the feature fusion module integrates feature maps at different levels, effectively combining low-level detailed information with high-level semantic information, thereby improving the model's performance in multi-scale object detection tasks. This module establishes information interaction between feature maps of different resolutions through upsampling and downsampling operations, ensuring the model maintains consistent feature representation capabilities when detecting targets of different sizes. In one optional implementation, the feature fusion module can employ a multi-scale feature fusion structure, such as a Feature Pyramid Network (FPN) and a Path Aggregation Network (PAN) working together to enhance feature transfer and fusion.

[0137] The Feature Pyramid Network (FPN) structure is primarily used to transmit semantic information from top to bottom to address the detection challenges of targets at different scales. To address the issue of small targets being easily missed and large targets having blurred details in drowning detection, FPN achieves the fusion of deep and shallow features by constructing a feature pyramid. Deep features provide high-level semantic information such as human dynamics and floating status, while shallow features retain detailed features such as limb movements and edges. The resulting multi-scale feature map, after fusion, can simultaneously improve the detection capability for both small-scale targets (such as struggling hands) and large-scale targets (such as floating human bodies).

[0138] The PAN structure further introduces a bottom-up feature transfer path on top of FPN, thus forming a bidirectional information flow. Through this top-down and bottom-up bidirectional fusion method, PAN can more fully integrate high-level semantic features and low-level spatial features, enabling the model to maintain high detection robustness under complex lighting, water ripples, and occlusion conditions.

[0139] Through the synergistic effect of FPN and PAN, the feature fusion module can achieve comprehensive information fusion among features at different scales, effectively improving the model's detection accuracy for drowning targets.

[0140] S304. The detection output module is the final stage of the network, used to complete the classification and regression tasks of object detection. This module predicts objects based on the multi-scale feature maps output by the feature fusion module, generating object class probabilities and bounding box locations. In one implementation, the output module can employ a multi-scale prediction mechanism, outputting detection results from feature maps of different sizes to simultaneously ensure detection performance for both distant small objects and close-range large objects. To reduce duplicate detections, the output can further incorporate a non-maximum suppression strategy to filter multiple candidate boxes, retaining the object prediction result with the highest confidence.

[0141] Furthermore, the detection output module can be designed with specific loss functions to balance target classification accuracy and localization accuracy. The detection model network uses the Complete Intersection over Union (CIoU) loss function, which is an improvement on the traditional Intersection over Union (IoU) loss function. CIoU not only considers the overlap area between the predicted bounding box and the ground truth bounding box, but also incorporates distance and shape factors to better optimize the target localization task. The CIoU loss function can more accurately constrain the position and shape of the predicted bounding box, reducing accuracy loss caused by predicted bounding box position offset or shape mismatch. The complete CIoU loss function consists of the Generalized Intersection over Union Loss (GIoU_Loss) and the Distance Intersection over Union Loss (DIoU_Loss), among others.

[0142] The formula for calculating GIoU_Loss is:

[0143]

[0144] in, A , B These represent the detection box and the target box, respectively. , This represents the cross-union ratio loss function value.

[0145] DIoU_Loss The calculation formula is:

[0146]

[0147] according to DIoU_Loss The method for calculating the function, and the method used in the detection model network. CIoU_Loss The function adds an influence factor, taking into account the aspect ratio of the target box and the predicted box. CIoU_Loss Recorded as Fcou_Loss The calculation formula is as follows:

[0148]

[0149] in, d 0 represents the Euclidean distance between the center points of the target box and the predicted box. d c The distance is the diagonal of the target bounding box. v It is a parameter that measures the aspect ratio, and its definition is:

[0150]

[0151] Among them w gt and h gt These correspond to the width and height of the actual target bounding box, respectively. w p and h p These correspond to the width and height of the predicted bounding box, respectively. The CIoU_Loss function further considers factors such as overlap area, center point distance, and aspect ratio, making the calculation of the loss function more accurate.

[0152] S4. The fused image obtained in step S2 is fed into the detection model network trained in step S3 to detect drowning behavior.

[0153] Once the fusion network achieved the expected results, it was applied to fuse the intensity and polarization maps of the acquired swimmers. The fused image then served as input to a trained detection model network to determine whether a swimmer was drowning, thus achieving drowning detection. The detection model network is as follows: Figure 3 As shown.

[0154] In another embodiment of the present invention, a drowning detection system combining polarization multimodal methods is provided. This system can be used to implement the above-mentioned drowning detection method combining polarization multimodal methods. Specifically, the drowning detection system combining polarization multimodal methods includes an acquisition module, a network module, a training module, and a drowning module.

[0155] The acquisition module is used to acquire polarization maps and intensity maps.

[0156] Network module: Used to store the pre-trained polarization fusion network, receive the polarization degree map and intensity map output by the acquisition module, and perform feature fusion through the polarization fusion network to generate a fused image;

[0157] Training module: Used to collect and label videos and images of simulated drowning scenes captured by a polarization camera, and use the videos and images as training data to train a target detection network adapted to polarization fusion features, thereby obtaining a trained target detection network;

[0158] The target detection network includes:

[0159] The input module is used to perform preprocessing operations on the polarization fusion image, including adaptive image scaling and adaptive anchor frame configuration, to maintain the target proportion and adapt to multi-scale detection.

[0160] The feature extraction module employs a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used to enhance the extraction of human limb and floating object contour features, the CSP structure is used to suppress water surface background interference, and the SPPF structure is used to capture multi-scale contextual information.

[0161] The feature fusion module adopts a bidirectional fusion structure that combines a feature pyramid network and a path aggregation network. The feature pyramid network is used to pass the deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features. The path aggregation network is used to pass the shallow spatial features from bottom to top and fuse them with the deep semantic features again.

[0162] The detection output module employs a multi-scale prediction mechanism to predict the target location, category, and confidence level from the feature maps of different scales output by the feature fusion module. The categories include normal swimming and drowning. The final detection result is output through non-maximum suppression.

[0163] Drowning Module: Inputs the fused image into the target detection network to detect drowning behavior and determine whether the person is drowning.

[0164] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or function. The processor described in this embodiment can be used in conjunction with a polarization multimodal drowning detection method, including:

[0165] A polarization degree map and intensity map of a swimming scene are acquired using a polarization camera, and the polarization degree map and intensity map are obtained based on the Stokes vector method. The obtained polarization degree map and intensity map are input into a pre-trained polarization fusion network for feature fusion to obtain a fused image. Videos and images of simulated drowning scenes captured by the polarization camera are collected and labeled as training data to train a target detection network adapted to the polarization fusion features, resulting in a trained target detection network. The target detection network includes: an input module for performing preprocessing operations on the polarization fusion image, including image adaptive scaling and anchor box adaptive configuration, to maintain the target proportion and adapt to multi-scale detection; and a feature extraction module that uses a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used for... Enhanced extraction of human limb and floating object contour features: a CSP structure is used to suppress water surface background interference, and an SPPF structure is used to capture multi-scale contextual information; a feature fusion module employs a bidirectional fusion structure combining a feature pyramid network and a path aggregation network; the feature pyramid network is used to pass deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features; the path aggregation network is used to pass shallow spatial features from bottom to top and fuse them again with deep semantic features; a detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from feature maps of different scales output by the feature fusion module, with categories including normal swimming and drowning state, and outputs the final detection result through non-maximum suppression; the obtained fused image is fed into a trained target detection network to detect drowning behavior and determine whether the person is drowning.

[0166] Please see Figure 5 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the polarization-multimodal drowning detection method described in this embodiment. To avoid repetition, details are omitted here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the polarization-multimodal drowning detection system described in this embodiment. To avoid repetition, details are omitted here.

[0167] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 5This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.

[0168] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0169] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.

[0170] Furthermore, the memory 62 may include both internal storage units of the computer device 60 and external storage devices. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.

[0171] Please see Figure 6 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.

[0172] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.

[0173] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.

[0174] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0175] Bus 630 can represent one or more of several bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0176] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0177] Example 4

[0178] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0179] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.

[0180] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0181] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the drowning detection method combining polarization multimodality in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0182] A polarization degree map and intensity map of a swimming scene are acquired using a polarization camera, and the polarization degree map and intensity map are obtained based on the Stokes vector method. The obtained polarization degree map and intensity map are input into a pre-trained polarization fusion network for feature fusion to obtain a fused image. Videos and images of simulated drowning scenes captured by the polarization camera are collected and labeled as training data to train a target detection network adapted to the polarization fusion features, resulting in a trained target detection network. The target detection network includes: an input module for performing preprocessing operations on the polarization fusion image, including image adaptive scaling and anchor box adaptive configuration, to maintain the target proportion and adapt to multi-scale detection; and a feature extraction module that uses a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used for... Enhanced extraction of human limb and floating object contour features: a CSP structure is used to suppress water surface background interference, and an SPPF structure is used to capture multi-scale contextual information; a feature fusion module employs a bidirectional fusion structure combining a feature pyramid network and a path aggregation network; the feature pyramid network is used to pass deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features; the path aggregation network is used to pass shallow spatial features from bottom to top and fuse them again with deep semantic features; a detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from feature maps of different scales output by the feature fusion module, with categories including normal swimming and drowning state, and outputs the final detection result through non-maximum suppression; the obtained fused image is fed into a trained target detection network to detect drowning behavior and determine whether the person is drowning.

[0183] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0184] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0185] Simulation Experiment

[0186] I. Experimental Design

[0187] Dataset: The dataset consists of the publicly available polarized image dataset PIF (containing 1000 sets of polarized data of water surface scenes) and a self-made simulated drowning dataset (polarized camera footage of swimming pool and bathing beach scenes, including 1000 videos / 5000 images of normal swimming and 500 videos / 2500 images of simulated drowning, labeled with drowning status). The dataset is divided into a training set (1200 videos / 6000 images) and a test set (300 videos / 1500 images) in an 8:2 ratio.

[0188] Comparison methods: Existing Poseidon drowning alarm system, OpenPose-based human posture drowning detection method, and YOLOv5 detection method based on RGB images.

[0189] Evaluation metrics: detection accuracy, false alarm rate, false negative rate, and processing speed (FPS, frames per second).

[0190] Experimental environment: CPU Intel i7-12700K, GPU NVIDIA RTX 3090, 32GB memory, operating system Windows 10.

[0191] II. Experimental Results

[0192]

[0193] III. Results Analysis

[0194] Accuracy advantages: The accuracy of this invention reaches 95.6%, which is 17.4 percentage points higher than the Poseidon system (due to the elimination of reflection by polarization multimodal and the improvement of feature quality by fusion network); it is 14.1 percentage points higher than the OpenPose-based method (it does not rely on complete key points and is suitable for underwater low visibility scenarios); and it is 12.3 percentage points higher than RGB-YOLOv5 (polarization data has stronger resistance to water surface ripples and reflection interference).

[0195] Reliability advantages: The false alarm rate of this invention is 3.5% and the false alarm rate is 2.1%, which are far lower than the comparison methods. The false alarm rate is 12.1 percentage points lower than that of Poseidon because polarization-guided fusion reduces background interference; the false alarm rate is 8.4 percentage points lower than that of OpenPose because multi-scale prediction covers small targets.

[0196] Real-time advantages: Processing speed of 28FPS, meeting the requirements of real-time detection (≥25FPS), which is 10FPS higher than Pose and 16FPS higher than OpenPose. Due to the Focus structure and depth-separable convolution, the amount of computation is reduced.

[0197] In summary, this invention presents a drowning detection method and system combining polarization multimodal imaging. It utilizes a multi-angle polarization camera to simultaneously acquire polarization degree images and intensity images, and through fusion processing, deeply integrates the information from both types of images. While eliminating water surface reflections and environmental noise, it effectively enhances the overall information content of the image, thereby significantly reducing the impact of noise interference on detection. The polarization fusion drowning detection method proposed in this invention enhances detection performance through innovative image fusion technology. Specifically, the designed fusion network can deeply fuse the basic visual information of the intensity image with the scene detail information of the polarization degree image, significantly enhancing the image's representational ability and providing richer and more reliable feature support for drowning target recognition. Simultaneously, this invention employs an optimized detection model network, achieving high-precision target localization and status judgment while ensuring real-time processing speed. Combined with specially designed drowning detection rules, it can quickly and accurately identify drowning victims, improving the overall safety and reliability of detection.

[0198] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0199] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0200] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0201] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0202] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0203] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0204] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random-access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0205] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0206] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0207] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0208] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A drowning detection method combining polarization multimode, characterized in that, Includes the following steps: S1. Collect polarization and intensity maps of the swimming scene, which are obtained based on the Stokes vector method; S2. Input the polarization degree map and intensity map obtained in step S1 into the pre-trained polarization fusion network for feature fusion to obtain the fused image; The polarization fusion network includes a feature extraction module, a feature fusion module, and a feature reconstruction module; The feature extraction module employs a dual-stream encoder structure to perform targeted feature extraction on the intensity map and polarization map, respectively, as follows: For the intensity map: First, shallow edge features are captured using a 3×3 standard convolution to preserve brightness and texture details; then, deep features are captured using two parallel branches. One branch includes a 5×5 convolution and a 3×3 standard convolution connected in sequence, and the other branch includes a 7×7 standard convolution and a 3×3 standard convolution connected in sequence. The output features of the two branches are concatenated and then concatenated with the original intensity map to obtain the feature map corresponding to the intensity map. For the polarization degree map: first extract the initial polarization features through 3×3 standard convolution; Then, deep feature capture is performed through two parallel branches. One branch includes a 5×5 depthwise separable convolution and a 3×3 standard convolution connected in sequence, and the other branch includes a 7×7 depthwise separable convolution and a 3×3 standard convolution connected in sequence. The output features of the two branches are concatenated and then concatenated with the original polarization map to obtain the feature map corresponding to the polarization map. The dual-stream encoder introduces a dense connection mechanism on both sides of the branch, which directly transmits the features of each layer to all subsequent layers through skip connections; The feature fusion module employs a combined structure of multi-scale processing and polarization-guided attention, specifically including: The features output by the feature extraction module are processed at 2x and 4x scales respectively. After performing upsampling, convolution, and downsampling operations on the features at each scale, they are fused together. The intensity feature map and polarization feature map after multi-scale processing and fusion are input into the global max pooling layer and the global average pooling layer, respectively, to obtain the global statistical features of the two. The global statistical features of the polarization degree map are introduced into the attention generation branch of the intensity map, and an adaptively adjustable weight matrix is ​​learned through convolutional layers. Generate polarization-guided channel attention weight vectors This enables channel-level fusion of cross-modal features; Gradient features of polarization direction changes are extracted by convolutional layers to generate a polarization degree map. Spatial weight mask The spatial weight mask Multiplying with the channel-weighted features achieves adaptive fusion of spatial dimensions; The feature reconstruction module achieves feature reconstruction through three layers of convolution: The first 3×3 convolution layer is used to integrate the correlation of cross-scale features and filter redundant information. The second 3×3 convolution layer is used to enhance local details and global structure; The third 1×1 convolutional layer is used to compress the channel dimension, mapping high-dimensional features to feature dimensions that match the input image, and generating the fused image. Preset loss function Multi-scale average structural similarity loss gradient loss With strength loss The weighted sum; S3. Collect and label videos and images of simulated drowning scenes taken by polarization cameras, and use them as training data to train a target detection network that adapts to polarization fusion features, and obtain a trained target detection network. The target detection network includes: The input module is used to perform preprocessing operations on the polarization fusion image, including adaptive image scaling and adaptive anchor frame configuration, to maintain the target proportion and adapt to multi-scale detection. The feature extraction module employs a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used to enhance the extraction of human limb and floating object contour features, the CSP structure is used to suppress water surface background interference, and the SPPF structure is used to capture multi-scale contextual information. The feature fusion module adopts a bidirectional fusion structure that combines a feature pyramid network and a path aggregation network. The feature pyramid network is used to pass the deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features. The path aggregation network is used to pass the shallow spatial features from bottom to top and fuse them with the deep semantic features again. The detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from the feature maps of different scales output by the feature fusion module. The categories include normal swimming and drowning state. The final detection result is output through non-maximum suppression. S4. The fused image obtained in step S2 is fed into the target detection network trained in step S3 to detect drowning behavior and determine whether the person is in a drowning state.

2. The drowning detection method combining polarization multimode according to claim 1, characterized in that, In step S1, assuming the difference between right-handed and left-handed circular polarization intensities is 0, the simplified form of the Stokes vector method is as follows: in, For intensity maps, The difference in polarization intensity between the 0° and 90° directions. The difference in polarization intensity is between the 45° and 135° directions. For Stokes vectors; The polarization degree diagram is a parameter that measures the degree of linear polarization of light waves, and its value ranges from 0 to 1.

3. The drowning detection method combining polarization multimode according to claim 1, characterized in that, In step S2, the pre-training process of the polarization fusion network includes: A publicly available polarization image dataset is used as the basic training data, and the basic training data is divided into a training set and a test set. The polarization and intensity maps in the training set are split into non-overlapping image blocks, and image enhancement operations are performed on the image blocks. The polarization fusion network, which includes a feature extraction module, a feature fusion module, and a feature reconstruction module, is trained using the training set after image enhancement. Under the constraint of a preset loss function, the polarization degree map and intensity map in the test set are input into the polarization fusion network during the training phase to obtain the fusion result corresponding to the test set. The pre-training is completed when the fusion result meets the preset accuracy requirements.

4. The drowning detection method combining polarization multimode according to claim 1, characterized in that, In step S3, the input module specifically comprises: The sample distribution is expanded through diversified strategies; the original proportions of the swimmer targets are maintained, and subtle boundary features and posture information are preserved; the anchor box size is dynamically adjusted by statistically annotating the bounding boxes of the samples based on the true size distribution of the swimmer targets in the training data.

5. The drowning detection method combining polarization multimode according to claim 1, characterized in that, In step S3, the feature extraction module specifically comprises: Focus structure: The pixels of the input image are divided into four spatial locations: top left, top right, bottom left, and bottom right. After increasing the number of channels, convolution processing is performed to preserve the human body posture and floating object contour features. CSP structure: The feature map is divided into two parts. One part performs a regular convolution operation, and the other part is fused with the first part of the feature map after passing through the residual path, thus filtering out redundant information of the water surface background. SPPF structure: Extracts multi-scale contextual information by combining max pooling operations of different scales.

6. The drowning detection method combining polarization multimode according to claim 1, characterized in that, In step S3, the training process of the target detection network uses the CIoU loss function to optimize the target localization accuracy. for: in, The crossover ratio loss function value, Let Euclidean distance be the center point of the target box and the predicted box. The distance is the diagonal of the target bounding box. It is a parameter that measures the aspect ratio.

7. A drowning detection system combining polarization multimode, characterized in that, include: Acquisition module: used to acquire polarization and intensity maps in swimming scenarios, which are obtained based on the Stokes vector method; Network module: Used to store the pre-trained polarization fusion network, receive the polarization degree map and intensity map output by the acquisition module, and perform feature fusion through the polarization fusion network to generate a fused image; The polarization fusion network includes a feature extraction module, a feature fusion module, and a feature reconstruction module; The feature extraction module employs a dual-stream encoder structure to perform targeted feature extraction on the intensity map and polarization map, respectively, as follows: For the intensity map: First, shallow edge features are captured using a 3×3 standard convolution to preserve brightness and texture details; then, deep features are captured using two parallel branches. One branch includes a 5×5 convolution and a 3×3 standard convolution connected in sequence, and the other branch includes a 7×7 standard convolution and a 3×3 standard convolution connected in sequence. The output features of the two branches are concatenated and then concatenated with the original intensity map to obtain the feature map corresponding to the intensity map. For the polarization degree map: first extract the initial polarization features through 3×3 standard convolution; Then, deep feature capture is performed through two parallel branches. One branch includes a 5×5 depthwise separable convolution and a 3×3 standard convolution connected in sequence, and the other branch includes a 7×7 depthwise separable convolution and a 3×3 standard convolution connected in sequence. The output features of the two branches are concatenated and then concatenated with the original polarization map to obtain the feature map corresponding to the polarization map. The dual-stream encoder introduces a dense connection mechanism on both sides of the branch, which directly transmits the features of each layer to all subsequent layers through skip connections; The feature fusion module employs a combined structure of multi-scale processing and polarization-guided attention, specifically including: The features output by the feature extraction module are processed at 2x and 4x scales respectively. After performing upsampling, convolution, and downsampling operations on the features at each scale, they are fused together. The intensity map feature map and polarization map feature map after multi-scale processing and fusion are input into the global average pooling layer and the global max pooling layer respectively to obtain the global statistical features of the two. The global statistical features of the polarization degree map are introduced into the attention generation branch of the intensity map, and a learnable weight matrix is ​​used. Achieve channel-level fusion of cross-modal features to generate polarization-guided channel attention weight vectors. ; Gradient features of polarization direction changes are extracted by convolutional layers to generate a polarization degree map. Spatial weight mask The spatial weight mask Multiplying with the channel-weighted features achieves adaptive fusion of spatial dimensions; The feature reconstruction module achieves feature reconstruction through three layers of convolution: The first 3×3 convolution layer is used to integrate the correlation of cross-scale features and filter redundant information. The second 3×3 convolution layer is used to enhance local details and global structure; The third 1×1 convolutional layer is used to compress the channel dimension, mapping high-dimensional features to feature dimensions that match the input image, and generating the fused image. Preset loss function Multi-scale average structural similarity loss gradient loss With strength loss The weighted sum; Training module: Used to collect and label videos and images of simulated drowning scenes captured by a polarization camera, and use the videos and images as training data to train a target detection network adapted to polarization fusion features, thereby obtaining a trained target detection network; The target detection network includes: The input module is used to perform preprocessing operations on the polarization fusion image, including adaptive image scaling and adaptive anchor frame configuration, to maintain the target proportion and adapt to multi-scale detection. The feature extraction module employs a hierarchical convolutional network containing Focus, CSP, and SPPF structures to extract multi-level and multi-scale visual features from the preprocessed polarization fusion image. The Focus structure is used to enhance the extraction of human limb and floating object contour features, the CSP structure is used to suppress water surface background interference, and the SPPF structure is used to capture multi-scale contextual information. The feature fusion module adopts a bidirectional fusion structure that combines a feature pyramid network and a path aggregation network. The feature pyramid network is used to pass the deep semantic features output by the feature extraction module from top to bottom and fuse them with shallow detail features. The path aggregation network is used to pass the shallow spatial features from bottom to top and fuse them with the deep semantic features again. The detection output module uses a multi-scale prediction mechanism to predict the target location, category, and confidence level from the feature maps of different scales output by the feature fusion module. The categories include normal swimming and drowning state. The final detection result is output through non-maximum suppression. Drowning Module: Inputs the fused image into the target detection network to detect drowning behavior and determine whether the person is drowning.

Citation Information

Patent Citations

  • Cross-domain binocular stereo matching method based on multi-scale information dynamic fusion and feature deviation correction

    CN120635507A

  • Swimming pool drowning detection method and system based on improved YOLO11 network

    CN121121842A