Drowning alarm method and system based on fusion of multiple strategies

By installing infrared and visible light cameras on the swimming pool and performing multi-condition data fusion analysis, the problem of existing drowning monitoring systems performing poorly under different lighting conditions is solved, achieving higher drowning behavior detection accuracy and lower false alarm rates.

CN120183032APending Publication Date: 2025-06-20巨岩智能科技(杭州)有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126921.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing drowning monitoring system performs poorly under different lighting conditions, has a low recall rate, and lacks an effective multimodal data fusion mechanism, resulting in more false alarms and missed alarms.

Method used

A drowning alarm method based on the fusion of multiple strategies is adopted. Image data is obtained through infrared cameras and visible light cameras installed above the swimming pool, human object detection and tracking, three-dimensional position information is calculated, infrared temperature difference information is analyzed, and these results are subjected to multi-condition fusion analysis to obtain the final drowning judgment result.

Benefits of technology

It improves the accuracy and robustness of drowning behavior detection, reduces false alarm phenomenon, and provides more reliable drowning prevention and rescue support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183032A_ABST
    Figure CN120183032A_ABST
Patent Text Reader

Abstract

The invention discloses a drowning alarm method and system based on fusion of multiple strategies. The method comprises the following steps: acquiring image data shot by an infrared camera and a visible light camera mounted above a swimming pool to obtain initial data; carrying out human body target detection and tracking on the initial data to obtain detection results under different visual angles; matching detection results under different visual angles, and calculating three-dimensional position information; mapping the human body target in the matched detection result to a corresponding infrared visual angle, reading and processing temperature information, analyzing the temperature information, and speculating the current state of the human body target; inputting the initial data into a behavior speculation model to analyze the floating or sinking behavior of the target; fusing and analyzing all results; and outputting a final drowning judgment result. According to the method, the false alarm phenomenon can be reduced through accurate detection and behavior recognition of the human body target in the swimming pool, and more reliable technical support is provided for drowning prevention and rescue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically to a drowning alarm method and system based on the fusion of multiple strategies. Background Art

[0002] With the improvement of society's awareness of water activity safety, drowning monitoring technology has become an indispensable part of swimming pool management. Traditional drowning detection methods mainly rely on the manual observation of lifeguards. This method is not only affected by the individual abilities and states of lifeguards, but also difficult to achieve all-weather and all-round real-time monitoring. Especially at night or in poor light conditions, the effect of manual observation is more limited. In addition, when there are a large number of people in the swimming pool, lifeguards may not be able to detect potential drowning incidents in time, thus increasing the difficulty and risk of rescue.

[0003] In recent years, the development of computer vision technology and deep learning algorithms has provided new solutions for intelligent monitoring. Especially the fusion system based on infrared and visible light cameras can significantly improve the accuracy and robustness of drowning behavior detection through the complementary of multi-modal information. This system can not only overcome the limitations of single-modal data in specific environments, but also use the information provided by different types of sensors to enhance the adaptability of the system. For example, in low light conditions, infrared imaging technology can capture temperature differences to help identify potential drowning victims; while in bright environments, visible light images can provide rich texture details to help more accurately locate and track targets.

[0004] In terms of object detection, the YOLO series of algorithms are widely used in video surveillance and human detection tasks due to their high efficiency and real-time performance. As the latest version of this series, YOLOv5 has powerful feature extraction capabilities and can quickly identify and locate multiple targets in complex environments. At the same time, in order to ensure continuous tracking of multiple moving objects, researchers have proposed a multi-object tracking method based on the Hungarian algorithm. This method matches the targets in the current frame with the targets in the previous frame by minimizing the cost matrix, achieving a stable and efficient target tracking effect. However, existing drowning monitoring systems still have some deficiencies. First, most systems rely only on a single type of camera for monitoring, which limits their performance in different light conditions. Second, for the two typical drowning behaviors of floating and sinking, the recall rate of existing technologies is relatively low, especially in scenarios with large dynamic changes, false alarms or missed detections are likely to occur. Finally, due to the lack of an effective multi-modal data fusion mechanism, many systems fail to fully utilize the correlation between infrared and visible light images, resulting in room for improvement in overall performance.

[0005] Therefore, it is necessary to design a new method to achieve accurate detection and behavior recognition of human targets in the swimming pool, and reduce the occurrence of false alarms through the optimized multi-condition fusion technology, providing more reliable technical support for drowning prevention and rescue. Summary of the Invention

[0006] The object of the present invention is to overcome the defects of the prior art and provide a drowning alarm method and system based on the fusion of multiple strategies.

[0007] To achieve the above object, the present invention adopts the following technical solutions: A drowning alarm method based on the fusion of multiple strategies, including:

[0008] Obtain the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data;

[0009] Perform human target detection and tracking on the initial data to obtain detection results from different perspectives;

[0010] Match the detection results from different perspectives and calculate the three-dimensional position information of the human target in the matched detection results to obtain a calculation result;

[0011] Map the human target in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information;

[0012] Input the initial data into the behavior speculation model to analyze the floating or sinking behavior of the target to obtain a speculation result;

[0013] Fusion analyze the calculation result, the infrared temperature difference information, and the speculation result to obtain a final drowning determination result;

[0014] Output the final drowning determination result.

[0015] A further technical solution thereof is: The total number of the infrared camera and the visible light camera is 16. One infrared camera and one visible light camera are combined to form a group of cameras, and each group of cameras observes the same area.

[0016] A further technical solution thereof is: The obtaining the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data includes:

[0017] Adopt multi-thread technology to parallel collect the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data.

[0018] A further technical solution thereof is: The performing human target detection and tracking on the initial data to obtain detection results from different perspectives includes:

[0019] Analyze the initial data using the pre-trained YOLOv5 model to obtain the bounding boxes and corresponding confidence levels of each human target, forming a detection result, where each human target includes human targets from different perspectives;

[0020] Based on the detection result, use the Hungarian algorithm to minimize the cost matrix to match the human targets in the current frame with those in the previous frame, and update the trajectory numbers of the human targets.

[0021] Its further technical solution is: matching the detection results from different perspectives and calculating the three-dimensional position information of the human targets in the matched detection results to obtain a calculation result, including:

[0022] Match the detection results from different perspectives through the epipolar matching algorithm;

[0023] Based on the triangulation method, calculate the three-dimensional position information of the human targets for the matched detection results, and calculate the distance between the human head and the water surface according to the three-dimensional position information to obtain a calculation result.

[0024] Its further technical solution is: mapping the human targets in the matched detection results to the corresponding infrared perspectives, reading and processing the temperature information, and analyzing the temperature information to obtain infrared temperature difference information, including:

[0025] Map the human targets in the matched detection results to the corresponding infrared perspectives, read and process the temperature information, and apply the Gaussian difference filtering algorithm to highlight the temperature change regions and infer the current state of the human targets to obtain infrared temperature difference information.

[0026] Its further technical solution is: the behavior inference model is obtained by training with the labeled drowning action video data containing floating and sinking states as the sample set, using the SlowOnly model as the basic architecture and combining it with a deep learning network.

[0027] Its further technical solution is: the behavior inference model is obtained by training with the labeled drowning action video data containing floating and sinking states as the sample set, using the SlowOnly model as the basic architecture and combining it with a deep learning network, including:

[0028] Obtain the drowning action video data containing floating and sinking states and perform type annotation to obtain a sample set;

[0029] Construct a two-stream network using the SlowOnly model as the basic architecture and combining it with a deep learning network to obtain an initial model;

[0030] The sample set is processed using a 3D convolutional neural network and a convolutional long short-term memory network to handle the spatio-temporal dynamics in video frames, and the processed RGB features and infrared features are combined together by means of early fusion or late fusion to obtain a feature result;

[0031] The initial model is trained using the feature result to obtain a behavior inference model.

[0032] A further technical solution thereof is: after fusing and analyzing the calculation result, the infrared temperature difference information, and the inference result to obtain a final drowning determination result, it includes:

[0033] When the final drowning determination result indicates the existence of a drowning behavior, based on the duration of the drowning behavior and the frame ratio, it is determined whether to generate a drowning warning message, and the generated drowning warning message is sent to the terminal for display.

[0034] The present invention also provides a drowning alarm system based on the fusion of multiple strategies, including:

[0035] An acquisition unit, configured to acquire the image data captured by an infrared camera and a visible light camera installed above the swimming pool to obtain initial data;

[0036] A detection and tracking unit, configured to perform human target detection and tracking on the initial data to obtain detection results from different perspectives;

[0037] A calculation unit, configured to match the detection results from different perspectives and calculate the three-dimensional position information of the human target in the matched detection results to obtain a calculation result;

[0038] A temperature analysis unit, configured to map the human target in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information;

[0039] A speculation unit, configured to input the initial data into the behavior speculation model to analyze the floating or sinking behavior of the target to obtain a speculation result;

[0040] A fusion analysis unit, configured to fuse and analyze the calculation result, the infrared temperature difference information, and the speculation result to obtain a final drowning determination result;

[0041] An output unit, configured to output the final drowning determination result.

[0042] The beneficial effects of the present invention compared with the prior art are as follows: Through the infrared and visible light cameras installed above the swimming pool, the system acquires initial image data, conducts human target detection and tracking on this data, and obtains detection results from different perspectives. Then, the system calculates the three-dimensional position information of the human target by matching the detection results from different perspectives, maps it to the infrared perspective, acquires and analyzes the temperature information, and obtains the infrared temperature difference information. After inputting the initial data into the behavior speculation model, the floating or sinking behavior of the target is analyzed to obtain the speculation result. Then, by combining the calculation result, the infrared temperature difference information, and the speculation result, multi-condition fusion analysis is carried out to reduce the false alarm phenomenon, and finally the drowning determination result is obtained. This technology provides more reliable support for drowning prevention and rescue.

[0043] The following further describes the present invention with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart of the drowning alarm method based on multi-strategy fusion provided by the embodiment of the present invention;

[0046] Figure 2 It is a schematic sub-flowchart of the drowning alarm method based on multi-strategy fusion provided by the embodiment of the present invention;

[0047] Figure 3 It is a schematic sub-flowchart of the drowning alarm method based on multi-strategy fusion provided by the embodiment of the present invention;

[0048] Figure 4 It is a schematic sub-flowchart of the drowning alarm method based on multi-strategy fusion provided by the embodiment of the present invention;

[0049] Figure 5 It is a schematic block diagram of the drowning alarm system based on multi-strategy fusion provided by the embodiment of the present invention;

[0050] Figure 6 It is a schematic block diagram of the computer device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0053] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0054] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0055] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a drowning alarm method based on the fusion of multiple strategies provided by an embodiment of the present invention. This method is applied to a server, which interacts with terminals, infrared cameras, and visible light cameras. By combining infrared and visible light cameras, image data in the pool is obtained, and the YOLOv5 model is used for human target detection and tracking to accurately identify targets from different perspectives and calculate their three-dimensional positions. Then, the spatial information of the target is calculated through the epipolar matching algorithm, and the human state is speculated by combining temperature analysis to provide a preliminary basis for judging drowning. The floating or sinking behavior of the target is analyzed through a behavior speculation model based on deep learning to further enhance the accuracy of drowning detection. The detection results, state speculation, and behavior analysis are fused with multiple strategies to effectively reduce false alarms and provide more accurate drowning alarms. Finally, when it is determined that drowning has occurred, the system generates a warning message according to the duration to ensure timely alarm and assist in rescue.

[0056] Specifically, first, the system covers the pool area by arranging 16 cameras (8 infrared and 8 visible light cameras), and uses multi-threaded technology to synchronously collect image data of the two modalities. Then, the YOLOv5 algorithm is used for image feature extraction and human target detection, and the Hungarian algorithm is used to achieve multi-target tracking. Subsequently, the epipolar matching algorithm and triangulation method are used to accurately calculate the three-dimensional position information of the human target, and combined with the horizontal data of the pool, the distance between the human head and the water surface is accurately measured. To improve the monitoring accuracy, by combining visible light and infrared data, the Gaussian difference filtering algorithm is used to extract the temperature difference information to evaluate the current state of the target. Finally, based on a large amount of labeled drowning action video data, combined with the SlowOnly model and deep learning network, multi-modal fusion analysis is carried out to identify two drowning behaviors of floating and sinking. The experimental results show that this method can efficiently and accurately detect drowning behaviors and generate alarms in real time, providing strong technical support for drowning prevention and rescue.

[0057] Figure 1 is a schematic flowchart of the drowning alarm method based on the fusion of multiple strategies provided by the embodiments of the present invention. As Figure 1 shown, the method includes the following steps S110 to S170.

[0058] S110. Obtain the image data captured by the infrared camera and visible light camera installed above the pool to obtain the initial data.

[0059] In this embodiment, the initial data includes the image data captured by different groups of infrared cameras and visible light cameras, that is, infrared images and visible light images from different perspectives.

[0060] Specifically, the total number of the infrared cameras and the visible light cameras is 16. One infrared camera and one visible light camera form a group of cameras, and each group of cameras observes the same area.

[0061] To ensure full coverage of the standard pool area (50 m × 25 m), 16 cameras are arranged around the pool, including 8 infrared cameras and 8 visible light cameras. Each pair of infrared cameras and visible light cameras observes the same area, thus realizing the monitoring of the entire pool area.

[0062] Specifically, the Unity 3D software is used to simulate the installation perspective of the cameras. According to the installation distance and layout requirements of the actual venue, the best camera installation method is designed to ensure the coverage effect and monitoring accuracy.

[0063] After the camera installation is completed, the calibration work is carried out next. The purpose of calibration is to obtain the internal and external parameters of the infrared camera and visible light camera. The specific steps are:

[0064] Use the left and right cameras to take multiple images containing the calibration board respectively.

[0065] Keep the camera position fixed and take images at different angles by changing the position and angle of the calibration board.

[0066] For each calibration image, apply the corner detection algorithm to extract the corners on the calibration board.

[0067] Each corner on the calibration board has known actual coordinates because the size of the calibration board is known.

[0068] By matching the image coordinates captured by the camera with the known actual coordinates on the calibration board, calculate the internal parameters (such as focal length, principal point, etc.) and external parameters (such as rotation matrix, translation vector, etc.) of the camera.

[0069] In this embodiment, multi-threading technology is adopted to parallelly collect the image data captured by the infrared camera and visible light camera installed above the pool to obtain the initial data.

[0070] To achieve the parallel collection of infrared camera and visible light camera data and ensure the temporal consistency during the acquisition of different modality images, multi-threading technology is adopted. During this process, two independent thread pools are created for the infrared camera and visible light camera respectively.

[0071] Thread pool 1 (infrared camera data acquisition): The size of this thread pool is dynamically determined according to the number of infrared cameras. Each thread is responsible for periodically obtaining images from the infrared camera and saving the images as infrared perspective images. To ensure that the image acquisition is not affected by thread blocking, the threads in the thread pool work in parallel to improve the data acquisition efficiency.

[0072] Thread pool 2 (visible light camera data acquisition): The size of thread pool 2 is determined by the number of visible light cameras. Each thread is responsible for obtaining image data from the visible light camera and saving the images as visible light perspective images. In addition, high-resolution visible light data needs to be obtained additionally to ensure that this data can provide accurate visual information in subsequent analysis. To improve the speed and quality of data acquisition, the threads in thread pool 2 also work in parallel.

[0073] In this way, the two thread pools are respectively responsible for the data acquisition of different cameras, ensuring that the data of the infrared and visible light cameras can be obtained efficiently and synchronously, thereby providing high-quality raw data for subsequent image analysis and processing.

[0074] Before the camera data acquisition starts, it is necessary to initialize the camera interface and set the relevant parameters of the camera. This process includes:

[0075] Initialize the infrared camera and visible light camera to ensure that the hardware interfaces and communication protocols of each camera can work properly, and configure the camera parameters (such as exposure time, resolution, image format, etc.) according to actual needs.

[0076] After the camera initialization is completed, start the work of two thread pools, which are responsible for starting the image acquisition of the infrared camera and visible light camera respectively. The threads in the two thread pools will work simultaneously, thus realizing parallel data acquisition.

[0077] During the acquisition process of each frame of image, adopt a timestamp recording mechanism to mark the acquisition time of the image. The timestamp can be accurate to the millisecond level and is used for subsequent image synchronization processing.

[0078] In order to ensure the timing consistency between the infrared image and the visible light image, use a third thread pool (thread pool 3) to calibrate the timestamps of the acquired images. Thread pool 3 will judge whether the infrared image and the visible light image are regarded as the same frame according to the difference in timestamps. If the timestamp difference between the two images is within 200 milliseconds, it is considered that the two images belong to the same frame. Only the images that meet this timing requirement will be considered synchronized and then proceed with subsequent analysis and processing.

[0079] In this way, it is ensured that the images acquired by the infrared camera and the visible light camera are highly consistent in timing, avoiding data distortion or errors caused by excessive acquisition time differences.

[0080] This method improves the parallelism and efficiency of data acquisition by introducing multi-thread technology to manage the image acquisition tasks of the infrared camera and the visible light camera respectively using thread pools. At the same time, by using timestamps to mark and calibrate the acquisition time of each frame of image, the timing consistency of different modality images is ensured. This provides high-quality and highly synchronized data support for subsequent image registration, analysis and processing.

[0081] S120. Perform human target detection and tracking on the initial data to obtain detection results from different perspectives.

[0082] In this embodiment, the detection results from different perspectives refer to the positions and confidence levels of human targets detected from different groups of visible light images and infrared images.

[0083] In this embodiment, in order to perform human target detection and tracking on the detection results from different perspectives, that is, the positions and confidences of human targets captured by different groups of visible light images and infrared images, step S120 is crucial. This step aims to process the initial data to obtain accurate detection results and ensure that these results can accurately reflect the status of each monitored object at each time point. The following will describe this process in detail, including two main sub-steps: S121 and S122.

[0084] In one embodiment, refer to Figure 2 , the above step S120 may include steps S121 to S122.

[0085] S121. Apply the pre-trained YOLOv5 model to analyze the initial data to obtain the bounding box of each human target and the corresponding confidence, forming a detection result, where each human target includes human targets from different perspectives.

[0086] In this embodiment, the pre-trained YOLOv5 algorithm model is used to analyze the initial image data collected from cameras (including visible light cameras and infrared cameras). This model can efficiently identify and locate human targets in the image, generate a bounding box for each detected target, and output the associated confidence value. Specifically, the YOLOv5 model will output a list containing information about each detected target, such as [x_min, y_min, x_max, y_max, class_id, confidence]. Here, (x_min, y_min) and (x_max, y_max) define the upper left and lower right coordinates of the bounding box, class_id represents the detected object category, and confidence represents the confidence of the model in detecting this target. Through this process, we can obtain the bounding boxes of human targets from different perspectives and their corresponding confidences, forming a preliminary detection result.

[0087] S122. Based on the detection result, use the Hungarian algorithm to minimize the cost matrix to match the human targets in the current frame with the human targets in the previous frame, and update the trajectory numbers of the human targets.

[0088] In this embodiment, after the preliminary detection is completed, the next problem to be solved is the multi-target tracking problem, that is, to ensure that each detected human target can be continuously tracked even if they move in different image frames or appear in different perspectives. To this end, the Hungarian algorithm is used to minimize the cost matrix, so as to best match the human targets detected in the current frame with the human targets in the previous frame. Each element in the cost matrix represents the matching cost between a certain target in the current frame and the target in the previous frame, which is usually calculated based on the relative positions between the targets. Through the Hungarian algorithm, we can not only find the best matching scheme that minimizes the total cost, but also update the trajectory number of each detected target every time a target is detected to maintain the consistency of tracking. For the targets that cannot be directly detected, the Kalman filter is used to predict their possible positions to maintain the continuity and integrity of the tracking chain.

[0089] In summary, step S120 combines advanced object detection and multi-target tracking technologies, ensuring the effective detection and stable tracking of human targets in image data obtained from different perspectives, providing a solid foundation for subsequent behavior analysis.

[0090] S130. Match the detection results from different perspectives and calculate the three-dimensional position information of the human targets in the matched detection results to obtain a calculation result.

[0091] In this embodiment, the calculation result refers to the distance between the human head and the water surface.

[0092] In one embodiment, please refer to Figure 3 the above step S130 may include steps S131 to S132.

[0093] S131. Match the detection results from different perspectives through the epipolar line matching algorithm.

[0094] In this embodiment, the epipolar line matching algorithm is used to accurately match the pixel points of the same human target in images from two different perspectives. This process first requires calculating the epipolar line equation based on the internal parameters (such as focal length, principal point, etc.) and external parameters (such as rotation and translation vectors) of the camera. The epipolar line equation defines the line on which the corresponding point of a given point in one perspective lies in the other perspective. Then, the human bounding boxes detected in the two perspectives are used to find the matching targets through the constraints of the epipolar line and the mask. This means that for each target detected in the left perspective, the most likely corresponding point along the epipolar line is searched for in the right perspective, ensuring that these two points belong to the same object in the real world.

[0095] S132. Calculate the three-dimensional position information of the human target based on the triangulation method for the matched detection results, and calculate the distance between the human head and the water surface according to the three-dimensional position information to obtain a calculation result.

[0096] In this embodiment, once the matching detection results are found, the next step is to use triangulation to calculate the three-dimensional position information of these human targets. Specifically, after the matching is completed, the coordinates of the matching points in two perspectives and the projection matrices of the left and right cameras can be input. The projection matrix contains information about the internal and external parameters of the camera, which are used to project points in three-dimensional space onto the two-dimensional image plane. The pixel coordinates (x L , x R ) of the matching points combined with the projection matrices (P L , P R ) of the cameras can be used to solve for the three-dimensional coordinates (X, Y, Z) of the target point, that is: x L = P L ·X, x R = P R ·X; where X = (X, Y, Z, 1) T is the homogeneous coordinate of the target point in three-dimensional space. Through the above formula, the true three-dimensional coordinates of the target point can be recovered from the pixel coordinates of the two images.

[0097] Finally, using the pre-established horizontal plane equation, when the three-dimensional coordinates of the target point are calculated, the position (X, Y, Z) of the target person's head can be obtained. Then, according to the preset horizontal plane information of the swimming pool, the vertical distance between the target person's head and the water surface can be calculated; the calculation formula is: d head-water = Z head- - Z water ; Z head is the coordinate of the calculated head on the Z-axis, and Z water represents the position of the water surface. This distance value d head-water is the required calculation result, which can reflect the height of the head relative to the water surface, and this is crucial for monitoring the safety of swimmers.

[0098] In summary, step S130, by combining the epipolar line matching algorithm and triangulation, realizes the accurate matching and three-dimensional positioning of the human target detection results from different perspectives, and finally calculates the distance between the head and the water surface, providing key data support for subsequent applications.

[0099] S140. Map the human targets in the matched detection results to the corresponding infrared perspectives, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information.

[0100] In this embodiment, the infrared temperature difference information refers to the infrared temperature difference information after Gaussian filtering processing, which can be used to infer the state of the human target, such as sinking or floating.

[0101] Specifically, map the detected human targets in the matched detection results to the corresponding infrared view, read and process the temperature information, and apply the Gaussian difference filtering algorithm to highlight the temperature change regions, and infer the current state of the human targets to obtain the infrared temperature difference information.

[0102] In this embodiment, first, based on the result of epipolar matching, map the pixel point (x, y) in the visible light image to the corresponding point (x', y') in the infrared image. This process can be completed by the following transformation formula: Here, K represents the internal parameter matrix of the camera, R is the rotation matrix, and t is the translation vector. This formula is used to convert the point in the visible light camera coordinate system to the infrared camera coordinate system to ensure that the target previously identified in the visible light image can be accurately located in the infrared image.

[0103] Once the coordinate transformation is completed, read the temperature information of the target area from the perspective of the infrared camera to obtain the temperature image. This temperature image represents the thermal distribution of the area where the target is located and provides the basic data for further analysis.

[0104] Next, use the Gaussian difference filtering algorithm to process the temperature information in the infrared image, aiming to highlight the regions with significant temperature changes, which helps to more accurately analyze the dynamic state of the target. Specifically, the Gaussian function is defined as follows: Where σ is the standard deviation of the Gaussian kernel, controlling the degree of blurring. The Gaussian difference filtering is the difference between two Gaussian filters with different standard deviations σ1 and σ2: DoG(x,y) = G(x,y,σ1) - G(x,y,σ2); where σ1 and σ2 are the standard deviations of two different scales respectively.

[0105] For the temperature data in the infrared image, applying the Gaussian difference filter means performing two Gaussian blur processes on the infrared image, each time using different standard deviations σ1 and σ2, and then calculating the difference between the two blurred images to obtain the temperature difference image I DoG ,

[0106] The regions with larger differences in the temperature difference image indicate the places where the temperature changes rapidly, and these places are usually associated with human activities or state changes.

[0107] Finally, apply threshold segmentation to the temperature difference image I DoG Set a threshold T for the temperature change thresh , and mark the regions with a temperature difference greater than this threshold as high-temperature regions. By this method, the temperature change regions of concern can be locked. The temperature difference change can be used to analyze the current state of the target, such as whether it is in an active state or whether there are obvious body temperature fluctuations, etc.

[0108] Combined with the established temperature change model, according to the temperature difference change situation in the target area, the state of the target can be further speculated. For example, if a certain area shows a rapidly rising temperature, it may mean that intense movement is taking place in that part; while a relatively stable temperature may indicate that the target is in a stationary state.

[0109] In summary, through a series of carefully designed steps, the target mapping from visible light images to infrared images, as well as the reading and processing of temperature information, are achieved. Finally, the temperature change area is highlighted through the Gaussian difference filtering algorithm, and the current state of the human target is speculated based on the temperature difference change analysis, obtaining the infrared temperature difference information. This series of operations provides solid data support for subsequent behavior judgment.

[0110] S150. Input the initial data into the behavior speculation model to analyze the floating or sinking behavior of the target, so as to obtain a speculation result.

[0111] In this embodiment, the speculation result refers to the result obtained by speculating the floating or sinking behavior of the human target.

[0112] Specifically, the behavior speculation model is obtained by using the drowning action video data with annotations including floating and sinking states as a sample set, training with the SlowOnly model as the basic architecture and combining it with a deep learning network.

[0113] In one embodiment, please refer to Figure 4 , the above-mentioned behavior speculation model is obtained by using the drowning action video data with annotations including floating and sinking states as a sample set, training with the SlowOnly model as the basic architecture and combining it with a deep learning network, and may include steps S151 to S154.

[0114] S151. Obtain the drowning action video data including floating and sinking states, and perform type annotation to obtain a sample set.

[0115] S152. Construct a two-stream network using the SlowOnly model as the basic architecture and combining it with a deep learning network to obtain an initial model;

[0116] S153. Use a 3D convolutional neural network and a convolutional long short-term memory network to process the spatio-temporal dynamics in the video frames of the sample set, and combine the processed RGB features and infrared features together through the method of early fusion or late fusion to obtain a feature result;

[0117] S154. Use the feature result to train the initial model to obtain a behavior speculation model.

[0118] In this embodiment, first, a large amount of video data of drowning actions including floating and sinking states is collected. This video data should be type - labeled, that is, each video frame should have a clear label indicating whether the target in it is in a floating or sinking state. The purpose of this step is to create a comprehensive and diverse sample set to ensure that the model trained subsequently has good generalization ability and can accurately classify drowning behaviors in different scenarios in practical applications.

[0119] Next, based on the above - mentioned sample set, an initial model with a two - stream network structure is constructed. This model is based on the SlowOnly model as the basic architecture and combines a custom deep - learning network to optimize performance. The two - stream network processes RGB (Red, Green, Blue) data stream and infrared data stream respectively, aiming to improve the recognition accuracy of the target floating or sinking behavior through the method of multi - modal fusion. At this stage, the network architecture is designed, including determining the number and parameter settings of components such as convolutional layers, pooling layers, and fully - connected layers, and at the same time, appropriate activation functions and loss functions are selected.

[0120] In this step, a 3D convolutional neural network (3D CNN) is used to process the temporal information in the video frames, and a convolutional long short - term memory network (ConvLSTM) is used for time - series modeling to capture the temporal dependencies between frames. For the data of both RGB and infrared modalities, the method of early fusion or late fusion is adopted to combine them. Specifically, in early fusion, RGB and infrared features are merged at the input end of the network; while in late fusion, RGB and infrared data pass through their respective networks, and finally their features are concatenated or weighted and summed. The purpose of doing this is to strengthen the information complementarity between the two modalities, thereby enhancing the system's ability to analyze complex drowning behaviors.

[0121] Finally, the feature results obtained in step S153 are used to train the initial model to obtain the final behavior inference model. During this process, the cross - entropy loss function is used to evaluate the difference between the probability distribution predicted by the model and the true label, and this is used as the optimization target to adjust the model weights. In addition, we also introduce the Softmax activation function for the final classification prediction to obtain the probability distribution of the target behavior categories. In particular, considering the different contribution degrees of RGB and infrared image features, the features of each modality are weighted by weighting coefficients α and β (satisfying α + β = 1) to ensure the fused feature F fused = αF RGB + βF IR , where F RGB and F IRThey are the features of RGB visible light and infrared images respectively. α and β are weighting coefficients, satisfying α + β = 1. The fused features are used for target behavior classification. The final classification prediction is carried out through a fully connected layer, and the Softmax activation function is used to obtain the probability distribution: P(behavior) = Softmax(W·F fused + b), where W is the weight matrix, F fused is the fused feature, and b is the bias term.

[0122] In summary, through the above steps, an efficient and accurate behavior inference model is constructed. This model can effectively identify the floating and sinking states in drowning behavior, providing important technical support for rescuing drowning victims in time. Throughout the process, from the preparation of the sample set to the training of the model, then to feature extraction and fusion, and finally to the final behavior classification prediction, each link is closely connected and indispensable.

[0123] S160. Fuse and analyze the calculation result, the infrared temperature difference information, and the inference result to obtain the final drowning determination result.

[0124] In this embodiment, the final drowning determination result refers to the prediction result of whether there is a drowning behavior for the human target.

[0125] In addition, when the final drowning determination result indicates a drowning behavior, based on the duration of the drowning behavior and the frame ratio, determine whether to generate a drowning warning message, and send the generated drowning warning message to the terminal for display.

[0126] In this embodiment, in order to achieve accurate judgment of drowning behavior, the system adopts a multi-modal feature information fusion strategy, comprehensively analyzes key conditions from different steps, including the distance between the target person's head and the water surface, the infrared temperature difference information after Gaussian filtering, and the confidence of the behavior model in the floating or sinking state of the target. These information are integrated to optimize the drowning judgment and effectively filter out false alarms, ensuring the reliability and accuracy of the monitoring system.

[0127] Specifically, fuse and analyze the calculation result (i.e., the distance between the target person's head and the water surface), the infrared temperature difference information (i.e., the infrared temperature difference information after Gaussian filtering), and the inference result (i.e., the confidence of the behavior model in the floating or sinking state of the target) to obtain the final drowning determination result.

[0128] For the successfully matched human target: Select an appropriate head distance alarm threshold according to the confidence of the behavior model, and combine the Gaussian temperature difference information, and use the preset constraint conditions to judge whether the target is in a dangerous state.

[0129] When the confidence level is greater than 0.6, a relatively loose distance alarm threshold (-0.6m to 0.1m) is adopted; if the target is within this range, the Gaussian temperature difference information is further checked (if the target Gaussian temperature difference is less than 0.2, the target is considered safe, otherwise it is marked as dangerous).

[0130] When the confidence level is less than 0.6, a more stringent distance alarm threshold (-0.6m to 0.06m) is adopted, and if the confidence level is below 0.2, the target is directly considered safe; otherwise, the Gaussian temperature difference information is continued to be checked (if the confidence level of the behavior is less than 0.2, the target is considered safe, otherwise the constraint of "if the target Gaussian temperature difference is less than 0.2, the target is considered safe, otherwise it is marked as dangerous" is carried out).

[0131] For the unmatched human target: Since the distance between the head and the water surface cannot be calculated, the system will mainly rely on the confidence level of the behavior model and the Gaussian temperature difference to make a judgment.

[0132] If the confidence level exceeds 0.7, the next step is to check the Gaussian temperature difference (that is, for the case where the Gaussian temperature difference is less than 0.2, the target is also considered safe; in other cases, it is marked as dangerous); conversely, if the confidence level is low, the target is defaulted to be safe.

[0133] For the case where the Gaussian temperature difference is less than 0.2, the target is also considered safe; in other cases, it is marked as dangerous.

[0134] For the target under the infrared view: First, evaluate whether the temperature difference between the head and the water surface exceeds the danger threshold, and then confirm whether there is a reliable measurement of the distance of the human head. If not, a decision is made based on the confidence level of the behavior model.

[0135] Once the above fusion analysis concludes that the target may be in a drowning state, the system will further determine whether to generate a drowning warning message based on the duration of the drowning behavior and the proportion of frames. Specifically:

[0136] If the time when the target is judged to be in a dangerous state exceeds 30 seconds; or, during the 30 seconds of continuously monitoring the target, the proportion of frames marked as dangerous in the total number of frames is greater than four-fifths.

[0137] When either of the above conditions is met, the system will generate a drowning warning message and send it to the terminal for display to promptly notify the lifeguard to take necessary rescue measures. The entire process aims to ensure that any potential drowning incident can be responded to promptly while minimizing false alarms, thereby improving the safety level of public swimming pools.

[0138] Specifically, for the successfully matched human target, calculate the distance between the head and the water surface, and select the head distance alarm threshold according to the confidence level of the behavior model. The following constraint conditions determine whether the target is in a dangerous or safe state.

[0139] Condition 1: If the confidence level of the behavior model is greater than 0.6, then select the distance alarm threshold as -0.6m - 0.1m. If the target is within the alarm threshold, perform the constraint of Condition 4; otherwise, consider the target safe.

[0140] Condition 2: If the confidence level of the behavior is less than 0.6, then select the distance alarm threshold as -0.6m - 0.06m. If the target is within the alarm threshold, perform the constraint of Condition 3; otherwise, consider the target safe.

[0141] Condition 3: If the confidence level of the behavior is less than 0.2, then consider the target safe; otherwise, perform the constraint of Condition 4.

[0142] Condition 4: If the Gaussian temperature difference of the target is less than 0.2, then consider the target safe; otherwise, mark it as dangerous.

[0143] For the unmatched human targets, the distance between the head and the water surface cannot be calculated, but the target status will be judged according to the confidence level of the behavior model and the Gaussian temperature difference. The following constraint conditions are used to determine whether the target is in a dangerous or safe state.

[0144] Condition 1: If the confidence level of the behavior model is greater than 0.7, perform the constraint of Condition 3; otherwise, consider the target safe.

[0145] Condition 2: If the confidence level of the behavior is less than 0.7, then consider the target safe.

[0146] Condition 3: If the Gaussian temperature difference of the target is less than 0.2, then consider the target safe; otherwise, mark it as dangerous.

[0147] For the targets under the infrared view, first judge whether the temperature difference between the head of the human target and the water surface exceeds the danger threshold, then check whether the target is successfully matched to obtain the distance between the head and the water surface, and further judge whether it is within the safe range. If neither meets the safety standards, then judge the confidence level of the behavior model. The following constraint conditions are used to determine whether the target is in a dangerous or safe state.

[0148] Condition 1: If the Gaussian temperature difference is within the alarm threshold range, perform the constraint of Condition 2; otherwise, consider the target safe.

[0149] Condition 2: If the target has the distance between the head and the water surface, then judge whether the head distance is within the safe threshold range. If it is, then consider the target safe; otherwise, perform the constraint of Condition 3.

[0150] Condition 3: If the confidence level of the behavior is less than 0.2, then consider the target safe; otherwise, mark it as dangerous.

[0151] Taking into account the judgment results of the above three situations, when the following constraint conditions are met, the system will send a drowning danger rescue request to the lifeguard.

[0152] Condition 1: The target is in a dangerous state for more than 30 seconds.

[0153] Condition 2: Within 30 seconds of focusing on the target, the number of dangerous frames is greater than four-fifths of the total number of frames.

[0154] S170. Output the final drowning determination result.

[0155] Output the final drowning determination result to the terminal for display.

[0156] The method of this embodiment realizes the efficient recognition of drowning behavior by combining the data characteristics of an infrared camera and a visible light camera. This method not only considers physical parameters such as human body temperature and the distance between the human head and the water surface, but also introduces a confidence evaluation based on a behavior model to enhance the ability to judge potential drowning events.

[0157] Compared with traditional single-mode monitoring means, the method of this embodiment has the following significant advantages:

[0158] Higher recognition rate of dangerous targets: By integrating information from different modalities (such as infrared and visible light), the system can capture target features more comprehensively, thereby improving the detection accuracy of drowning behavior.

[0159] Lower false alarm rate: Using multi-condition fusion technology, including but not limited to the change of human body temperature, position information, and behavior pattern analysis of the target, effectively filters out non-real drowning situations and reduces unnecessary alarms.

[0160] Real-time and intelligent: This system supports instant data processing and intelligent decision-making, can detect and respond to drowning risks in the first time, and ensures timely rescue.

[0161] Broad application prospects: Especially in the safety management and intelligent monitoring of swimming pools, the solution provided by the present invention can significantly improve the safety performance of the venue and escort public health.

[0162] Particularly, the method of this embodiment focuses on optimizing the recognition recall rate of two typical drowning behaviors, floating and sinking. Through the deep fusion of multi-modal data and the support of advanced algorithms, a more accurate and reliable judgment of the drowning state is achieved. This progress is crucial for improving the overall efficiency of the intelligent monitoring system and brings an innovative breakthrough to related fields.

[0163] The above drowning alarm method based on the fusion of multiple strategies uses infrared and visible light cameras installed above the swimming pool. The system obtains initial image data, performs human target detection and tracking on this data, and obtains detection results from different perspectives. Then, the system calculates the three-dimensional position information of the human target by matching the detection results from different perspectives, maps it to the infrared perspective, obtains and analyzes the temperature information, and obtains the infrared temperature difference information. After inputting the initial data into the behavior speculation model, it analyzes the floating or sinking behavior of the target to obtain the speculation result. Then, by combining the calculation result, the infrared temperature difference information, and the speculation result, it performs multi-condition fusion analysis to reduce false alarms and finally obtains the drowning determination result. This technology provides more reliable support for drowning prevention and rescue.

[0164] Figure 5 FIG. 4 is a schematic block diagram of a drowning alarm system 300 based on the fusion of multiple strategies provided by an embodiment of the present invention. As Figure 5 shown, corresponding to the above drowning alarm method based on the fusion of multiple strategies, the present invention also provides a drowning alarm system 300 based on the fusion of multiple strategies. The drowning alarm system 300 based on the fusion of multiple strategies includes units for executing the above drowning alarm method based on the fusion of multiple strategies, and this system can be configured in a server. Specifically, please refer to Figure 5 , the drowning alarm system 300 based on the fusion of multiple strategies includes an acquisition unit 301, a detection and tracking unit 302, a calculation unit 303, a temperature analysis unit 304, a speculation unit 305, a fusion analysis unit 306, and an output unit 307.

[0165] The acquisition unit 301 is used to acquire the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain the initial data; the detection and tracking unit 302 is used to perform human target detection and tracking on the initial data to obtain detection results from different perspectives; the calculation unit 303 is used to match the detection results from different perspectives and calculate the three-dimensional position information of the human target in the matched detection results to obtain the calculation result; the temperature analysis unit 304 is used to map the human target in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and analyze the temperature information to obtain the infrared temperature difference information; the speculation unit 305 is used to input the initial data into the behavior speculation model to analyze the floating or sinking behavior of the target to obtain the speculation result; the fusion analysis unit 306 is used to perform fusion analysis on the calculation result, the infrared temperature difference information, and the speculation result to obtain the final drowning determination result; the output unit 307 is used to output the final drowning determination result.

[0166] In one embodiment, the obtaining unit 301 is configured to collect image data captured by an infrared camera and a visible light camera installed above the swimming pool in parallel using multi-threading technology to obtain initial data.

[0167] In one embodiment, the detection and tracking unit 302 includes:

[0168] A detection subunit, configured to analyze the initial data using a pre-trained YOLOv5 model to obtain the bounding box and corresponding confidence level of each human target, forming a detection result, where each human target includes human targets from different perspectives; a tracking subunit, configured to minimize the cost matrix based on the detection result using the Hungarian algorithm to match the human targets in the current frame with the human targets in the previous frame, and update the trajectory numbers of the human targets.

[0169] In one embodiment, the calculation unit 303 includes:

[0170] A matching subunit, configured to match the detection results from different perspectives through an epipolar matching algorithm; a distance calculation subunit, configured to calculate the three-dimensional position information of the human target based on triangulation for the matched detection results, and calculate the distance between the human head and the water surface according to the three-dimensional position information to obtain a calculation result.

[0171] In one embodiment, the temperature analysis unit 304 is configured to map the human targets in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and apply a Gaussian difference filtering algorithm to highlight the temperature change region, and infer the current state of the human targets to obtain infrared temperature difference information.

[0172] In one embodiment, the drowning alarm system 300 based on multi-strategy fusion further includes: a training unit, configured to:

[0173] Obtain drowning action video data including floating and sinking states, and perform type annotation to obtain a sample set; construct a two-stream network using the SlowOnly model as the basic architecture combined with a deep learning network to obtain an initial model; process the spatio-temporal dynamics in the video frames using a 3D convolutional neural network and a convolutional long short-term memory network for the sample set, and combine the processed RGB features and infrared features together through early fusion or late fusion methods to obtain a feature result; use the feature result to train the initial model to obtain a behavior inference model.

[0174] In one embodiment, the drowning alarm system 300 based on multi-strategy fusion further includes:

[0175] An information generation unit, configured to determine whether to generate a drowning warning message based on the duration and frame ratio of the drowning behavior when the final drowning determination result indicates the existence of a drowning behavior, and send the generated drowning warning message to a terminal for display.

[0176] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above-mentioned drowning alarm system 300 based on the integration of multiple strategies and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and conciseness of description, they will not be elaborated here.

[0177] The above-mentioned drowning alarm system 300 based on the integration of multiple strategies can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 6 shown.

[0178] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a computer device provided by an embodiment of the present application. This computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.

[0179] Refer to Figure 6 , this computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.

[0180] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. This computer program 5032 includes program instructions, and when these program instructions are executed, the processor 502 can be made to execute a drowning alarm method based on the integration of multiple strategies.

[0181] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0182] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When this computer program 5032 is executed by the processor 502, the processor 502 can be made to execute a drowning alarm method based on the integration of multiple strategies.

[0183] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 6The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. Specifically, the computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0184] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps:

[0185] Obtain the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data; perform human target detection and tracking on the initial data to obtain detection results from different perspectives; match the detection results from different perspectives, and calculate the three-dimensional position information of the human target in the matched detection results to obtain a calculation result; map the human target in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information; input the initial data into the behavior speculation model to analyze the floating or sinking behavior of the target to obtain a speculation result; perform fusion analysis on the calculation result, the infrared temperature difference information, and the speculation result to obtain a final drowning determination result; output the final drowning determination result.

[0186] Among them, the total number of the infrared camera and the visible light camera is 16. One infrared camera and one visible light camera form a group of cameras, and each group of cameras observes the same area.

[0187] The behavior speculation model is obtained by using the drowning action video data with annotations including floating and sinking states as a sample set to train a model that uses the SlowOnly model as the basic architecture and combines a deep learning network.

[0188] In one embodiment, when the processor 502 implements the step of obtaining the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data, the following steps are specifically implemented:

[0189] Adopt multi-thread technology to parallelly collect the image data captured by the infrared camera and the visible light camera installed above the swimming pool to obtain initial data.

[0190] In one embodiment, when the processor 502 implements the step of performing human target detection and tracking on the initial data to obtain detection results from different perspectives, the following steps are specifically implemented:

[0191] Analyze the initial data using the pre-trained YOLOv5 model to obtain the bounding box and corresponding confidence level of each human target, forming a detection result, where each human target includes human targets from different perspectives; based on the detection result, use the Hungarian algorithm to minimize the cost matrix to match the human targets in the current frame with those in the previous frame, and update the trajectory numbers of the human targets.

[0192] In one embodiment, when the processor 502 implements the step of matching the detection results from different perspectives and calculating the three-dimensional position information of the human targets in the matched detection results to obtain a calculation result, the specific implementation is as follows:

[0193] Match the detection results from different perspectives through the epipolar matching algorithm; calculate the three-dimensional position information of the human targets based on the triangulation method for the matched detection results, and calculate the distance between the human head and the water surface according to the three-dimensional position information to obtain a calculation result.

[0194] In one embodiment, when the processor 502 implements the step of mapping the human targets in the matched detection results to the corresponding infrared perspective, reading and processing the temperature information, and analyzing the temperature information to obtain infrared temperature difference information, the specific implementation is as follows:

[0195] Map the human targets in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and apply the Gaussian difference filtering algorithm to highlight the temperature change region and infer the current state of the human targets to obtain infrared temperature difference information.

[0196] In one embodiment, when the processor 502 implements the step that the behavior inference model is obtained by training with the video data of drowning actions with floating and sinking states with annotations as the sample set, using the SlowOnly model as the basic architecture and combining it with a deep learning network, the specific implementation is as follows:

[0197] Obtain the video data of drowning actions with floating and sinking states and perform type annotation to obtain a sample set; construct a two-stream network using the SlowOnly model as the basic architecture and combining it with a deep learning network to obtain an initial model; use a 3D convolutional neural network and a convolutional long short-term memory network to process the spatio-temporal dynamics in the video frames for the sample set, and combine the processed RGB features and infrared features together through the method of early fusion or late fusion to obtain a feature result; use the feature result to train the initial model to obtain a behavior inference model.

[0198] In one embodiment, after the processor 502 implements the step of fusing and analyzing the calculation result, the infrared temperature difference information, and the speculation result to obtain the final drowning determination result, the following steps are further implemented:

[0199] When the final drowning determination result indicates the existence of a drowning behavior, based on the duration of the drowning behavior and the frame ratio, determine whether to generate a drowning warning message, and send the generated drowning warning message to the terminal for display.

[0200] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0201] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0202] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the following steps:

[0203] Obtain the image data captured by the infrared camera and visible light camera installed above the pool to obtain initial data; perform human target detection and tracking on the initial data to obtain detection results from different perspectives; match the detection results from different perspectives and calculate the three-dimensional position information of the human targets in the matched detection results to obtain a calculation result; map the human targets in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information; input the initial data into a behavior speculation model to analyze the floating or sinking behavior of the target to obtain a speculation result; fuse and analyze the calculation result, the infrared temperature difference information, and the speculation result to obtain a final drowning determination result; output the final drowning determination result.

[0204] Among them, the total number of the infrared camera and the visible light camera is 16. One infrared camera and one visible light camera form a group of cameras, and each group of cameras observes the same area.

[0205] The behavior speculation model is obtained by using the drowning action video data with annotations including floating and sinking states as a sample set, training with the SlowOnly model as the basic architecture and combining it with a deep learning network.

[0206] In one embodiment, when the processor executes the computer program to implement the step of obtaining the image data captured by the infrared camera and visible light camera installed above the pool to obtain initial data, the specific implementation is as follows:

[0207] Adopt multi-thread technology to parallelly collect the image data captured by the infrared camera and visible light camera installed above the pool to obtain initial data.

[0208] In one embodiment, when the processor executes the computer program to implement the step of performing human target detection and tracking on the initial data to obtain detection results from different perspectives, the specific implementation is as follows:

[0209] Apply the pre-trained YOLOv5 model to analyze the initial data to obtain the bounding box and corresponding confidence of each human target, forming a detection result. Among them, each human target includes human targets from different perspectives; based on the detection result, use the Hungarian algorithm to minimize the cost matrix to match the human targets in the current frame with the human targets in the previous frame, and update the trajectory number of the human targets.

[0210] In one embodiment, when the processor executes the computer program to implement the step of matching the detection results from different perspectives and calculating the three-dimensional position information of the human target in the matched detection results to obtain a calculation result, the specific implementation is as follows:

[0211] Match the detection results from different perspectives through the epipolar matching algorithm; calculate the three-dimensional position information of the human target based on triangulation for the matched detection results, and calculate the distance between the human head and the water surface according to the three-dimensional position information to obtain a calculation result.

[0212] In one embodiment, when the processor executes the computer program to implement the step of mapping the human target in the matched detection results to the corresponding infrared perspective, reading and processing the temperature information, and analyzing the temperature information to obtain infrared temperature difference information, the specific implementation is as follows:

[0213] Map the human target in the matched detection results to the corresponding infrared perspective, read and process the temperature information, and apply the Difference of Gaussians (DoG) filtering algorithm to highlight the temperature change region and infer the current state of the human target to obtain infrared temperature difference information.

[0214] In one embodiment, when the processor executes the computer program to implement the step that the behavior inference model is obtained by training a SlowOnly model as the basic architecture combined with a deep learning network using the annotated drowning action video data containing floating and sinking states as a sample set, the specific implementation is as follows:

[0215] Obtain the drowning action video data containing floating and sinking states and perform type annotation to obtain a sample set; construct a two-stream network using the SlowOnly model as the basic architecture combined with a deep learning network to obtain an initial model; use a 3D convolutional neural network and a convolutional long short-term memory network to process the spatio-temporal dynamics in the video frames for the sample set, and combine the processed RGB features and infrared features together through the method of early fusion or late fusion to obtain a feature result; use the feature result to train the initial model to obtain a behavior inference model.

[0216] In one embodiment, after the processor executes the computer program to implement the step of fusing and analyzing the calculation result, the infrared temperature difference information, and the inference result to obtain a final drowning determination result, the following steps are further implemented:

[0217] When the final drowning determination result indicates the existence of a drowning behavior, determine whether to generate a drowning warning message based on the duration and frame ratio of the drowning behavior, and send the generated drowning warning message to the terminal for display.

[0218] The storage medium may be a variety of computer-readable storage media such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., which can store program codes.

[0219] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0220] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0221] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0223] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A drowning alarm method based on the fusion of multiple strategies, characterized in that: include: Obtain image data captured by an infrared camera and a visible light camera installed above the swimming pool to obtain initial data; Performing human target detection and tracking on the initial data to obtain detection results under different viewing angles; Matching the detection results under different viewing angles, and calculating the three-dimensional position information of the human target in the matched detection results to obtain a calculation result; Mapping the human target in the matched detection result to the corresponding infrared viewing angle, reading and processing the temperature information, and analyzing the temperature information to obtain infrared temperature difference information; Inputting the initial data into a behavior inference model to analyze the floating or sinking behavior of the target to obtain an inference result; The calculation result, the infrared temperature difference information and the inference result are integrated and analyzed to obtain a final drowning determination result; The final drowning determination result is outputted.

2. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: The total number of the infrared cameras and the visible light cameras is 16. One infrared camera and one visible light camera are combined to form a group of cameras, and each group of cameras observes the same area.

3. The drowning alarm method based on the fusion of multiple strategies according to claim 2 is characterized in that: The step of obtaining image data captured by an infrared camera and a visible light camera installed above the swimming pool to obtain initial data includes: Multi-threading technology is used to collect image data taken by the infrared camera and visible light camera installed above the swimming pool in parallel to obtain initial data.

4. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: The performing human target detection and tracking on the initial data to obtain detection results under different viewing angles includes: Applying a pre-trained YOLOv5 model to analyze the initial data to obtain a bounding box of each human target and a corresponding confidence level to form a detection result, wherein each human target includes human targets under different perspectives; Based on the detection result, the human target in the current frame is matched with the human target in the previous frame by minimizing the cost matrix using the Hungarian algorithm, and the trajectory number of the human target is updated.

5. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: The matching of the detection results under different viewing angles and calculating the three-dimensional position information of the human target in the matched detection results to obtain a calculation result includes: Matching the detection results under different viewing angles by using an epipolar matching algorithm; For the matched detection results, the three-dimensional position information of the human target is calculated based on the triangulation method, and the distance between the human head and the water surface is calculated according to the three-dimensional position information to obtain the calculation result.

6. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: The human body target in the matched detection result is mapped to the corresponding infrared viewing angle, the temperature information is read and processed, and the temperature information is analyzed to obtain infrared temperature difference information, including: The human target in the matched detection result is mapped to the corresponding infrared viewing angle, the temperature information is read and processed, and the Gaussian difference filtering algorithm is applied to highlight the temperature change area, and the current state of the human target is inferred to obtain the infrared temperature difference information.

7. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: The behavior inference model is obtained by training a model composed of a SlowOnly model as a basic architecture combined with a deep learning network using annotated drowning action video data containing floating and sinking states as a sample set.

8. The drowning alarm method based on the fusion of multiple strategies according to claim 7 is characterized in that: The behavior inference model is obtained by training a model composed of a SlowOnly model as a basic architecture combined with a deep learning network using annotated drowning action video data containing floating and sinking states as a sample set, including: Obtain drowning action video data including floating and sinking states, and perform type annotation to obtain a sample set; Build a two-stream network using the SlowOnly model as the infrastructure combined with a deep learning network to obtain an initial model; The sample set is processed using a 3D convolutional neural network and a convolutional long short-term memory network to process the spatiotemporal dynamics in the video frame, and the processed RGB features and infrared features are combined by an early fusion method or a late fusion method to obtain a feature result; The initial model is trained using the feature results to obtain a behavior inference model.

9. The drowning alarm method based on the fusion of multiple strategies according to claim 1 is characterized in that: After the calculation result, the infrared temperature difference information and the inference result are integrated and analyzed to obtain the final drowning determination result, the method includes: When the final drowning determination result is that there is drowning behavior, it is determined whether to generate drowning warning information based on the duration and frame ratio of the drowning behavior, and the generated drowning warning information is sent to the terminal for display.

10. The drowning alarm system based on the fusion of multiple strategies is characterized by: include: An acquisition unit, used to acquire image data captured by an infrared camera and a visible light camera installed above the swimming pool to obtain initial data; A detection and tracking unit, used to perform human target detection and tracking on the initial data to obtain detection results under different viewing angles; A calculation unit, used for matching the detection results under different viewing angles, and calculating the three-dimensional position information of the human target in the matched detection results to obtain a calculation result; A temperature analysis unit is used to map the human body target in the matched detection result to the corresponding infrared viewing angle, read and process the temperature information, and analyze the temperature information to obtain infrared temperature difference information; An inference unit, used for inputting the initial data into a behavior inference model to analyze the floating or sinking behavior of the target to obtain an inference result; A fusion analysis unit, used for fusing and analyzing the calculation result, the infrared temperature difference information and the inference result to obtain a final drowning determination result; An output unit is used to output the final drowning determination result.