Multi-vision lifesaving system and lifesaving method thereof
By constructing a 3D depth model and recognizing attitude and motion features using a multi-view vision system, and combining it with GNSS collaborative positioning to plan rescue routes, the problem of monitoring lag and slow response in water drowning monitoring has been solved, enabling all-weather high-precision drowning identification and rapid rescue.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DONGGUAN UNIV OF TECH
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies for drowning monitoring in water areas suffer from problems such as monitoring lag, slow response, and poor positioning. Traditional manual lookout mode is difficult to meet the needs of all-weather and wide coverage. Monocular vision is easily affected by water surface reflection and obstacles, resulting in poor recognition robustness. GNSS positioning accuracy is unstable. The segmented architecture of rescue systems is time-consuming and has a high misjudgment rate.
A multi-view vision system is used to collect image information from different directions, construct a 3D depth model, and identify drowning victims by combining attitude-motion features. The multi-view vision system is used to track, locate, and plan rescue routes. Extended Kalman filtering is used to achieve GNSS and multi-view vision collaborative positioning. A deep learning model is constructed to improve the recognition accuracy. An autonomous navigation algorithm is used to control the lifeboat for precise rescue.
It enables all-weather, high-precision monitoring and rapid response for drowning identification and rescue, reduces the misjudgment rate, ensures that lifeboats accurately and quickly reach the drowning point, and improves rescue efficiency.
Smart Images

Figure CN121990138A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emergency rescue technology, and in particular to a multi-vision rescue system and rescue method thereof. Background Technology
[0002] Drowning accidents are frequent in open waters (lakes, reservoirs, etc.), with most missing the golden rescue period due to "lagging monitoring, slow response, and poor positioning." Traditional manual lookout methods are insufficient to meet the needs of wide coverage and all-weather operation. Existing technologies have significant shortcomings in the entire process: In terms of water monitoring, low-cost monocular vision solutions are susceptible to water surface reflection and obstacle interference, resulting in low image signal-to-noise ratio and loss of target details. Single-station coverage is small, and multi-station deployment requires weekly manual calibration due to temperature and wind load external parameter drift, making it difficult to adapt to unattended operation. Point cloud solutions such as lidar are expensive and unsuitable for rainy or foggy weather. The point cloud has large gaps, resulting in poor robustness in recognition. It relies on a single feature, making it difficult to distinguish between playing in the water and drowning, leading to a high false positive rate. Biosensors need to be actively worn and have limited coverage. In terms of rescue positioning and dispatch, GNSS positioning accuracy drops from 1-3 meters to 10-20 meters or even interruptions of more than 10 seconds under bridges and trees, increasing search time and causing inertial navigation errors to accumulate, making long-term high-precision positioning impossible. In terms of full-process collaboration, the existing system has a segmented architecture of "monitoring-alarm-manual dispatch-rescue". Manual input of coordinates and setting of parameters is time-consuming, has a slow response, and is prone to failure due to operational errors. Summary of the Invention
[0003] The purpose of this invention is to overcome the above-mentioned defects in the prior art and provide a multi-view vision rescue system and its rescue control method. This invention has a low-cost pure image-driven solution, realizes all-weather wide-coverage monitoring, high-precision positioning in all scenarios, automatic ejection rescue and return, fast response throughout the process, and improves rescue efficiency.
[0004] To achieve the above objectives, the present invention provides a multi-view vision lifesaving system, comprising:
[0005] Image information of the monitored water area is acquired from different directions using a multi-view vision system;
[0006] A three-dimensional depth model of the entire water area is constructed based on image information, and drowning identification is performed to determine whether there are people who have fallen into the water. When it is determined that there are people who have fallen into the water, their location is determined.
[0007] The rescue boat is activated based on the location of the person in the water, and the rescue boat and the person in the water are tracked and located in real time through a multi-view vision system.
[0008] By combining a 3D depth model to plan the optimal rescue route, the rescue boat is directed to accurately reach the designated location to carry out the rescue.
[0009] Furthermore, the image information is subjected to reflection suppression preprocessing, specifically: each image information is divided into blocks, and the average gray level of each block is calculated. When the average gray level of a block is greater than a threshold, it is determined to be a reflective block. Then, the average gray level of the reflective block is compressed so that the average gray level of the compressed reflective block is less than or equal to the threshold. Finally, Gaussian filtering is performed to eliminate noise interference.
[0010] Furthermore, during the tracking and positioning process, it is necessary to automatically correct the extrinsic parameters of the multi-view system. Specifically, fixed markers on the surface of the lifeboat are used as dynamic calibration objects. The coordinates are extracted from the captured images of the markers and combined with the known relative positions of the marker points to complete the automatic correction of the extrinsic parameters of the multi-view vision system.
[0011] Furthermore, the three-dimensional depth model is constructed based on the principle of multi-view geometry using images simultaneously collected from no less than three monitoring stations. It achieves spatial separation of the effective water area, the water background, and obstacles, thus forming environmental constraints for tracking and positioning.
[0012] Furthermore, when identifying whether a person has fallen into the water, it is necessary to identify the target's posture and motion features. By integrating the dual features of the human torso posture angle and limb movement frequency, an identification confidence model is constructed. When the output value of the confidence model exceeds a set threshold, it is determined to be a person who has fallen into the water, triggering an alarm and initiating a rescue process. At the same time, the drowning point is located, and the coordinates of the monitoring station of the multi-view vision system are obtained. Combining the monitoring station coordinates with the constraints of the 3D depth model, non-water background targets are first filtered out, and then the world coordinates of the drowning point are optimized through weighted fusion.
[0013] Furthermore, the real-time location of the drowning victim is predicted based on historical location data, and the rescue boat is guided to conduct the rescue based on the latest predicted real-time location.
[0014] Furthermore, multi-view vision and GNSS positioning are used for collaborative positioning of the lifeboats. Multi-view vision marker measurement, target vision measurement and GNSS data are integrated, and extended Kalman filtering is used to achieve accurate positioning in all scenarios to ensure the continuity of lifeboat navigation.
[0015] Furthermore, the optimal rescue route planning includes: detecting obstacle locations through multi-view vision, constructing a grid map, and planning a collision-free path based on the cost function of the improved A* algorithm to ensure that the lifeboat safely and quickly reaches the drowning point.
[0016] Furthermore, a large deep learning model is constructed. Based on the 3D deep model, the stereoscopic image of the current scene and contextual information are input in real time to achieve end-to-end drowning recognition. The contextual information includes normal swimming, abnormal swimming, or intermittent water play.
[0017] The present invention also provides a multi-view vision rescue system, including a lifeboat, wherein a plurality of fixed markings are provided on the surface of the lifeboat for providing a positioning reference in multi-view vision monitoring;
[0018] The multi-view vision system includes several vision monitoring stations. Each vision monitoring station includes a battery module, a camera, a positioning module, an inertial navigation module, and a data acquisition module. The vision monitoring stations are arranged in a ring around the monitored water area, so that the cameras of each vision monitoring station can work together to achieve multi-angle collaborative shooting and realize full coverage monitoring of the rescue water area. The cameras can obtain image information of the rescue water area in real time. The data acquisition module is used to collect spatial information acquired by the positioning module and the inertial navigation module in real time.
[0019] The processing system receives and processes image and spatial information, calculates the position data of each target in the environment using multi-view vision measurement methods, and identifies drowning people in water targets using stereo perception algorithms. When a drowning person is detected, it issues a command to control the lifeboat to start, calculates the position of the lifeboat through multi-view vision measurement algorithms, and controls the lifeboat to navigate to the drowning location to carry out the rescue using autonomous navigation algorithms.
[0020] The communication system is used to enable real-time two-way communication between the multi-view vision system, the processing system, and the lifeboat.
[0021] This invention sets up no fewer than three monitoring stations in the rescue water area to form a multi-view vision system. The multi-view vision system collects multi-directional image information in the rescue water area and constructs a three-dimensional depth model.
[0022] The processing system acquires data from each visual monitoring station through the communication system, calculates the position and orientation of the cameras at each visual monitoring station, and then uses multi-view visual measurement methods to calculate the position data of each target in the environment. At the same time, it uses stereo perception algorithms to identify the various states of targets in the water area, thereby detecting drowning victims.
[0023] When a drowning person is detected, a command is sent via the communication system to activate the lifeboat. After the processing system detects the lifeboat, it calculates the lifeboat's position using a multi-view vision measurement algorithm and simultaneously uses an autonomous navigation algorithm to guide the lifeboat to the location of the drowning to carry out the rescue, thereby improving the lifesaving system's efficiency.
[0024] The system optimizes the drowning recognition model through a continuous learning mechanism and dynamically adjusts rescue strategies by combining meteorological and water flow data to ensure response accuracy in complex environments. The lifeboat is equipped with a self-stabilizing propulsion device, enabling it to accurately and quickly reach the drowning site for rescue based on system instructions. The deep integration of a multi-view vision system and an inertial navigation module effectively improves the robustness of 3D positioning, allowing the lifeboat to accurately approach the drowning site even with obstacles.
[0025] Meanwhile, the present invention uses abnormal patterns of limb movement characteristics and water surface interaction behavior as the criteria for identifying drowning victims, combined with a human posture estimation model for dynamic analysis, to ensure an accuracy rate of over 98%. The system continuously transmits data through a visual monitoring station, updates the drowning victim's location in real time, and predicts their drift trend with the water flow, planning the optimal path for the lifeboat.
[0026] This invention constructs an all-weather, three-dimensional detection, identification, positioning, and rescue system, providing effective protection for water rescue. Attached Figure Description
[0027] To more clearly illustrate the technology in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0028] Figure 1 This is a flowchart illustrating the multi-view vision-based lifesaving method of the present invention;
[0029] Figure 2 This is a schematic diagram showing the distribution of the visual monitoring station and life-saving device of the present invention.
[0030] The diagram includes:
[0031] 1. Visual monitoring station; 2. Lifeboat; 3. Server. Detailed Implementation
[0032] The technology of this embodiment of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiment is one embodiment of the present invention, and not all embodiments thereof. Based on this embodiment of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0034] Furthermore, if the embodiments of the present invention involve descriptions such as "first" or "second", such descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated.
[0035] This invention provides a multi-view vision-based lifesaving method, comprising the following:
[0036] like Figure 1 and Figure 2 As shown, multiple visual monitoring stations are first deployed around the water area to form a multi-view visual system to ensure coverage of the entire monitored water area. Each monitoring station is powered by solar energy and uses cameras to collect water surface image information in real time. The multi-view visual system collects image information of the monitored water area from different directions.
[0037] The acquired image information needs to undergo reflection suppression preprocessing to avoid interference from water surface reflections on image recognition and improve image quality. Specifically:
[0038] First, reflective block identification is performed: the acquired original image I(x, y) is divided into blocks, each block being of size [size missing]. And calculate the average gray level of each block. When the average gray level of the blocks is greater than the threshold, It is then determined to be a reflective block, as mentioned above. The reflective grayscale threshold, The grayscale mean is... This represents the maximum grayscale value.
[0039] Next, adaptive grayscale correction is performed on the reflector according to the formula. Gray-scale compression is applied to the average gray value of the reflector, where k is a correction coefficient. Ensure the average grayscale of the reflective blocks after correction. Non-reflective blocks retain their original grayscale value.
[0040] Finally, Gaussian filtering is applied to the corrected image to eliminate noise interference; reflection suppression preprocessing is completed. This processing method can effectively suppress image interference caused by water surface reflection, preserve the key visual features of the drowning person and lifeboat 2, and provide high-quality image input for subsequent identification and positioning.
[0041] Based on the collected image information, a three-dimensional depth model of the entire water area is constructed. To improve the accuracy of the constructed model, it is preferable to deploy no fewer than three visual monitoring stations 1 in the rescue water area. The time synchronization protocol between the stations ensures that the image acquisition sequence is consistent. The image data collected synchronously by each monitoring station is fused using the principle of multi-view geometry to generate a high-resolution three-dimensional depth model. This enables spatial separation of the effective water area, the water background, and obstacles, forming environmental constraints for subsequent tracking and positioning, as detailed below:
[0042] First, feature matching is performed on images captured by multiple monitoring stations, with each station's camera image matched according to frame rate. Simultaneous image acquisition, optimizing time synchronization error Each image is extracted using the ORB feature algorithm. Calculate feature matching pairs between images from different monitoring stations using 1 feature point. ,in Here are the coordinates of the feature points in the image at station k. For the coordinates of the matching point at station l, the RANSAC algorithm is used to remove false matches, with a retention rate of [missing information]. ;
[0043] Secondly, the relative pose of the cameras between monitoring stations is estimated, and the relative rotation matrix between station k and station l is solved by essential matrix decomposition based on the matching feature pairs. With translation vector Define the reprojection error function Perspective projection function , Assuming an initial depth hypothesis, where M is the number of matching pairs, the following requirements are made: ,in Pixel error threshold;
[0044] Then, a global depth map is generated. Each monitoring station uses the SGBM algorithm to generate a disparity map based on images acquired by binocular cameras, with the disparity range... Internal reference via camera (or other camera)
[0045] Convert to depth map The depth values of overlapping areas are weighted and fused according to "image sharpness", and the formula is as follows:
[0046] in, For sharpness weights, calculation is based on image gradient magnitude. ;
[0047] Finally, the water area model is segmented based on the depth map. With grayscale Construct a dual threshold constraint, where depth grayscale Areas that meet the constraints are considered "valid water areas".
[0048] Based on this, continuous monitoring of dynamic targets within the effective water area is conducted, and drowning identification is performed within the effective water area. When identifying whether a person has fallen into the water, it is necessary to identify the target's posture and motion characteristics. By fusing the dual features of the human torso posture angle and limb movement frequency, an identification confidence model is constructed. When the output value of the confidence model exceeds a set threshold, it is determined that a person has fallen into the water, triggering an alarm and initiating a rescue process. Specifically:
[0049] First, key points and features of the human body are extracted. Human regions are extracted from the image through semantic segmentation, and key points of the torso (shoulder, shoulder, etc.) are labeled. ,waist Hip ) and key points of the limbs (hands) ,foot ), calculate the torso posture angle (Angle between the torso and the water surface) , The water surface normal vector is obtained from background segmentation.
[0050] Then, calculate the limb movement frequency f and the statistical time window. The formula for the number of inner hand swings is: , For indicator functions, The motion threshold;
[0051] Finally, the recognition confidence score is constructed by weighted fusion of features after normalization using the Sigmoid function: ,in For feature weights, For the Sigmoid function, It was determined to be drowning.
[0052] Once a person is confirmed to have fallen into the water, the drowning point is located. The coordinates of the monitoring stations of the multi-view vision system are obtained. Combining the monitoring station coordinates with the constraints of the 3D depth model, non-aquatic background targets are first filtered out. Then, the world coordinates of the drowning point are optimized through weighted fusion. This solves the problems of background interference and multi-station data redundancy in traditional positioning, ensuring high-precision positioning based on pure image signals. The specific steps are as follows:
[0053] First, depth-constrained filtering is implemented. When a suspected target is detected, the target image coordinates are... Projected onto the generated depth map Obtain target depth ,like If it is determined to be background, it is excluded and does not require positioning; Then proceed to the fusion process;
[0054] Next, the world coordinates of multiple stations are transformed by converting the relative coordinates of each monitoring station to world coordinates using a homogeneous transformation matrix. The transformation matrix is
[0055] in, The IMU attitude angles of the cameras at each monitoring station. (GNSS coordinates of the monitoring station).
[0056] Then, the weighted coordinates are merged by introducing "distance weights". ( (GNSS coordinates of the monitoring station) and "depth confidence weight" ,in (where the average depth of the water area is ), the fusion formula is: .
[0057] Preferably, in this implementation, to improve the accuracy of drowning identification and minimize false positives or false negatives, thereby increasing the response speed and accuracy of rescue efforts, a large deep learning model can be constructed. Based on the 3D deep model, a 3D image of the current scene and contextual information are input in real time to achieve end-to-end drowning identification. The contextual information includes time-series features of behavioral patterns such as normal swimming, abnormal swimming, or intermittent water play. The identification threshold is dynamically adjusted in conjunction with parameters such as ambient light, wind speed, and waves. Specifically:
[0058] First, multimodal training data was constructed. Based on video / real-scene images collected by multi-view visual monitoring station 1 (N≥3), a 3D modeling image of the drowning scene was generated using a multi-view dense reconstruction algorithm of a 3D depth model. M≥500 reconstructed images from different perspectives were generated for each scene, including the 3D pose coordinates of the human body. Synchronously collect contextual information, which includes: 1. Temporal sequence containing human pose. 2. Environmental parameters ,in For light intensity, For wind speed and 3. Wave height; 4. Status of monitoring station equipment ,in For frame rate, To achieve a certain resolution, a labeled dataset was constructed, with the following labeling categories: drowning, normal swimming, and obstacles. The dataset size is... sample;
[0059] Then, we conducted deep learning large-scale model training to build a fusion network: using a CNN (Convolutional Neural Network) module to extract spatial features from the stereo modeling images. The temporal features of the pose time series are learned using an LSTM (Long Short-Term Memory) network module. Multimodal features are obtained by fusing spatial, temporal, and environmental features using the Transformer module. The confidence level is identified by the output of the fully connected layer. The training process uses the cross-entropy loss function. , For the true labels of the samples, Iterative training for batch size Round-robin validation set accuracy False positive rate ;
[0060] Finally, model deployment and real-time recognition are achieved by deploying the trained large model to the processing system and inputting a stereoscopic image of the current scene in real time based on the 3D depth model. With context information The model outputs recognition results end-to-end. ( The system determines the cause of death as drowning by integrating redundant confidence scores from multiple stations, thus arriving at a final decision. ( (As model weights), ensuring low recognition latency and adapting to the second-level response requirements of rescue operations.
[0061] The aforementioned multi-station redundancy confidence level improves identification reliability through weighted voting, avoiding false positives or false negatives caused by single-station failures. Specifically:
[0062] First, identification results from multiple stations were collected, and the confidence level for drowning identification was calculated at each monitoring station. And report it to the processing system;
[0063] Secondly, calculate the inter-station weights. Image sharpness from the monitoring station to the drowning point is determined by calculating the gradient magnitude. (W, H are the image width and height, (where the gradient is in the x / y direction), the weight formula is: ;
[0064] Finally, the fusion confidence score. If the rescue is triggered in time, a manual review will be initiated otherwise.
[0065] After determining the coordinates of the drowning point, lifeboat 2 is activated. The processing system uses a multi-view vision system to track and locate the drowning victim and lifeboat 2 in real time. During the tracking and positioning process, the extrinsic parameters of the multi-view system need to be automatically corrected, eliminating the need for manual calibration. Specifically:
[0066] First, design the markers and extract the image coordinates, then fix them on the deck of lifeboat 2. Each marker point has its world coordinates relative to the center of lifeboat 2. It is known that The markers include LED lights with a flashing frequency of [frequency missing]. Each monitoring station uses a "color threshold" + Frequency matching extracts the image coordinates of marker points ;
[0067] Secondly, the extrinsic parameter correction values are calculated based on the correspondence between the "known world coordinates - image coordinates" of the marker points, and the extrinsic parameters of monitoring station k are corrected. Satisfying the perspective projection model
[0068]
[0069] in, As a scale factor, As camera intrinsic parameters, the reprojection error is constructed using the L2 norm.
[0070] ,
[0071] This optimization problem is optimized using an optimization algorithm, with the error target being [value missing]. ;
[0072] Finally, the calibration process is triggered, and during regular calibration, it is performed periodically. Triggered, lifeboat 2 cruises to the area of overlap with the monitoring station's field of view; during emergency calibration, if the IMU detects a change in attitude... Single-station calibration is triggered immediately, taking [time]. And it does not affect the rescue of drowning people.
[0073] As a preferred approach, addressing the issue of the time lag between the launch and arrival of lifeboat 2, which causes the drowning location to shift, the identification module can predict the location of the drowning person. By analyzing historical location data of the drowning person, it predicts the real-time location and guides lifeboat 2 to the destination accurately, reducing search time. Specifically:
[0074] First, the movement speed of drowning victims is fitted based on historical data. Frame location data (time interval) Fit the velocity vector using the least squares method. The formula is , Similarly;
[0075] Secondly, construct a Kalman prediction model and define the state vector. The prediction equation is
[0076]
[0077] in, Estimated arrival time of lifeboat 2;
[0078] Finally, the prediction error is evaluated. ,in, For velocity fitting error, To ensure positioning error .
[0079] Finally, the optimal rescue path is planned by combining the three-dimensional depth model, and the rescue boat 2 is directed to accurately reach the designated location to carry out the rescue. Specifically, the location of obstacles is detected by multi-view vision, a grid map is constructed, and a collision-free path is planned based on the cost function of the improved A* algorithm to ensure that the rescue boat 2 arrives at the drowning point safely and quickly.
[0080] First, obstacles are detected and a map is constructed. The multi-view vision system extracts the obstacle image coordinates through image segmentation and converts them into world coordinates. Build a raster map, raster size Obstacle grids are marked as 1, and feasible grids are marked as 0.
[0081] Secondly, we design an improved A* cost function, the cost function being... Where g(n) is the actual distance from the starting point to grid n, and h(n) is the Euclidean distance from grid n to the drowning point. ), o(n) is the obstacle penalty term (o(n) = 1 if the grid n is a distance from the obstacle). Otherwise, o(n) = 0. This is the penalty coefficient;
[0082] Finally, the search and smoothing process involves finding the minimum-cost path from the starting point to the drowning point and smoothing the path using a B-spline curve to ensure the turning angle of lifeboat 2. Path planning time .
[0083] Preferably, to address the issue of a sharp drop in positioning accuracy for lifeboat 2 in GNSS obstruction scenarios, this embodiment employs multi-view vision and GNSS positioning collaborative positioning for lifeboat 2. It integrates multi-view marker measurements, target visual measurements, and GNSS data, and uses extended Kalman filtering to achieve accurate positioning across all scenarios, ensuring the continuity of lifeboat 2's navigation.
[0084] First, construct a multi-source measurement system, retaining "visual target measurement". (Multi-station visual inspection coordinates, accuracy) GNSS Measurement (Submarine-borne GNSS coordinates, accuracy) The new "Image Identification Measurement" option has been added. (The coordinates of lifeboat 2 were calculated based on monocular ranging.) To indicate the actual size, To indicate dimensions and precision in the image. The second step is to design the filtering model, including the state vector. The prediction phase is based on motion models.
[0085]
[0086] in, For the filter period, For process noise, Variance; Update phase measurement vector Measurement matrix Noise measurement (Weights are allocated according to precision)
[0087] Finally, the measurement selection was performed, and GNSS was normal. ) time fusion GNSS obstruction ( ) time fusion When the identifier is not visible, it degenerates into... .
[0088] The present invention also provides a multi-view vision rescue system, including a lifeboat 2. Several fixed markers are provided on the surface of the lifeboat 2 to provide positioning references in multi-view vision monitoring. Preferably, the fixed markers include LED lights. The lifeboat 2 consists of a hull, a power system, a lifeboat 2 communication module, a positioning module, a control system, and a wireless charging module. The power system can be a propeller propulsion, water jet propulsion, ducted propulsion, or other power systems. Through the lifeboat 2 communication module, positioning module, and control module, the system can receive and process trajectory data for autonomous control.
[0089] In this embodiment, the lifeboat 2 is launched using an ejection device, which includes a wireless charging module, a solar panel, a battery, and an ejection mechanism. The lifeboat 2 is mounted on the ejection device. Preferably, multiple lifeboats 2 can be loaded simultaneously in the ejection device. The ejection mechanism can be a spring ejection mechanism, an electromagnetic ejection mechanism, or a guide rail slider mechanism, thereby enabling the lifeboat 2 to be launched by means of spring ejection, electromagnetic ejection, guide rail slider, etc. The solar panel supplies power to the battery, and the battery provides power to each component. When the lifeboat 2 is loaded on the ejection device, it is charged using the wireless charging module.
[0090] The multi-view vision system includes several vision monitoring stations 1. Each vision monitoring station 1 includes a battery module, a camera, a positioning module, an inertial navigation module, a data acquisition module, and a monitoring and communication module. The camera is a binocular camera, and the camera can be a visible light, infrared, or a combination of both. The battery module also consists of a solar panel and a battery, which powers the camera, positioning module, inertial navigation module, data acquisition module, and monitoring and communication module.
[0091] The visual monitoring stations 1 are arranged in a ring around the monitored water area, so that the cameras of each visual monitoring station 1 can form a multi-angle collaborative shooting to achieve full coverage monitoring of the rescue water area and obtain image information of the rescue water area in real time through the cameras; the data acquisition module is used to collect spatial information of the camera obtained by the positioning module and the inertial navigation module in real time. This spatial information includes spatial position information such as the coordinates of the camera, the camera angle, the relative attitude, and the IMU attitude angle.
[0092] The processing system consists of server 3, recognition software, control software, communication software, and a system communication module. It is used to receive and process image information and spatial information, calculate the position data of each target in the environment using multi-view vision measurement methods, and identify drowning people in the water using stereo perception algorithms. When a drowning person is detected, a command is issued to control the lifeboat 2 to start. The position of the lifeboat 2 is calculated using multi-view vision measurement algorithms, and the autonomous navigation algorithm is used to control the lifeboat 2 to navigate to the drowning site to carry out the rescue.
[0093] The communication system of the present invention, which is composed of the above-mentioned lifeboat 2 communication module, monitoring communication module and system communication module, can form a wired or wireless communication network to realize real-time two-way communication between the multi-view vision system, the processing system and the lifeboat 2.
[0094] This invention enables stereoscopic modeling of the environment after calibration using a multi-view vision system, and then uses the established stereoscopic model to detect drowning victims, thereby improving detection efficiency.
[0095] The processing system acquires data from each visual monitoring station 1 through the communication system, calculates the position and orientation of the cameras at each visual monitoring station 1, and then uses multi-view vision measurement methods to calculate the position data of each target in the environment. At the same time, it uses stereo perception algorithms to identify the various states of targets in the water area, thereby detecting drowning persons.
[0096] When a drowning person is detected, a command is sent to the ejection device via the communication system, and the ejection device launches lifeboat 2.
[0097] After detecting lifeboat 2, the processing system calculates its position using a multi-view vision measurement algorithm and simultaneously uses an autonomous navigation algorithm to guide lifeboat 2 to the drowning site for rescue, thereby improving the rescue system. This invention constructs an all-weather, three-dimensional detection, identification, positioning, and rescue system, providing effective protection for water rescue.
[0098] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-vision rescue method, characterized in that, include: Image information of the monitored water area is acquired from different directions using a multi-view vision system; A three-dimensional depth model of the entire water area is constructed based on image information, and drowning identification is performed to determine whether there are people who have fallen into the water. When it is determined that there are people who have fallen into the water, their location is determined. The lifeboat (2) is activated based on the location of the person who fell into the water, and the person and the lifeboat (2) are tracked and located in real time through a multi-view vision system. Combine the three-dimensional depth model to plan the optimal rescue route and direct the lifeboat (2) to accurately reach the designated location to carry out the rescue.
2. The multi-view vision-based lifesaving method according to claim 1, characterized in that, The image information is subjected to reflection suppression preprocessing, specifically: each image information is divided into blocks, and the average gray value of each block is calculated. When the average gray value of a block is greater than a threshold, it is judged as a reflection block. Then, the average gray value of the reflection block is compressed so that the average gray value of the compressed reflection block is less than or equal to the threshold. Finally, Gaussian filtering is performed to eliminate noise interference.
3. The multi-vision rescue method according to claim 1, characterized in that, During the tracking and positioning process, it is necessary to automatically correct the extrinsic parameters of the multi-view system. Specifically, the fixed markings on the surface of the lifeboat (2) are used as dynamic calibration objects. The coordinates are extracted by capturing the marking images and combined with the known relative positions of the marking points to complete the automatic correction of the extrinsic parameters of the multi-view system.
4. The multi-view vision-based lifesaving method according to claim 1, characterized in that, The three-dimensional depth model is constructed from images simultaneously collected by no fewer than three monitoring stations and based on the principle of multi-view set. It achieves spatial separation of effective water area, water background and obstacles, forming environmental constraints for tracking and positioning.
5. A multi-view vision-based lifesaving method according to claim 4, characterized in that, When identifying whether a person has fallen into the water, it is necessary to identify the target's posture and motion features. By fusing the human torso posture angle and limb movement frequency as dual features, a recognition confidence model is constructed. When the output value of the confidence model exceeds a set threshold, it is determined to be a person who has fallen into the water, triggering an alarm and initiating a rescue process. At the same time, the drowning point is located, and the coordinates of the detection station of the multi-view vision system are obtained. Combining the detection station coordinates with the constraints of the 3D depth model, non-water background targets are first filtered out, and then the world coordinates of the drowning point are optimized through weighted fusion.
6. A multi-view vision-based lifesaving method according to claim 1, characterized in that, The real-time location of the drowning person is predicted based on the historical location data, and the lifeboat (2) is guided to carry out the rescue based on the latest predicted real-time location.
7. A multi-view vision-based lifesaving method according to claim 1, characterized in that, Multi-view vision and GNSS positioning are used to coordinate positioning of the lifeboat (2). Multi-view vision identification measurement, target vision measurement and GNSS data are integrated. Through extended Kalman filtering, accurate positioning in the whole scene is achieved to ensure the navigation continuity of the lifeboat (2).
8. A multi-view vision-based lifesaving method according to claim 7, characterized in that, The optimal rescue route planning includes: detecting the location of obstacles through multi-view vision, constructing a grid map, and planning a collision-free route based on the cost function of the improved A* algorithm, so that the lifeboat (2) can safely and quickly reach the drowning point.
9. A multi-view vision-based lifesaving method according to claim 4, characterized in that, A large deep learning model is constructed, and the current scene's stereoscopic image and contextual information are input in real time based on the 3D deep model to achieve end-to-end drowning recognition. The aforementioned contextual information includes normal swimming, abnormal swimming, or intermittent water play.
10. A multi-view vision lifesaving system, characterized in that, Lifeboat (2), with several fixed markings on its surface for providing positioning references in multi-view visual monitoring; A multi-view vision system, comprising several vision monitoring stations (1), each of which includes a battery module, a camera, a positioning module, an inertial navigation module, and a data acquisition module. The vision monitoring stations (1) are arranged in a ring around the monitored water area, so that the cameras of each vision monitoring station (1) can form a multi-angle collaborative shooting to achieve full coverage monitoring of the rescue water area, and obtain image information of the rescue water area in real time through the cameras; the data acquisition module is used to collect spatial information obtained by the positioning module and the inertial navigation module in real time. The processing system is used to receive and process image information and spatial information, calculate the position data of each target in the environment using multi-view vision measurement method, and identify drowning people in the water area using stereo perception algorithm. When a drowning person is detected, the system issues a command to control the lifeboat (2) to start, calculates the position of the lifeboat (2) using multi-view vision measurement algorithm, and controls the lifeboat (2) to navigate to the drowning location to carry out the rescue using autonomous navigation algorithm. The communication system is used to realize real-time two-way communication between the multi-view vision system, the processing system and the lifeboat (2).
Citation Information
Patent Citations
Water rescue system and method
CN110576951A
Anti-drowning intelligent monitoring and rescue method
CN118323395A
Diver distress signal response type automatic cruise positioning lifeboat system and method
CN120370951A
Drowning detection method and device and storage medium
CN120997765A
Water rescue method and system based on radar camera
CN121133954A