Multi-domain collaborative ocean emergency rescue search and rescue method

By using drones, unmanned vessels, and underwater robots in collaborative operations, combined with a multimodal drowning scoring mechanism and high-precision sensors, the problems of short sailing time and low identification accuracy in marine search and rescue have been solved. This has enabled efficient and intelligent multi-domain collaborative search and rescue, improving the efficiency and accuracy of deep-sea rescue operations.

CN120942522APending Publication Date: 2025-11-14HARBIN INST OF TECH AT WEIHAI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511299205.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing drones, unmanned vessels, and underwater robots have problems in marine search and rescue, such as short sailing time, small operating range, and low accuracy in identifying victims, especially in the open ocean, deep sea, and complex waters.

Method used

A multi-domain collaborative marine emergency rescue approach is adopted, utilizing drones, unmanned vessels, and underwater robots to work together. Drones conduct reconnaissance, unmanned vessels transport underwater robots, and underwater robots collaboratively identify and deploy life-saving devices. Combined with a multimodal drowning scoring mechanism and high-precision sensor fusion, accurate drowning identification and rescue are achieved.

Benefits of technology

It has expanded the search and rescue scope, improved the rescue effect, realized efficient and intelligent search and rescue in complex waters, enhanced the motion control accuracy and autonomous operation capability of underwater robots, and improved the intelligence level and efficiency of search and rescue missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120942522A_ABST
    Figure CN120942522A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-domain collaborative ocean emergency rescue method, which solves the technical problems of short endurance, small operation range and low recognition accuracy of persons in distress when an unmanned aerial vehicle or an unmanned ship or an underwater robot is adopted to carry out a search and rescue task in the prior art, and combines the unmanned aerial vehicle, the unmanned ship and the underwater robot for mutual cooperation. The unmanned aerial vehicle carries out investigation traversal in the air to find a search and rescue target, the unmanned aerial vehicle gets close to the search and rescue target and sends the position information of the target to the control center, the unmanned ship sails to the position of the target, the at least two underwater robots are transferred along with the unmanned ship, and the unmanned ship controls the at least two underwater robots to work cooperatively. The method can be widely applied to deep ocean emergency rescue, and an'air-sea-submerged 'multi-domain cooperative new normal form is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of search and rescue robot technology, and more specifically, to a multi-domain collaborative marine emergency rescue and search method. Background Technology

[0002] With the increasing frequency of natural disasters, emergencies, and extreme weather events, the demand for emergency search and rescue in marine and aquatic environments is growing. Especially in emergencies such as floods, maritime accidents, and underwater collapses, traditional manual search and rescue methods suffer from slow response, high risk, and low efficiency, necessitating more intelligent, efficient, and unmanned technologies to improve emergency response and rescue capabilities.

[0003] Currently, the main search and rescue equipment includes three categories: drones, unmanned surface vessels (USVs), and underwater robots. Marine drones face the challenge of short endurance. Underwater robots suffer from operational difficulties; shore-based control of underwater robots limits their operating range to a small extent based on cable length, necessitating improved motion accuracy to facilitate search and rescue operations. The accuracy and intelligence of identifying missing persons also need improvement. These issues are particularly prominent in emergency rescue operations in the open ocean, deep sea, and complex waters. Summary of the Invention

[0004] This application aims to address the technical problems of short flight time, small operating range, and low accuracy in identifying victims in existing technologies that use drones, unmanned vessels, or underwater robots for search and rescue missions. It provides a multi-domain collaborative marine emergency rescue and search method with a large operating range and improved rescue effectiveness.

[0005] This invention provides a multi-domain collaborative marine emergency rescue and search method, including a collaborative emergency search and rescue system, which includes unmanned aerial vehicles (UAVs), unmanned vessels, underwater robots, and a control center.

[0006] The multi-domain collaborative marine emergency rescue and search method includes the following steps:

[0007] The first step is for the drone to conduct reconnaissance and search in the air to find the search and rescue target. The drone then approaches the search and rescue target and sends its own position information or the position information obtained by the drone from locating the target to the control center.

[0008] The second step is for the unmanned vessel to sail to the target location, where at least two underwater robots are located in the gap between the unmanned vessel's catamaran structure, and the underwater robots move with the unmanned vessel.

[0009] The third step involves at least two underwater robots working together.

[0010] Step 1: The unmanned vessel releases the underwater robot, which then navigates in the water.

[0011] Step 2: The first underwater robot performs a traversal search, finds the target, and sends the target's location coordinates to the second underwater robot.

[0012] Step 3: The second underwater robot travels to the target location and then performs target identification.

[0013] Step S301: The underwater binocular camera mounted on the second underwater robot collects video of the target;

[0014] Step S302: The video is decomposed into continuous static image frames at a fixed frame rate, and then each frame image is input into the human key point fitting detection model to obtain the detection result image and identify the human key points.

[0015] Step S303: Calculate the indicators based on the key points;

[0016] Step S304: Determine whether the target is drowning based on the indicators;

[0017] Step 4: After confirming drowning, the second underwater robot sails close to the target and releases the lifebuoy via a towing device or deploys the life-saving device via an automatic delivery device.

[0018] Step 5: After the lifebuoy is released by the towing device in step 4, the second underwater robot travels towards the shore and drags the target to the shore.

[0019] Preferably, the unmanned vessel has a catamaran structure, and the hull of the unmanned vessel is equipped with a drone landing platform. The drone landing platform is connected to a wireless charging device, and the gap in the middle of the unmanned vessel is used as an underwater robot recovery dock; Step 1, the unmanned vessel controls the underwater robot through cables.

[0020] Preferably, the key points of the human body obtained in step S302 include the left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, left eye, right eye, nose, left ear, right ear, and mouth. The left eye includes the key points at the left corner of the eye, the right corner of the eye, the center key point of the upper eyelid, and the center key point of the lower eyelid. The right eye includes the key points at the left corner of the eye, the right corner of the eye, the center key point of the upper eyelid, and the center key point of the lower eyelid. The mouth includes the key points at the left corner of the mouth, the right corner of the mouth, the center key point of the upper lip, and the center key point of the lower lip.

[0021] Step S303: Calculate one or more of the following indicators;

[0022] (1) The angle between the elbow and the shoulder joint is θ, θ = arccos(a·b / |a||b|), where a is the line segment between the right shoulder point and the right wrist point, and b is the line segment between the right elbow point and the right wrist point;

[0023] (2) The dominant frequency of the arm swing is fdominant, and the two-dimensional image coordinates of the key point of the right wrist in frame t are pt=(x t ,y t ), calculate the displacement sequence of the right wrist key points along the y-axis in N consecutive frames, forming a discrete signal S = {y1, y2, ..., y N}, where y N This represents the absolute y-axis coordinate of the wrist keypoint in the t-th frame image. The displacement sequence signal is then preprocessed to obtain signal S. filtered Then the filtered signal S filtered Perform a Fast Fourier Transform to convert it from the time domain to the frequency domain:

[0024]

[0025] Then, its power spectral density (PSD) distribution is obtained, and the frequency of the maximum energy peak in the power spectral density (PSD) distribution is defined as the dominant frequency fdominant of the arm swing.

[0026] (3) Quantification of cross-correlation coefficient: Where L and R represent the joint angle θ sequences of the left and right wrists over a period of time, respectively, Cov(L,R) measures the synergy of their changes, and σ L and σ R Then it is the standard deviation of each sequence;

[0027] (4) Aspect ratio of the eye region r eye r eye =h eye / w eye h eye The distance is defined as: the vertical distance between the center key points of the upper and lower eyelids of the left eye, and the vertical distance between the center key points of the upper and lower eyelids of the right eye; the average of the two vertical distances is taken. eye The horizontal distance between the key points at the left and right corners of the eye;

[0028] (5) Opening coefficient k mouth k mouth =h mouth / w mouth Width w mouth To determine the horizontal distance between the key points at the left and right corners of the mouth, the height h is... mouth The vertical distance between the center key point of the upper lip and the center key point of the lower lip;

[0029] In step S304, one of the following methods is selected to determine whether the target is drowning:

[0030] The first determination method is: if the angle θ between the elbow and shoulder joint is greater than the set threshold, and the angle θ for 30 consecutive frames is greater than the set threshold, then it is determined to be a drowning state.

[0031] The second method of judgment: if the dominant frequency of the arm swing is in the range of 3 to 5 Hz, it is judged as drowning;

[0032] The third method of determination: if ρ LR If the value is approximately 0, then it is determined to be drowning;

[0033] The fourth method of determination: the aspect ratio r of the eye region. eye If the value exceeds the threshold calculated based on the baseline value and persists for more than 5 frames, it is determined to be drowning.

[0034] The fifth determination method: opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning.

[0035] The sixth method of determination: Count the total time T during which the key nasal point is below the water surface within T=10s. submerged and the longest single duration T max If T submerged >6s or T max If the time exceeds 3 seconds, it is considered drowning;

[0036] The seventh method of judgment: Calculate the frequency of bubbles appearing near the nose and mouth. If the frequency of bubbles appears continuously for more than 2 Hz and lasts for more than 3 seconds, it is judged as drowning.

[0037] The ⑧th determination method: based on the dominant frequency of arm swing, fdominant, ρ LR The aspect ratio of the eye area is r eye Opening coefficient k mouth Total duration T submerged Longest single duration T max Multiple indicators, such as the frequency of bubbles appearing near the nose and mouth, are selected and input into a multivariate time series prediction model. The multivariate time series prediction model outputs a score, and then drowning is determined based on the score.

[0038] The 9th determination method: based on the dominant frequency of arm swing, fdominant, ρ LR The aspect ratio of the eye area is r eye Opening coefficient k mouth Multiple indicators are selected and input into a multivariate time series prediction model. The multivariate time series prediction model outputs a score, and then drowning is determined based on the score.

[0039] The 10th determination method: the aspect ratio r of the eye region eye The value is greater than the threshold and lasts for more than 5 frames, and the opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning.

[0040] Preferably, in step 3, the human body key point fitting and detection model is the YOLOv8 target detection model.

[0041] Preferably, the process of creating the human body key point fitting and detection model is as follows:

[0042] Step (1): Establish a dataset of key point detection images of drowning victims;

[0043] Step (2): Use the drowning crowd key point detection image dataset established in step (1) to train the YOLOv8 target detection model and obtain the human body key point fitting detection model.

[0044] The present invention also provides a multi-domain collaborative marine emergency rescue and search method, including a collaborative emergency rescue system, wherein the collaborative emergency rescue system includes unmanned aerial vehicles, unmanned vessels, underwater robots, and a control center;

[0045] The multi-domain collaborative marine emergency rescue and search method includes the following steps:

[0046] The first step is for the drone to conduct reconnaissance and search in the air to find the search and rescue target. The drone then approaches the search and rescue target and sends its own position information or the position information obtained by the drone from locating the target to the control center.

[0047] The second step is for the unmanned vessel to sail to the target location, where at least two underwater robots are located in the gap between the unmanned vessel's catamaran structure, and the underwater robots move with the unmanned vessel.

[0048] The third step involves at least two underwater robots working together.

[0049] Step 1: The unmanned vessel releases the underwater robot, which then navigates in the water.

[0050] Step 2: The first underwater robot performs a traversal search, finds the target, and sends the target's location coordinates to the second underwater robot.

[0051] Step 3: The second underwater robot travels to the target location and then performs target identification.

[0052] Step S301: The underwater binocular camera mounted on the second underwater robot collects video of the target;

[0053] Step S302: The video is decomposed into continuous static image frames at a fixed frame rate, and then each frame image is input into the human key point fitting detection model to obtain the detection result image and identify the human key points.

[0054] Step S303: Calculate the indicators based on the key points;

[0055] Step S304: Determine whether the target is drowning based on the indicators;

[0056] Step 4: The drone is positioned at the precise location coordinates provided by the first underwater robot, and then the drone deploys the rescue device.

[0057] The beneficial effects of this disclosure are: combining the advantages of drones, unmanned vessels, and underwater robots to achieve a comprehensive and collaborative search and rescue approach. It allows for a large search and rescue area and enables long-term search and rescue missions, thus improving search and rescue effectiveness.

[0058] It can be used not only for target identification, environmental detection, and life search and rescue at disaster sites, enabling collaborative operation and rapid deployment of multiple devices, but also greatly improving rescue efficiency and safety in complex aquatic environments. It is also suitable for emergency rescue in the open ocean and deep sea, forming a new paradigm of multi-domain collaboration encompassing air, sea, and underwater.

[0059] The underwater robot's motion control is more precise and efficient, achieving highly reliable obstacle avoidance and improving its autonomous operation capabilities, which is beneficial for precise and efficient search and rescue operations.

[0060] By combining human key point detection with behavioral semantic understanding, this approach breaks through the limitations of traditional visual algorithms that rely solely on appearance features. It proposes a multimodal drowning scoring mechanism that integrates spatiotemporal features and motion context, enhancing the algorithm's robustness in complex underwater environments. It achieves full autonomy from identification and localization to towing and rescue, significantly improving the intelligence level and execution efficiency of complex underwater search and rescue tasks.

[0061] Further features and aspects of this disclosure will be clearly described in the following detailed description with reference to the accompanying drawings. Attached Figure Description

[0062] Figure 1 This is a schematic diagram of the collaborative emergency search and rescue system.

[0063] Figure 2 This is a schematic diagram of the various modules configured in the drone;

[0064] Figure 3 This is a structural diagram of an unmanned vessel;

[0065] Figure 4 It is a flowchart for calculating quantitative indicators based on key points of the human body to further assess the condition.

[0066] Figure 5 It is a flowchart that calculates quantitative indicators based on key points of the human body to further determine whether drowning has occurred.

[0067] Figure 6 This is a diagram illustrating the angle between the elbow and shoulder joints;

[0068] Figure 7 yes Figure 6 Partial view of the right elbow point, right wrist point, and right shoulder point;

[0069] Figure 8 This is a diagram illustrating the height difference between the wrist and shoulder.

[0070] Figure 9 This is a schematic diagram illustrating the calculation of the eye's aspect ratio.

[0071] Figure 10 This is the architecture diagram of the improved YOLOv8 object detection model.

[0072] Explanation of symbols in the diagram:

[0073] 1. Unmanned vessel, 1-1. Unmanned aerial vehicle landing platform, 1-2. Gap, 2. Unmanned aerial vehicle. Detailed Implementation

[0074] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] The specific embodiments described below are merely preferred embodiments of this application, and the scope of protection of this application is not limited thereto. Those skilled in the art can make modifications or variations based on the principles, concepts, and spirit of this application, and the resulting technical solutions should all be covered within the scope of protection of this application.

[0076] The drone is equipped with: an M10N GPS module, a flight controller, an onboard depth camera, an onboard millimeter-wave radar, an inertial measurement unit (IMU), a wireless charging unit, and an onboard delivery device. The M10N GPS module ensures accurate positioning capabilities in complex environments. The onboard depth camera can be used for visual guidance for precise landing. The drone's fuselage can utilize an X-shaped quadcopter frame. This design not only enhances structural stability but also provides wave and wind resistance, enabling stable operation in harsh marine environments. This design is particularly important for search and rescue operations in vast ocean areas, as it ensures the drone remains stable even when facing waves and strong winds, thereby improving search and rescue efficiency and success rates.

[0077] Integrating high-precision millimeter-wave radar, a depth camera, and a high-sampling-rate inertial measurement unit (IMU), the system significantly improves accuracy and reliability through deep fusion of multimodal sensor information. Specifically, it employs a multi-sensor fusion framework centered on an extended Kalman filter (EKF). This framework is based on an improvement to the Lidar-Visual-Inertial SLAM (LVI-SAM) algorithm, particularly strengthening the weighting of visual and radar point cloud information in state estimation, thereby achieving high-frequency, high-precision estimation of the six-degree-of-freedom (6-DoF) pose of the UAV. The specific process is as follows:

[0078] Step (1): Obtain the UAV's position observation information in the global coordinate system through the M10N GPS module; the M10N GPS module receives signals from multiple satellites and outputs raw observation data containing latitude, longitude, and altitude information. Let the raw observation provided at time t be:

[0079] zGPS, t=[φ,λ,h] T

[0080] In the formula, φ represents latitude, λ represents longitude, and h represents altitude. Through coordinate transformation, it is converted to a three-dimensional position pGPS = [X,Y,Z]^T in the Earth-centered Earth-fixed (ECEF) coordinate system. The transformation process can be expressed as:

[0081] P GPS =T(φ,λ,h)

[0082] In the above equation, T is the transformation function from the geodetic coordinate system to ECEF. This observation will be used as a global position reference input to the fusion filter.

[0083] Step (2): Obtain the measured values ​​of angular velocity ω and acceleration a through an inertial measurement unit (IMU); the IMU outputs the three-axis angular velocity ω in the body coordinate system at a high frequency (usually ≥200Hz). m and triaxial acceleration a m Its measurement model can be expressed as:

[0084] ω m =ω t +b ω +n ω

[0085] a m =a t +b a +R(q) T g+n a

[0086] Where, ω t and a t For the actual angular velocity and acceleration, bω and b a For the zero bias of the gyroscope and accelerometer, n ω and n a To measure white noise, R(q) is the rotation matrix corresponding to the current attitude q, and g = [0,0,g]^T is the gravity vector in the global coordinate system. These raw measurements will be used for state prediction and zero-bias estimation.

[0087] Step (3) involves obtaining visual-radar joint relative pose observation information using the improved LVI-SAM algorithm; and analyzing the image sequence I acquired by the airborne depth camera. k Extract multi-scale ORB feature points to obtain the feature point set: F(V) = fv, i∈R 2 ×S.

[0088] Obtain the descriptor.

[0089] Feature extraction based on geometric distribution is performed on the point cloud Pr acquired by 3D radar to obtain structural features such as edge points ε and planar points S.

[0090] The core improvement lies in designing an adaptive weight allocation mechanism based on information entropy. First, the inter-frame matching success rate p of visual features is calculated separately. v The overlap ratio p between the radar point cloud and the radar point cloud r :

[0091]

[0092] Where d is the matching distance of the points in the point cloud.

[0093] Furthermore, according to p v and p r Instantaneous weights w for computer vision and radar v w r :

[0094] w r =1-w v

[0095] In the formula, ηv and η r The prior confidence coefficients are calibrated offline based on sensor characteristics. Finally, the relative pose transformation is solved by optimizing the joint cost function using weighted least squares:

[0096]

[0097] Where Cv and Cr are the visual reprojection error and radar point cloud matching error, respectively. This relative pose observation T_{k-1,k} will serve as an important measurement input for EKF.

[0098] Step (4) is to achieve multi-sensor fusion pose estimation by extended Kalman filter (EKF).

[0099] Step 1), State prediction, based on angular velocity ω measured by IMU. m and acceleration a m The state is recursively derived using a nonlinear dynamic model. The system state vector is defined as x = [p, v, q, bω, b...]. a ] T .

[0100] The prediction process is based on the following dynamic equations:

[0101]

[0102] in, This represents quaternion multiplication, where n bw and n ba The noise is zero-biased random walk. The Runge-Kutta method (RK4) is used to discretely integrate the above differential equation to obtain the state prediction xk|k-1 and the corresponding prediction covariance Pk|k-1.

[0103] Step 2), Observation Update. The GPS-provided position observation zGPS and the relative pose observation T_{k-1,k} provided by the improved LVI-SAM algorithm are jointly input into the EKF update stage. An adaptive observation noise covariance matrix R is designed:

[0104]

[0105] Where α is the normalization factor and γ is the signal-to-noise ratio threshold. The state estimate xk|k and covariance Pk|k are updated by fusing observation information through Kalman gain calculation.

[0106] Step (5) outputs a high-frequency, high-precision six-degree-of-freedom pose estimation result x = [p, v, q, b]. ω ,ba] T The final output state estimation frequency is consistent with the IMU frequency (≥200Hz). Position p and velocity v are expressed in a global coordinate system, and attitude q is output in quaternion form. IMU error estimates bω and ba are also provided. This result serves as the core state input for the UAV's autonomous flight and mission execution.

[0107] It is evident that this fusion architecture effectively suppresses the drift problem inherent in single GPS positioning, enhances the system's robustness in complex electromagnetic environments and under cover conditions, and provides a reliable pose estimation basis for maritime search and rescue and wide-area reconnaissance missions.

[0108] It achieves high-precision positioning and multi-source satellite map fusion.

[0109] As a key component of search and rescue systems, unmanned surface vessels (USVs) are designed and functioned to improve search and rescue efficiency and success rates. Equipped with high-precision RTK GPS modules, USVs achieve centimeter-level satellite positioning, a technology that provides extremely high accuracy crucial for precise search and rescue operations in vast sea areas. Simultaneously, USVs are also equipped with IMUs (Inertial Measurement Units), LiDAR, and a Realsense435i binocular camera. The combined use of these sensors not only enhances the USV's autonomous navigation capabilities but also provides a relay center for information transmission to unmanned aerial vehicles (UAVs). USVs are also equipped with wheel speedometers.

[0110] The RTK GPS module, combined with IMU and LiDAR, employs existing loosely coupled algorithms—RTK+IMU+wheel velocity measurement—to provide absolute position via GPS, high-frequency attitude via IMU, and wheel velocity measurement assistance. This not only enhances the autonomous navigation capabilities of the unmanned surface vessel (USV) but also provides a relay center for information transmission. The USV's autonomous berthing based on multi-sensor fusion SLAM can accurately complete port berthing tasks, thereby improving the USV's environmental perception accuracy and positioning capabilities. This solves the problems of large satellite positioning deviations and difficulty in acquiring attitude information. It utilizes the real-time depth information perceived by the USV and matches it with prior point cloud maps to obtain accurate pose information.

[0111] The control unit of the unmanned surface vessel (USV) can be equipped with ArduRovert. Simultaneously, the USV is also equipped with an ultra-long-range, high-precision remote control system, which not only improves operational flexibility but also enhances its wave-resistant self-stabilizing control capabilities, ensuring the stability and maneuverability of the USV's landing platform in complex sea conditions.

[0112] The specific structure of the unmanned surface vessel can adopt a catamaran structure, such as... Figure 3 As shown, a drone landing platform 1-1 is installed directly above the hull of the unmanned surface vessel (USV). Wireless charging equipment is installed within the USV landing platform 1-1. This design serves two purposes: providing a temporary storage location for the drones and charging them, ensuring their endurance and enabling them to conduct extended search and rescue operations. This design takes into account the potential power limitations faced by drones during missions. When a drone lands on the USV landing platform 1-1, it can quickly replenish its power via wireless charging without returning to base, thus extending its working time in the mission area. Through wireless charging recharging from the USV, drones can achieve ultra-long-range cruising distances exceeding 1000 km and 8 hours.

[0113] Furthermore, the unmanned surface vessel (USV) is equipped with a differential high-power thruster, enabling it to conduct high-speed and efficient rescue operations. In emergencies, rapid response is crucial, and the high-power thruster significantly enhances the USV's maneuverability, allowing it to quickly reach the rescue site. Simultaneously, the USV is also equipped with an ultra-long-range, high-precision remote control system, which not only improves operational flexibility but also enhances its wave-resistant self-stabilizing control capabilities, ensuring stability and maneuverability in complex sea conditions. Clearly, this significantly improves the efficiency and reliability of the search and rescue system, enabling search and rescue missions to be conducted over a wider area of ​​the sea with higher precision and faster speed, thereby increasing the likelihood of successful rescues.

[0114] The unmanned surface vessel (USV) can adopt a catamaran structure, which not only ensures stability but also provides space for the storage and operation of underwater robots. A large gap is left between the floating materials of the catamaran (this space is equivalent to an underwater robot ROV recovery dock), such as... Figure 3 The gaps 1-2 shown allow the underwater robot to be safely stored and deployed without affecting the stability of the hull. The underwater robot towing and carrying device is installed in the gap. The underwater robot is connected to the towing and carrying device, which can transport the underwater robot to the work area for deployment and retrieval. After retrieval, the underwater robot can be recharged. The charging device adopts a wireless charging solution to ensure the convenience and safety of operation.

[0115] The underwater robot cables mounted on the upper part of the unmanned surface vessel (USV) are key to achieving efficient communication and low-latency control. These cables transmit not only power but also data, enabling the USV to control its underwater robots in real time for precise underwater exploration and operations. This design significantly expands the operational range and flexibility.

[0116] It is evident that unmanned surface vessels (USVs) not only provide a stable and efficient platform for storing and operating underwater robots, but also enable precise control of these robots through advanced communication technologies. This design offers new possibilities for underwater exploration and search and rescue operations, especially in situations requiring long-term, large-scale operations, where its advantages are particularly evident, overcoming the limitations of traditional underwater robots that can only be controlled from shore and operate within a limited range based on cable length.

[0117] As the main execution unit of the search and rescue system, the underwater robot can adopt an eight-motor frame arrangement, which has powerful thrust and flexible control capabilities, enabling the robot to carry out efficient search and rescue operations in complex underwater environments.

[0118] The underwater robot is equipped with an underwater binocular camera, an ultrasonic ranging sensor, and a depth sensor. These sensors provide the underwater robot with comprehensive perception capabilities, enabling it to acquire information about its surrounding environment in real time and providing data support for posture adjustment and path planning. The underwater binocular camera provides stereo vision, which helps the robot to accurately identify and locate targets; the ultrasonic ranging sensor provides accurate distance information, which helps the robot avoid obstacles; and the depth sensor provides accurate depth information, which helps the robot to perform depth control.

[0119] The underwater robot's main control board utilizes a combination of Orange Pi 5PLUS and STM32, which not only provides powerful computing capabilities but also features low power consumption and high reliability. Through power line carrier communication and Mavlink communication, the underwater robot can achieve four movement modes: handle control, anti-current mode, forward roll mode, and bullet-shaped movement mode. These movement modes enable the underwater robot to adapt to more diverse and complex environments, improving the flexibility and efficiency of search and rescue operations.

[0120] Underwater robots can be controlled using sliding mode control algorithms, making their motion control more precise and efficient. Sliding mode control is a nonlinear control method that, through the design of a sliding surface and a sliding mode controller, enables the system to quickly and accurately reach the desired state, thereby improving the robot's maneuverability.

[0121] Furthermore, this system, based on the open-source KIMI or DeepSeek large model architecture, constructs a highly efficient perception and decision-making system for search and rescue missions through domain adaptation and algorithm enhancement. Firstly, in the basic model optimization stage, supervised fine-tuning (SFT) technology is employed to selectively adjust the parameters of the pre-trained model using multimodal data from search and rescue scenarios, including dense smoke, fires, rubble, and water obstacles. This process significantly improves the model's semantic understanding of extreme environments and its ability to identify abnormal states.

[0122] Secondly, to further adapt to embedded platform deployment, model pruning and quantization strategies are introduced for large KIMI or DeepSeek models: pruning removes redundant attention heads and feedforward network parameters, reducing model complexity; quantization converts FP32 weights to FP16 or INT8 format, significantly compressing model size and computational latency while maintaining inference accuracy.

[0123] Then, at the perception level, an ultrasonic-visual cross-modal fusion mechanism is proposed. An underwater binocular camera acquires visual images, while an ultrasonic ranging sensor detects distance data. The algorithm processes the visual images and ultrasonic ranging data separately through a dual-branch structure: the visual branch extracts semantic, edge, and depth features from the environment based on a lightweight CNN or ViT architecture; the ultrasonic branch provides high-precision near-field distance information, playing a crucial role, especially when the optical sensor is occluded or interfered with. In the feature fusion stage, a cross-modal attention mechanism is used to achieve feature alignment and interaction between different modalities, generating a unified Bird's Eye View (BEV) representation, thus maintaining stable perception of obstacle contours and distances even under visual degradation conditions.

[0124] Finally, the decision-making system integrates a low-latency hybrid planning architecture. This engine, based on Model Predictive Control (MPC) and combined with Reinforcement Learning (RL) strategies, predicts dynamic obstacle behavior and constructs risk maps, generating safe and smooth motion trajectories in extremely short cycles. Its core advantage lies in its trade-off between optimality and real-time performance, with an average single inference time controlled within 50 milliseconds and supporting control frequencies above 10Hz, significantly improving the robot's responsiveness and robustness in unknown and dynamic environments.

[0125] In summary, through end-to-end optimization of "pre-trained model fine-tuning - multimodal perception fusion - lightweight real-time decision-making", the search and rescue robot has achieved high-reliability obstacle avoidance and autonomous operation capabilities in extreme environments, demonstrating the effectiveness and scalability of the incremental innovation path of "general model + vertical scenario".

[0126] It is evident that underwater robots possess significant advantages in search and rescue scenarios. They not only have powerful thrust and flexible maneuverability, but also achieve precise and efficient search and rescue operations through advanced sensors and intelligent control algorithms. Their advantages are particularly pronounced in situations requiring long-term, large-scale operations.

[0127] The underwater robot, equipped with a multispectral high-brightness illumination device and an underwater binocular camera, can maintain clear image quality in murky or low-light waters. In the task of identifying drowning victims, the YOLOv8 target detection model is used. The specific process is as follows:

[0128] Step (1): Establish a dataset of key point detection images of drowning people.

[0129] The key technical goal in constructing the image dataset is to build a diverse training dataset that covers complex underwater scenes and solve the problem of real underwater image degradation. The following methods are used to construct the dataset.

[0130] Degradation scenario simulation: Using physical model synthesis technology, underwater color distortion is simulated (red light attenuation, blue-green light dominance: I) R (d)=I R0 ·e -αRd I B (G)=I R0 ·e -αBG The red light attenuation coefficient αRd is taken as 0.2–0.5 m⁻¹, and the blue-green light attenuation coefficient αBG is taken as 0.05–0.1 m⁻¹. -1 ), fog-like blur (scattering effect), point spread function modeled with Gaussian kernel ( σ is taken as 3–7 pixels), and suspended particle interference (Mie scattering: Anisotropy factor g is set to 0.8–0.95, thereby generating a synthetic image that is highly similar to the real underwater environment.

[0131] Occlusion and distortion enhancement: Occlusion is simulated by randomly adding suspended objects such as aquatic plants and rocks, with the occlusion area accounting for 10%–30%, combined with an optical distortion model (radial distortion model: r d =r u (1+k1r u 2 +k2r u 4 (Radial distortion coefficients k1 = -0.3 to 0.3, k2 = 0 to 0.1) generate nonlinear deformed images, improving the model's robustness to partially occluded targets.

[0132] Illumination normalization: Applying multi-scale Gaussian filtering (filter function: Where σ = 5 to 15 pixels, the image with uneven lighting is divided into blocks for normalization to alleviate the problem of local overexposure / underexposure caused by light refraction.

[0133] Adversarial example training: based on FGSM (Fast Gradient Sign Method) Adversarial examples are generated with perturbation strength ε = 0.02 to 0.05, forcing the model to learn more robust feature representations.

[0134] The dataset comprises 20,000 synthetic images and 5,000 real underwater images, annotated with 18 human keypoints (left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, left eye, right eye, nose, left ear, right ear, mouth, left hip, right hip, left knee, right knee, left ankle, and right ankle feature points), providing ample environmental diversity support for model training. The left eye keypoints include the left corner of the eye, right corner of the eye, upper eyelid center keypoint, and lower eyelid center keypoint. The right eye keypoints include the left corner of the eye, right corner of the eye, upper eyelid center keypoint, and lower eyelid center keypoint. The mouth keypoints include the left corner of the mouth, right corner of the mouth, upper lip center keypoint, and lower lip center keypoint.

[0135] Step (2): Train the human body key point fitting detection model.

[0136] The applied model can be either a standard YOLOv8 object detection model or an improved YOLOv8 object detection model.

[0137] To improve detection accuracy and robustness in complex underwater environments, the conventional YOLOv8 model can be architecturally improved and its algorithms optimized. The CSPDarknet53 backbone is retained, but a Spatial Pyramid Pooling Enhancement Module (SPPFF+) is introduced. This enhances the perception of targets at different scales through multi-scale pooling and feature fusion. Four pooling kernels (5×5, 9×9, 13×13, 17×17) are used to fuse multi-scale features, improving the detection capability of distant targets (such as a blurred human body 10 meters away). In the Neck part, the system uses a combination of a Path Aggregation Network (PANet) and Deformable Convolution to enhance feature extraction capabilities for human body contortion poses and partial occlusion. The core innovation of the algorithm lies in the deep fusion of keypoint detection and behavioral semantic understanding. The Neck enhancement design uses a bidirectional PANet path, passing high-level semantics (human contour) from top to bottom and fusing low-level details (limb movements) from bottom to top, resulting in a 12% improvement in keypoint detection mAP. Simultaneously, deformable convolution is used (the formula for calculating the output feature map of deformable convolution is: y(p)=∑ k w k ·x(p+p k +Δp k This paper introduces offset learning (stride 0.1) on top of a 3×3 convolutional kernel to adapt to non-rigid deformations of the human body (such as feature extraction when the arm twists at an angle of 45°). Finally, the Keypoints branch is expanded; this project adds an 18-channel output layer to the Head part of the neural network and uses the CIoU loss function to optimize keypoint localization accuracy. On the test set, the PCK@0.1 (keypoint correct prediction rate) reached 92.3%.

[0138] The improved YOLOv8 target detection model was trained using the drowning population key point detection image dataset established in step (1) to obtain the human body key point fitting detection model.

[0139] The training process uses deep learning to fit the visual features and spatial distribution relationship of human keypoints, ultimately achieving accurate detection and localization of 18 human keypoints in the input image. The entire training process is based on a supervised learning framework, using the aforementioned synthetic and real underwater images and their annotation information as the training sample set {(xi,yi,}_(i=1)^N, where xi is the input image and yi is the ground truth coordinate of the corresponding keypoint. The model optimizes the network parameters θ by minimizing the loss function L(θ)=Σ_i||f_θ(xi)-(yi)||2+λ·R(θ) between predicted and real keypoints, where f_θ is the forward computation process of the model, R(θ) is the regularization term, and λ is the hyperparameter. Gradient descent algorithm is used to iteratively update the parameters during training. Where η is the learning rate. During backpropagation, the gradient of the loss function with respect to the parameters of each layer is calculated using the chain rule. The Adam optimizer is used to adaptively adjust the learning rate of each parameter, accelerating convergence and improving training stability. The model learns the spatial prior distribution and local feature patterns of keypoints through a large number of training samples, ultimately enabling robust estimation of human joint positions under different underwater degradation, occlusion, and lighting conditions, providing reliable structured posture data for subsequent drowning behavior analysis.

[0140] The YOLOv8-Keypoints branch not only detects the human bounding box but also regresses 18 keypoints, including the left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, left eye, right eye, nose, left ear, right ear, mouth, left hip, right hip, left knee, right knee, left ankle, and right ankle. Furthermore, to address underwater optical distortion and suspended object interference, the model incorporates underwater image enhancement and adversarial example training strategies during the training phase, improving generalization performance under turbid and low-light conditions. Ultimately, the system achieves a drowning person detection accuracy exceeding 98%, a false alarm rate below 2%, and an inference speed of 35 FPS on the Jetson AGX Orin platform, meeting the high standards required for real-time underwater rescue.

[0141] Step (3): The underwater binocular camera acquires video and sends the video to the controller for further processing.

[0142] Step (4) involves decomposing the video into continuous static image frames at a fixed frame rate (e.g., 30 FPS). Each frame is then input into the human keypoint fitting detection model to obtain the detection result image, which identifies the key points of the human body. The detection result includes the coordinate data (x, y) of each human keypoint and the corresponding confidence score.

[0143] Step (5) constructs a spatiotemporal behavioral reasoning and analysis mechanism based on these key points: for example, firstly, the relative angles and displacements of each key point are calculated (such as the height difference between the elbow and shoulder, and whether the hands are continuously above the head), and then the movement patterns (such as high-frequency struggling, irregular movements) and facial states (such as whether the mouth and nose are continuously submerged in water, and the bubble overflow pattern). These features are input into a lightweight temporal reasoning network, which outputs a drowning probability score driven by rules and data.

[0144] Specifically, the change in joint angle is obtained by calculating the relative motion characteristics: the angle θ between the elbow and shoulder joints is calculated using the inverse cosine function, such as θ = arccos(a·b / |a||b|). Figure 6 and 7 As shown, line segment 'a' is the line segment between the right shoulder point and the right wrist point, and line segment 'b' is the line segment between the right elbow point and the right wrist point. Calculate the angle θ between line segments 'a' and 'b'. Similarly, the angle between the line segment formed by the left shoulder point and the left wrist point and the line segment formed by the left elbow point and the left wrist point can be calculated. Simultaneously, the swing frequency is calculated by analyzing the movement of wrist key points between consecutive frames. Specifically, refer to... Figure 5 Let the two-dimensional image coordinates of the key points of the wrist (right wrist or left wrist) in frame t be pt = (x t ,y t ), calculate the displacement sequence of the wrist key points in the vertical direction (y-axis) in N consecutive frames (e.g., N = 30 frames, corresponding to approximately 1 second of video data), forming a discrete signal S = {y1, y2, ..., y N}, where y N This represents the absolute y-coordinate of the wrist keypoint in the t-th frame of the image. For example, y1 is the y-coordinate of the wrist keypoint in the first frame, y2 is the y-coordinate of the wrist keypoint in the second frame, and so on. N This is the y-coordinate of the wrist key point in the Nth frame. The displacement sequence signal is then preprocessed (passed through a bandpass filter (typically with a passband of 1-6Hz to remove high-frequency noise and low-frequency drift) to obtain signal S. filtered Then the filtered signal S filtered Perform a Fast Fourier Transform (FFT) to transform it from the time domain to the frequency domain:

[0145]

[0146] Then, its power spectral density (PSD) distribution is obtained.

[0147] The frequency of the maximum energy peak in the power spectral density (PSD) distribution is defined as the dominant frequency fdominant of the arm swing. If the dominant frequency fdominant is in the range of 3 to 5 Hz, it is judged as high-frequency disordered swinging during drowning; if the dominant frequency fdominant is in the range of 1 to 2 Hz, it is judged as normal (such as the coordinated arm stroke during normal swimming).

[0148] It can also calculate the displacement velocity of the extremities, using the optical flow method (Farneback algorithm) to calculate the movement velocity of the wrist or ankle, and the vertical velocity v during drowning. vertical Fluctuations > 0.5 m / s. Finally, temporal feature analysis was used to perform temporal inference on two main parts: judging the degree of struggle of the inspectors and weighted evaluation of facial expressions.

[0149] Alternatively, a drowning scoring mechanism can be used, which is derived by weighted fusion and comprehensive analysis of a series of spatiotemporal features through a multimodal temporal inference network (such as GRU or Transformer). The higher the score, the more closely the current target's behavior matches the drowning pattern. It includes the disorder and high frequency of the struggle mentioned above, as well as the vertical posture and lack of limb coordination of the person directly fitting the results, mainly manifested in the continuous calculation of the height difference Δh between the "wrist" and the "shoulder", Δh=|y w -y s ∣, reference Figure 8 y w It is the Y-coordinate of the right or left wrist point along the Y-axis. s This refers to the Y-coordinate of the right or left shoulder point. A drowning person will instinctively try to extend their hand out of the water, but their legs are too weak to kick, leaving their body in a vertical position. It also allows analysis of whether the movements of the left and right limbs are synchronized, quantified using a cross-correlation coefficient. Where L and R represent the sequence of joint angles θ (or position coordinates) of the left and right wrists over a period of time, respectively, Cov(L,R) measures the synergy of their changes, and σ L and σ R This represents the standard deviation of each sequence, used to standardize the covariance. The calculated ρ... LR The value is used to objectively determine the motion pattern: if ρ LR ≈1 indicates that the arm movements are highly synchronized, corresponding to coordinated strokes in swimming styles such as breaststroke; if ρ LR If ρ ≈ -1, it indicates an alternation of movement patterns, such as the arm stroke in freestyle swimming; both are normal swimming behaviors. Conversely, if ρ LR A value approximately 0 indicates a lack of linear relationship between the left and right limb movements, exhibiting random and chaotic characteristics. This is a typical manifestation of the disordered struggling of the arms during drowning (such as the "climbing a ladder" phenomenon). Therefore, continuous monitoring of ρ... LRWhether the value remains consistently close to zero, combined with characteristics such as vertical posture and high-frequency oscillation, is used to comprehensively assess the risk of drowning. Furthermore, this ρ... LR As an important input to the multimodal temporal inference network, the indicator is used to improve recognition accuracy. At this point, the drowning assessment score in this part increases sharply. ρ LR Evaluation indicators account for approximately 70% of the weight.

[0150] Furthermore, by analyzing the relative positions and movement patterns of key points around the eyes, nose, and mouth, drowning-related facial expressions such as fear and pain can be inferred. A baseline feature profile of the user in a calm state is pre-established. This step involves collecting facial feature data within 10 seconds during the initial phase (usually when the person has just entered the water and is moving steadily), and calculating the eye aspect ratio r during this period. eye and mouth opening coefficient k mouth The aspect ratio of the eye (r) eye The calculation method is as follows: (Refer to) Figure 9 The horizontal distance between the key points at the left and right corners of the eye is L-1. The vertical distance between the center key points of the upper and lower eyelids of the left eye is H-1. The vertical distance between the center key points of the upper and lower eyelids of the right eye is H-2. The average value Hp of H-1 and H-2 is taken, and then the average value Hp is divided by the horizontal distance L-1, which is the aspect ratio r of the eye. eye =Hp / L-1. Then, for all eye aspect ratios r... eye Calculate the average value; this average value is the individual's personalized baseline. For example, the average aspect ratio of someone's eyes in a calm state is r. reyebaseline =0.3, then its abnormal threshold is 0.3 * 1.5 = 0.45; this method avoids individual physiological differences (such as some people being born with larger eyes), improving the accuracy of judgment. Mouth opening coefficient k mouth The horizontal distance between the key points of the left and right corners of the mouth is the width, and the vertical distance between the center key points of the upper and lower lips is the height. The height divided by the width is the mouth opening coefficient k. mouth Then calculate the mouth opening coefficient k for all mouths. mouth The average value is calculated and becomes the individual's personalized baseline value. For the actually detected eye keypoints, the boundaries of the Region of Interest (ROI) are defined using geometric rules based on the coordinates of the eye keypoints and their positions. (Refer to...) Figure 9 The width of the ROI (w eye The horizontal distance between the left and right corner key points is denoted as ROI height (h). eyeCalculate: the vertical distance between the center key points of the upper and lower eyelids of the left eye, and the vertical distance between the center key points of the upper and lower eyelids of the right eye. Take the average of the two vertical distances and further detect the open-eye state by changing the aspect ratio: Define the height h of the eye region. eye With width w eye The ratio r eye =h eye / w eye When a person's eyes widen in fear, the ratio r... eye It will increase significantly; if the current frame's r eye If the value exceeds 1.5 times the baseline value and lasts for more than 5 frames (200ms), it is considered an abnormal eye-opening state.

[0151] For detecting the degree of mouth opening and closing, the height h of the area surrounding key points of the mouth is used. mouth With width w mouth Calculate the opening coefficient k mouth =h mouth / w mouth The region of interest (ROI) around the mouth keypoints is directly defined by the coordinates of the mouth keypoints using geometric rules. The width w of the ROI can be obtained through the mouth keypoints. mouth : Calculate the horizontal distance between the key points at the left and right corners of the mouth. ROI height h mouth : Usually, the vertical distance between the center key point of the upper lip and the center key point of the lower lip is taken. , when k mouth Greater than a preset baseline value, for example, when k mouth A reading >1.8 indicates a significant mouth-opening state, potentially related to struggling breathing during drowning. Simultaneously, the system continuously tracks the positions of key points on the nose and mouth relative to the water surface (obtained through pre-defined water level segmentation). The total duration T for which the key points on the nose and / or mouth are below the water surface within a sliding time window T = 10 seconds is calculated. submerged and the longest single duration T max (where T) submerged It is the cumulative immersion time of the mouth and nose, measuring the overall risk; T max It represents the maximum duration of immersion of both mouth and nose, measuring the most extreme single-event risk; the two are complementary indicators. If T... submerged >6s or T max If the duration is greater than 3 seconds, it is considered an abnormal immersion.

[0152] Furthermore, by detecting continuous small bubble overflows near the nose and mouth area (using inter-frame differencing combined with connected component analysis; specifically, the coordinates of the nose and mouth are again used to define the region of interest (ROI), thus limiting the pixel range for subsequent processing to the area around the nose and mouth; within this ROI, the system calculates the pixel differences between consecutive video frames using inter-frame differencing to extract moving foreground pixel blocks representing newly appearing bubbles; subsequently, the connected component analysis algorithm clusters these foreground pixels into independent connected regions (i.e., candidate bubbles) and filters them based on their area, shape, and other features, ultimately calculating the frequency of bubble appearance; if the bubble appearance frequency consistently exceeds 2Hz and lasts for more than 3 seconds, it is recorded as a drowning support feature).

[0153] The facial expressions and immersion features mentioned above together constitute 30% of the drowning assessment score. After being fused with limb movement features, this data is input into a temporal reasoning network to complete the comprehensive decision. The final comprehensive score is a continuous score between 0 and 50. The comprehensive score f(t) = GRU θ ([F (t-k) ,F (t-k+1) ,...,F t ]), θ is the angle between the elbow and shoulder joints, F t It is the low-level feature vector at time step t, which contains all the low-level features (f) calculated around time t. dominant , ρ LR ,Δh,v vertical k mouth T submerged T max This low-level feature accounts for 70% of the weight, f dominant The dominant frequency of arm swing (3-5Hz is abnormal), ρ LR The left-right limb movement correlation coefficient (close to 0 is abnormal), Δh is the vertical height difference between the wrist and shoulder (consistently large difference is abnormal), v vertical Vertical velocity of the extremities (fluctuations >0.5 m / s are considered abnormal). Facial and immersion features (approximately 30% weight): the ratio of eye aspect ratio to baseline (>1.5 is considered abnormal). mouth Mouth opening / closing coefficient (>1.8 is abnormal). T submerged Total immersion time of mouth and nose (>6s is considered abnormal). maxThe system calculates the following metrics: longest single immersion time of mouth and nose (>3s is abnormal); frequency of bubble overflow (>2Hz is abnormal); and k is the size of the network backtracking time window. The system uses these calculated metrics as feature vectors, directly inputting them into a multivariate time-series prediction model such as GRU / Transformer. This network dynamically weights and fuses these metrics through its internal mechanism, ultimately outputting a score of 0-50. 0-10 points: indicates normal swimming or floating behavior. 11-30 points: abnormal behavior may occur; the system will continue to monitor but will not issue an alarm (possibly due to potential fatigue or cramps). 31-50 points: high-risk level. The system combines the duration of immersion with the score. If the score remains high and exceeds a set threshold (e.g., 5 seconds), a high-confidence drowning alarm is triggered, and subsequent rescue procedures are initiated.

[0154] It should be noted that various combinations of indicators can also be input into the multivariate time series forecasting model according to actual needs.

[0155] The following describes the overall working process of the collaborative emergency search and rescue system:

[0156] The first step involves staff on shore using a drone remote controller to operate the drone, which then conducts reconnaissance at altitudes of 70 meters and 5 meters to locate the search and rescue target. The drone then approaches the target and transmits its own position and attitude information to the control center; this information represents the location of the search and rescue target. This drone's reconnaissance capability is crucial for quickly locating people in distress, as it can provide critical information in the shortest possible time, thus initiating rescue operations.

[0157] It should be noted that the position information of the search and rescue target can also be obtained by using the drone's own pose information instead of the drone's own pose information.

[0158] It should be noted that the drone's own pose information can also be used instead of the target's location information. Instead, the drone can drop a disposable communication buoy over the target. The location of the disposable communication buoy is the location of the target.

[0159] The second step involves staff using a remote control to maneuver the unmanned vessel from a shore base, precisely guiding it to the search and rescue target location. The underwater robot is positioned in the gap between the unmanned vessel's catamaran structure, and it moves along with the vessel.

[0160] A rover station can be set up, communicating with a base station and with the unmanned surface vessel (USV). The rover station can be stationary or in motion. This makes it suitable for missions in the open ocean.

[0161] The third step involves three underwater robots working together.

[0162] Step 1: The towing device releases three underwater robots. The three underwater robots navigate in the water, communicating with the unmanned vessel via cables. The unmanned vessel controls the navigation of the underwater robots.

[0163] Step 2: The first underwater robot is responsible for conducting a full-coverage search of the target water area to locate any suspected victims. Once a suspected victim is detected (specifically through conventional technical means, such as capturing images with a camera and inputting the images into the corresponding YOLO target detection model to identify the human silhouette), the precise coordinates (including latitude, longitude, and depth information) and confidence level are immediately sent to the second underwater robot via the underwater acoustic communication module.

[0164] The first underwater robot can be equipped with multibeam forward-looking sonar and side-scan sonar, enabling it to efficiently acquire underwater terrain and point cloud data of suspected targets. The robot employs grid-based path planning combined with real-time localization and mapping (SLAM) technology to ensure no area is missed in the search.

[0165] It should be noted that an underwater acoustic communication module can be omitted. Instead, the first underwater robot sends its own position information to the unmanned vessel, which then sends the position information to the control center. The control center then sends the position information to the second underwater robot via the unmanned vessel.

[0166] Step 3: The second underwater robot, acting as the rescue execution unit, sails to the precise location coordinates provided by the first underwater robot, approaches the target, and then performs target identification.

[0167] The process of target recognition is as follows:

[0168] Step S301: The underwater binocular camera mounted on the second underwater robot collects video of the target.

[0169] In step S302, the video is decomposed into continuous static image frames at a fixed frame rate (e.g., 30 FPS), and then each frame image is input into the human key point fitting detection model to obtain the detection result image and identify the key points of the human body.

[0170] Step S303: Calculate the following indicators based on the key points: the angle θ between the elbow and shoulder joints, the height difference Δh between the wrist and shoulder.

[0171] Step S304: Determine that the target is in a drowning state.

[0172] The first determination method is: if the angle θ between the elbow and shoulder joint is greater than a set threshold, and the angle θ for 30 consecutive frames is greater than the set threshold, then it is determined to be a drowning state.

[0173] The second method of determination is: if the dominant frequency of the arm swing is in the range of 3 to 5 Hz, then it is determined to be drowning.

[0174] The third method of determination is: quantification through cross-correlation coefficient. Where L and R represent the sequence of joint angles θ (or position coordinates) of the left and right wrists over a period of time, respectively, Cov(L,R) measures the synergy of their changes, and σ L and σ R Then, ρ represents the standard deviation of each sequence, used to standardize the covariance. LR ≈1 indicates normal; if ρ LR ≈-1 indicates normal; if ρ LR If the value is approximately 0, then it is determined to be drowning.

[0175] The fourth method of determination is: the aspect ratio r of the eye region. eye If the value exceeds the threshold calculated based on the baseline value and persists for more than 5 frames, it is determined to be drowning.

[0176] The fifth method of determination is: the opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning.

[0177] The sixth method of determination is: to count the total time T during which the key points of the nose and / or mouth are below the water surface within T=10s. submerged and the longest single duration T max If T submerged >6s or T max If the time exceeds 3 seconds, it is considered drowning.

[0178] The seventh method of determination is to calculate the frequency of bubbles appearing near the nose and mouth. If the frequency of bubbles appears continuously for more than 2 Hz and lasts for more than 3 seconds, it is determined to be drowning.

[0179] The ⑧th method of determination is: based on the angle θ between the elbow and shoulder joints, the dominant frequency of arm swing fdominant, and ρ. LR The aspect ratio of the eye area is r eye Opening coefficient k mouth Total duration T submerged Longest single duration T max Multiple indicators (at least two) are selected from the frequency of bubbles appearing near the nose and mouth and input into the multivariate time series prediction model. The multivariate time series prediction model outputs a score, and then drowning is determined based on the score.

[0180] The ninth method of determination is: the aspect ratio r of the eye region. eye The value is greater than the threshold derived from the baseline and persists for more than 5 frames, and the opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning.

[0181] Step 4: After confirming drowning, the second underwater robot enters the visual servo control phase. Model predictive control (MPC) calculates the trajectory tracking error in real time and dynamically adjusts the thruster output to achieve precise approach in turbulent conditions. Upon approaching the target, the towing device installed on the second underwater robot releases the lifebuoy, which is attached to the towing device. The person in distress can then grab the lifebuoy. Besides the lifebuoy, other lifesaving devices can be used, such as drowning protectors. An automatic deployment device can be mounted on the second underwater robot. After approaching the target, the automatic deployment device deploys lifebuoys, drowning protectors, etc. In this case, the towing device is not used, and the person in distress waits for other vessels to arrive for rescue.

[0182] It should be noted that by equipping the drone with an onboard automatic drop device, and placing the water rescue device inside, the drone can be activated after the target location is confirmed to drop the water rescue device into the water for rescue purposes.

[0183] Step 5: The second underwater robot travels towards the shore and tows the victim to the shore. During the towing operation, a horizontal pulling force of 5-10N can be applied to avoid secondary injury. The system achieves real-time processing at 35FPS on the Orange Pie 5PLUS platform, meeting the stringent requirements of underwater rescue.

[0184] The unmanned boat may move or remain stationary depending on the actual distance.

[0185] Step 6: The third underwater robot serves as a system redundancy backup, always in standby mode. If either the first or second robot malfunctions or is damaged, the third robot immediately takes over its task.

[0186] It is evident that minute-level scanning can be completed in sea areas of tens of thousands of square kilometers, instantly locking onto distressed targets and generating centimeter-level coordinates.

[0187] It should be noted that after completing step 3, steps 4 and 5 are not performed. Instead, the drone is positioned at the precise location coordinates provided by the first underwater robot, and then the drone deploys lifebuoys, water rescue devices, and other life-saving equipment.

Claims

1. A multi-domain collaborative marine emergency rescue and search method, characterized in that, This includes a collaborative emergency search and rescue system, which comprises drones, unmanned vessels, underwater robots, and a control center. The multi-domain collaborative marine emergency rescue and search method includes the following steps: The first step is for the drone to conduct reconnaissance and search in the air to find the search and rescue target. The drone then approaches the search and rescue target and sends its own position information or the position information obtained by the drone from locating the target to the control center. The second step is for the unmanned vessel to sail to the target location, where at least two underwater robots are located in the gap between the unmanned vessel's catamaran structure, and the underwater robots move with the unmanned vessel. The third step involves at least two underwater robots working together. Step 1: The unmanned vessel releases the underwater robot, which then navigates in the water. Step 2: The first underwater robot performs a traversal search, finds the target, and sends the target's location coordinates to the second underwater robot. Step 3: The second underwater robot travels to the target location and then performs target identification. Step S301: The underwater binocular camera mounted on the second underwater robot collects video of the target; Step S302: The video is decomposed into continuous static image frames at a fixed frame rate, and then each frame image is input into the human key point fitting detection model to obtain the detection result image and identify the human key points. Step S303: Calculate the indicators based on the key points; Step S304: Determine whether the target is drowning based on the indicators; Step 4: After confirming drowning, the second underwater robot sails close to the target and releases the lifebuoy via a towing device or deploys the life-saving device via an automatic delivery device. Step 5: After the lifebuoy is released by the towing device in step 4, the second underwater robot travels towards the shore and drags the target to the shore.

2. The multi-domain collaborative marine emergency rescue and search method according to claim 1, characterized in that, The unmanned vessel has a catamaran structure. The hull of the unmanned vessel is equipped with a drone landing platform, which is connected to a wireless charging device. The central gap of the unmanned vessel is used as an underwater robot recovery dock. In step 1, the unmanned vessel controls the underwater robot via cable.

3. The multi-domain collaborative marine emergency rescue and search method according to claim 1, characterized in that, The key points of the human body obtained in step S302 include the left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, left eye, right eye, nose, left ear, right ear, and mouth. The left eye includes the key points at the left corner of the eye, the right corner of the eye, the center key point of the upper eyelid, and the center key point of the lower eyelid. The right eye includes the key points at the left corner of the eye, the right corner of the eye, the center key point of the upper eyelid, and the center key point of the lower eyelid. The mouth includes the key points at the left corner of the mouth, the right corner of the mouth, the center key point of the upper lip, and the center key point of the lower lip. In step S303, one or more of the following indicators are calculated; (1) The angle between the elbow and the shoulder joint is θ, θ = arccos(a·b / |a||b|), where a is the line segment between the right shoulder point and the right wrist point, and b is the line segment between the right elbow point and the right wrist point; (2) The dominant frequency of the arm swing is fdominant, and the two-dimensional image coordinates of the key point of the right wrist in frame t are pt=(x t ,y t ), calculate the displacement sequence of the right wrist key points along the y-axis in N consecutive frames, forming a discrete signal S = {y1, y2, ..., y N }, where y N This represents the absolute y-axis coordinate of the wrist keypoint in the t-th frame image. The displacement sequence signal is then preprocessed to obtain signal S. filtered Then the filtered signal S filtered Perform a Fast Fourier Transform to convert it from the time domain to the frequency domain: Then, its power spectral density (PSD) distribution is obtained, and the frequency of the maximum energy peak in the power spectral density (PSD) distribution is defined as the dominant frequency fdominant of the arm swing. (3) Quantification of cross-correlation coefficient: Where L and R represent the joint angle θ sequences of the left and right wrists over a period of time, respectively, Cov(L,R) measures the synergy of their changes, and σ L and σ R Then it is the standard deviation of each sequence. (4) Aspect ratio of the eye region r eye r eye =h eye / w eye h eye The distance is defined as: the vertical distance between the center key points of the upper and lower eyelids of the left eye, and the vertical distance between the center key points of the upper and lower eyelids of the right eye; the average of the two vertical distances is taken. eye The horizontal distance between the key points at the left and right corners of the eye; (5) Opening coefficient k mouth k mouth =h mouth / w mouth Width w mouth To determine the horizontal distance between the key points at the left and right corners of the mouth, the height h is... mouth The vertical distance between the center key point of the upper lip and the center key point of the lower lip; In step S304, one of the following methods is selected to determine whether the target is drowning: The first determination method is: if the angle θ between the elbow and shoulder joint is greater than the set threshold, and the angle θ for 30 consecutive frames is greater than the set threshold, then it is determined to be a drowning state. The second method of judgment: if the dominant frequency of the arm swing is in the range of 3 to 5 Hz, it is judged as drowning; The third method of determination: if ρ LR If the value is approximately 0, then it is determined to be drowning; The fourth method of determination: the aspect ratio r of the eye region. eye If the value exceeds the threshold calculated based on the baseline value and persists for more than 5 frames, it is determined to be drowning. The fifth determination method: opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning. The sixth method of determination: Count the total time T during which the key nasal point is below the water surface within T=10s. submerged and the longest single duration T max If T submerged >6s or T max If the time exceeds 3 seconds, it is considered drowning; The seventh method of judgment: Calculate the frequency of bubbles appearing near the nose and mouth. If the frequency of bubbles appears continuously for more than 2 Hz and lasts for more than 3 seconds, it is judged as drowning. The ⑧th determination method: based on the dominant frequency of arm swing, fdominant, ρ LR The aspect ratio of the eye area is r eye Opening coefficient k mouth Total duration T submerged Longest single duration T max Multiple indicators, such as the frequency of bubbles appearing near the nose and mouth, are selected and input into a multivariate time series prediction model. The multivariate time series prediction model outputs a score, and then drowning is determined based on the score. The 9th determination method: based on the dominant frequency of arm swing, fdominant, ρ LR The aspect ratio of the eye area is r eye Opening coefficient k mouth Multiple indicators are selected and input into a multivariate time series prediction model. The multivariate time series prediction model outputs a score, and then drowning is determined based on the score. The 10th determination method: the aspect ratio r of the eye region eye The value is greater than the threshold and lasts for more than 5 frames, and the opening coefficient k mouth If the water level exceeds a predetermined threshold, it is considered drowning.

4. The multi-domain collaborative marine emergency rescue and search method according to claim 1, characterized in that, In step 3, the human body key point fitting and detection model is the YOLOv8 target detection model.

5. The multi-domain collaborative marine emergency rescue and search method according to claim 4, characterized in that, The YOLOv8 object detection model is an improved YOLOv8 object detection model. In the backbone part, the improved YOLOv8 object detection model retains CSPDarknet53 as the backbone network and introduces the spatial pyramid pooling enhancement module SPPFF+. In the neck part, the path aggregation network PANet is combined with Deformable Convolution.

6. The multi-domain collaborative marine emergency rescue and search method according to claim 4 or 5, characterized in that, The process of creating the human body key point fitting and detection model is as follows: Step (1): Establish a dataset of key point detection images of drowning victims; Step (2): Use the drowning crowd key point detection image dataset established in step (1) to train the YOLOv8 target detection model and obtain the human body key point fitting detection model.

7. The multi-domain collaborative marine emergency rescue and search method according to claim 6, characterized in that, The following method was used to establish the image dataset for key point detection of drowning victims: Degradation scenario simulation: Using physical model synthesis technology, underwater color distortion is simulated, with red light attenuation and blue-green light dominance: I R (d)=I R0 ·e -αRd I B (G)=I R0 ·e -αBG The red light attenuation coefficient αRd is taken as 0.2–0.5m. -1 The blue-green light attenuation coefficient αBG is taken as 0.05–0.1m. -1 ; Fog blurring, point spread function modeled with Gaussian kernel Suspended particle interference is used to generate synthetic images that are highly similar to the real underwater environment; Occlusion and distortion enhancement: Occlusion is simulated by randomly adding suspended objects such as aquatic plants and rocks, with the occlusion area accounting for 10%–30%, and non-linear deformed images are generated by combining optical distortion models; Illumination normalization processing: Multi-scale Gaussian filtering is applied to block normalize the image with uneven illumination; Adversarial example training: generating adversarial examples based on FGSM; The dataset consists of 20,000 synthetic images and 5,000 real underwater images, with key points of the human body annotated.

8. The multi-domain collaborative marine emergency rescue and search method according to claim 2, characterized in that, The drone lands on the drone landing platform of the unmanned boat using visual guidance, and the wireless charging device on the unmanned boat charges the drone.

9. The multi-domain collaborative marine emergency rescue and search method according to claim 1, characterized in that, The UAV integrates a high-precision millimeter-wave radar, a depth camera, and a high-sampling-rate inertial measurement unit (IMU). It achieves high-frequency, high-precision estimation of the UAV's six-degree-of-freedom pose through deep fusion of multimodal sensor information. The specific process is as follows: Step (1): Obtain the UAV's position observation information in the global coordinate system through the M10N GPS module; let the original observation provided by it at time t be: zGPS,t=[φ,λ,h] T In the formula, φ is latitude, λ is longitude, and h is altitude. Through coordinate transformation, it is converted to the three-dimensional position pGPS=[X,Y,Z]^T in the geocentric coordinate system. The transformation process can be expressed as: P GPS =T(φ,λ,h) In the formula, T is the transformation function from the geodetic coordinate system to ECEF, and this observation will be used as a global position reference input to the fusion filter; Step (2): Output the three-axis angular velocities ω in the body coordinate system through the inertial measurement unit (IMU). m and triaxial acceleration a m : oh m =ω t +b ω +n ω a m =a t +b a +R(q) T g+n a Where, ω t and a t For the actual angular velocity and acceleration, b ω and b a For zero bias of the gyroscope and accelerometer, n ω and n a To measure white noise, R(q) is the rotation matrix corresponding to the current attitude q, and g = [0,0,g]^T is the gravity vector in the global coordinate system; Step (3): Obtain visual-radar joint relative pose observation information through the LVI-SAM algorithm; analyze the image sequence I acquired by the airborne depth camera. k Extract multi-scale ORB feature points to obtain the feature point set: F(V) = fv, i∈R 2 ×S; obtain the descriptor; Feature extraction based on geometric distribution is performed on the point cloud Pr acquired by 3D radar to obtain structural features such as edge points ε and planar points S; Calculate the inter-frame matching success rate p of visual features respectively v The overlap ratio p between the radar point cloud and the radar point cloud r : Where d is the matching distance of the points in the point cloud; Furthermore, according to p v and p r Instantaneous weights w for computer vision and radar v w r : In the formula, ηv and η r Using the prior confidence coefficients, the relative pose transformation is solved by optimizing the joint cost function using weighted least squares: Wherein, Cv and Cr are the visual reprojection error and radar point cloud matching error, respectively; the relative pose observation T_{k-1,k} will serve as an important measurement input for EKF; Step (4) is to achieve multi-sensor fusion pose estimation by extended Kalman filtering; Step 1), State prediction, based on angular velocity ω measured by IMU. m and acceleration a m The state is recursively derived using a nonlinear dynamic model, and the system state vector is defined as x=[p,v,q,bω,b a ] T ; The prediction process is based on the following dynamic equations: in, This represents quaternion multiplication, where n bw and n ba For zero-biased random walk noise, the Runge-Kutta method is used to discretely integrate the above differential equation to obtain the state prediction xk|k-1 and the corresponding prediction covariance Pk|k-1. Step 2), observation update: Input the position observation zGPS provided by GPS and the relative pose observation T_{k-1,k} provided by the LVI-SAM algorithm into the EKF update stage, and design the adaptive observation noise covariance matrix R: Where α is the normalization factor and γ is the signal-to-noise ratio threshold, the state estimate xk|k and covariance Pk|k are updated by fusing observation information through Kalman gain calculation; Step (5) outputs a high-frequency, high-precision six-degree-of-freedom pose estimation result x = [p, v, q, b]. ω ,ba] T .

10. A multi-domain collaborative marine emergency rescue and search method, characterized in that, This includes a collaborative emergency search and rescue system, which comprises drones, unmanned vessels, underwater robots, and a control center. The multi-domain collaborative marine emergency rescue and search method includes the following steps: The first step is for the drone to conduct reconnaissance and search in the air to find the search and rescue target. The drone then approaches the search and rescue target and sends its own position information or the position information obtained by the drone from locating the target to the control center. The second step is for the unmanned vessel to sail to the target location, where at least two underwater robots are located in the gap between the unmanned vessel's catamaran structure, and the underwater robots move with the unmanned vessel. The third step involves at least two underwater robots working together. Step 1: The unmanned vessel releases the underwater robot, which then navigates in the water. Step 2: The first underwater robot performs a traversal search, finds the target, and sends the target's location coordinates to the second underwater robot. Step 3: The second underwater robot travels to the target location and then performs target identification. Step S301: The underwater binocular camera mounted on the second underwater robot collects video of the target; Step S302: The video is decomposed into continuous static image frames at a fixed frame rate, and then each frame image is input into the human key point fitting detection model to obtain the detection result image and identify the human key points. Step S303: Calculate the indicators based on the key points; Step S304: Determine whether the target is drowning based on the indicators; Step 4: The drone is positioned at the precise location coordinates provided by the first underwater robot, and then the drone deploys the rescue device.

Citation Information

Cited By

  • Unmanned aerial vehicle closed-loop intelligent cooperative search and rescue method and device

    CN121386880A

  • Unmanned aerial vehicle closed-loop intelligent cooperative search and rescue method and device

    CN121386880B