Visual guidance method for ROV-assisted emergency hooking of a diving bell

CN122550993APending Publication Date: 2026-08-11CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]由于深海环境存在洋流扰动,悬浮于水体中的ROV本体与同样处于悬浮状态的失能潜水钟均会受到水流冲击,这种相对运动导致ROV摄像画面中挂接接口的位置跳变,增加了挂接失败的概率

Benefits of technology

[0008]通过上述技术方案,由于构建了全局空间映射矩阵与六自由度相对位姿序列,使得ROV能够在深海洋流扰动环境下定位摆动中的潜水钟应急挂接接口;由于引入了捕获区间机制,从而将机械臂对目标微幅晃动的响应限定在合理范围内,减少了机械臂频繁往复振荡导致的碰撞风险,提高了深海应急挂接作业的成功率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550993A_ABST
    Figure CN122550993A_ABST
Patent Text Reader

Abstract

This application relates to the field of marine engineering operations and discloses a visual guidance method for emergency attachment of an ROV-assisted diving bell. The method includes: acquiring an environmental image of the emergency attachment area on the surface of the diving bell; extracting a marking pattern based on the environmental image; parsing the marking pattern to construct a global spatial mapping matrix of the diving bell's emergency attachment interface relative to a wrist-mounted camera; calculating a six-degree-of-freedom relative pose sequence of the diving bell's emergency attachment interface in continuous time steps based on the global spatial mapping matrix; generating a capture range for the diving bell's emergency attachment interface based on the six-degree-of-freedom relative pose sequence; and triggering a pose compensation command when the end-effector enters the capture range to control the end-effector to lock the diving bell's emergency attachment interface. This technical solution reduces the collision risk caused by frequent reciprocating oscillations of the robotic arm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine engineering operations, and in particular to a visual guidance method for emergency attachment of an ROV auxiliary diving bell. Background Technology

[0002] In fields such as deep-sea oil and gas exploration, marine engineering operations, and deep-sea rescue, diving bells play a crucial role as pressure vessels that safely transport divers to the seabed and back to the surface. In extreme sea conditions or when a sudden accident causes the diving bell to lose power, its umbilical cable to break, or it to become trapped on the seabed, a remotely operated vehicle (ROV) is deployed to carry a rescue hook or cable to assist in emergency rendezvous operations. The disabled diving bell is then retrieved and brought to the surface using the winch system of the surface support vessel.

[0003] In the relevant emergency rescue execution logic, ROV operators usually rely on the camera equipment on the ROV to provide real-time video footage. After approaching the diving bell, the operator combines the video feedback and uses a joystick to control the ROV body to slowly approach and drive the robotic arm to extend. Then, the hook at the end of the robotic arm is aligned with the lifting ring or emergency attachment interface on the top of the diving bell to complete the rigging and locking.

[0004] Due to ocean current disturbances in the deep-sea environment, both the ROV body suspended in the water and the disabled diving bell, which is also suspended, are impacted by the water flow. This relative motion causes the position of the mounting interface in the ROV camera footage to change, increasing the probability of mounting failure.

[0005] Solving this technical problem is a technical challenge that needs to be overcome by those skilled in the art. Summary of the Invention

[0006] This application provides a visual guidance method for emergency attachment of an ROV auxiliary diving bell, which at least partially solves the above-mentioned technical problems.

[0007] To achieve the above objectives, this application provides a visual guidance method for emergency attachment of an ROV auxiliary diving bell, comprising: Acquire environmental images of the emergency attachment area located on the surface of the diving bell; Extract the identification pattern based on the environmental image; The identification pattern is analyzed to construct a global spatial mapping matrix of the diving bell emergency mounting interface relative to the wrist camera device; The six-degree-of-freedom relative pose sequence of the diving bell emergency mounting interface in continuous time steps is calculated based on the global spatial mapping matrix. A capture range for the emergency mounting interface of the diving bell is generated based on the six-degree-of-freedom relative pose sequence; When the end-mount tool enters the capture zone, a pose compensation command is triggered to control the end-mount tool to lock the diving bell emergency attachment interface.

[0008] Through the above technical solutions, by constructing a global spatial mapping matrix and a six-degree-of-freedom relative pose sequence, the ROV can locate the emergency docking interface of the swaying diving bell in the deep ocean current disturbance environment; by introducing a capture interval mechanism, the response of the robotic arm to the slight sway of the target is limited to a reasonable range, reducing the collision risk caused by the frequent reciprocating oscillation of the robotic arm and improving the success rate of deep-sea emergency docking operations.

[0009] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating the steps of a visual guidance method for emergency attachment of an ROV-assisted diving bell, as provided in an exemplary embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the protection scope of this application.

[0013] This application provides a visual guidance method for emergency attachment of an ROV auxiliary diving bell. Please refer to [link / reference]. Figure 1 The visual guidance method for emergency attachment of an ROV auxiliary diving bell provided in this application includes the following steps: Step 101: Acquire an environmental image of the emergency attachment area on the surface of the diving bell. Specifically, an environmental image of the emergency attachment area on the surface of the diving bell is acquired using a camera device installed near the end of the ROV robotic arm. The camera device can be a combination of a low-light camera, a structured light depth camera, or an infrared rangefinder camera and a black-and-white image sensor array. The environmental image referred to here is an image containing the attachment area that is exposed and output by the camera device in a deep-sea operating environment.

[0014] Step 102: Extract the identification pattern based on the environmental image. Specifically, perform multi-scale edge detection and depth gradient analysis on the environmental image to identify reflective coating patches on the diving bell shell; the identification pattern here refers to the set of pixels of the reflective coating patches stripped from the background of the environmental image.

[0015] Step 103: Analyze the marking pattern to construct a global spatial mapping matrix of the diving bell emergency mounting interface relative to the wrist camera device. Specifically, the two-dimensional pixel coordinates are back-projected to the three-dimensional camera coordinate system using the intrinsic parameter matrix of the camera device. Key corner points or contour centers of the marking pattern are extracted, a local rigid body reference system of the target object is established, and this reference system is mapped to the camera coordinate system. The global spatial mapping matrix mentioned here refers to a multi-dimensional linear algebraic matrix that covers the relative position and deflection attitude between the diving bell target and the wrist camera device, which can be represented by a 4×4 homogeneous transformation matrix. It can be understood that the global spatial mapping matrix establishes the spatial coordinate mapping relationship between the diving bell and the ROV.

[0016] Step 104: Calculate the six-degree-of-freedom relative pose sequence of the diving bell's emergency docking interface in continuous time steps based on the global spatial mapping matrix. Specifically, the global spatial mapping matrix is ​​iteratively calculated and cached at a preset image acquisition frame rate to construct a temporal matrix queue. The translation and rotation components of adjacent time steps are temporally differentially calculated to extract the relative velocity and relative angular velocity components. After removing high-frequency jitter noise, the pose sequence representing the actual spatial swing trajectory is reconstructed. The six-degree-of-freedom relative pose sequence mentioned here refers to a continuous temporal state set composed of three-dimensional spatial translation coordinate data and pitch, yaw, and roll Euler angle data over time in multiple discrete time steps. It can be understood that the six-degree-of-freedom relative pose sequence reflects the real-time motion trajectory of the diving bell in the water.

[0017] Step 105: Generate a capture range for the emergency docking interface of the diving bell based on the six-degree-of-freedom relative pose sequence. Specifically, the spatial offset extreme value is extracted by traversing the pose sequence and combined with the maximum follow-up response radius of the robotic arm under the current torque parameters to delineate the three-dimensional boundary range surrounding the emergency docking interface of the diving bell to generate the capture range. The capture range here refers to the threshold range within which the emergency docking interface is allowed to make slight swaying without the robotic arm end effector needing to immediately follow and correct it. The spatial offset extreme value here refers to the farthest three-dimensional offset boundary coordinates reached in the time domain by the disabled diving bell due to the swaying of ocean currents in the water. The maximum follow-up response radius refers to the longest spherical distance that the end effector of the robotic arm can cover and follow without exceeding the current rated current and limit of the motor. It can be understood that the capture range is the effective working range within which the robotic arm can stably follow the sway of the target.

[0018] Step 106: When the end effector enters the capture zone, a pose compensation command is triggered to control the end effector to lock the emergency docking interface of the diving bell. Specifically, when the spatial pose coordinates of the end effector fall into the capture zone, the main control calculation unit wakes up the follow-up locking process and monitors the distance deviation between the end effector and the emergency docking interface of the diving bell in real time. When the distance deviation is lower than the engagement threshold, a closing trigger electrical signal is output to drive the end effector to engage the emergency docking interface of the diving bell. The distance deviation mentioned here refers to the linear deviation between the center point of the end effector and the center point of the docking lock hole or lifting ring of the diving bell in the same three-dimensional spatial coordinate system. The engagement threshold is a pre-calibrated maximum allowable alignment error constant that ensures that the end effector jaws can slide smoothly into the inner hole of the interface without squeezing the edge of the interface and causing rigid jamming. It can be taken as 5 to 20 mm. It can be understood that when the distance between the end effector and the interface is small enough, the final engagement action is triggered.

[0019] In one implementation, the three-dimensional envelope shape of the capture interval can be a cuboid bounding box to facilitate quick determination of whether the end effector falls into the interval, or a directional bounding box with a gradually shrinking boundary can be used to gradually shrink the control gain as it approaches the boundary to achieve a smooth transition.

[0020] In another implementation, the engagement threshold can be adjusted according to the jaw size and the inner diameter of the interface of the end-mounted tool. The greater the difference between the jaw size and the inner diameter, the larger the engagement threshold can be set.

[0021] Through the above technical solutions, the ROV is able to locate the emergency docking interface of the diving bell in a swaying environment by using a global spatial mapping matrix and a six-degree-of-freedom relative pose sequence. By capturing the range, the response of the robotic arm to the slight sway of the target is limited to a reasonable range, reducing the collision risk caused by the frequent reciprocating oscillation of the robotic arm and improving the success rate of deep-sea emergency docking operations.

[0022] In some embodiments, extracting the identification pattern based on an environmental image includes: The depth gradient distribution features and edge contour features in the environmental image are extracted. Specifically, the Sobel or Canny operator can be used to perform convolution filtering on the input 2D grayscale image to extract pixel abrupt edges as edge contour features; the depth image is read and the spatial first-order partial derivatives of the depth corresponding to each pixel are calculated to generate depth gradient distribution features; the depth gradient distribution features mentioned here refer to the set of spatial first-order partial derivatives of the depth corresponding to each pixel calculated using the ranging data of the depth camera, reflecting the degree of spatial undulation; the edge contour features refer to the binary pixel abrupt edge matrix extracted by the edge detection operator; it can be understood that depth gradient and edge contour are two complementary feature extraction methods.

[0023] The initial feature pixel clusters are extracted by fusing the depth gradient distribution features and the edge contour features. Specifically, the binarized matrix of the edge contour features and the depth gradient vector are superimposed by matrix channels or tensor concatenated using the same coordinate index to form a composite feature cluster. Connected component analysis and clustering segmentation are then performed on the composite feature clusters to extract the initial feature pixel clusters. The initial feature pixel clusters referred to here are the suspicious target pixel regions that have been initially aggregated after shallow edge and depth convolution processing but have not yet undergone geometric screening; they can be understood as candidate regions for identifying patterns.

[0024] The initial feature pixel cluster is compared with a preset baseline geometric template using morphological matching. Specifically, the baseline geometric template stored in the main control computing unit is invoked. The baseline geometric template contains the standard geometric shape, size parameters, and edge features of the identifier pattern. A normalized cross-correlation function is used to calculate the similarity scalar between the initial feature pixel cluster and the baseline geometric template. The similarity value ranges from 0 to 1. The morphological matching degree mentioned here refers to the similarity value calculated using the normalized cross-correlation function or Hausdorff distance. It can be understood that the morphological matching degree reflects the degree of similarity between the candidate region and the standard template.

[0025] When the morphological matching degree is greater than a threshold, the corresponding initial feature pixel cluster is retained to form the identification pattern. Specifically, when the calculated similarity value exceeds a preset threshold, the initial feature pixel cluster is determined to be a valid identification pattern and is retained; the morphological matching degree threshold can be set to 0.75 to 0.85.

[0026] In one alternative approach, the morphological matching degree can be calculated using the Hausdorff distance metric method, which calculates the maximum and minimum distances between the edges of the initial feature pixel clusters and the edges of the reference geometric template; the smaller the distance, the higher the matching degree.

[0027] In another implementation, when there are multiple candidate pixel clusters in the environmental image, they can be sorted from high to low according to morphological matching degree, and the pixel cluster with the highest matching degree can be selected as the identification pattern.

[0028] The above technical solution employs multimodal composite filtering using depth gradient and edge contour. The depth gradient provides spatial undulation information, while the edge contour provides contour boundary information. The two complement each other in the filtering process, thereby reducing image artifact interference caused by light spot refraction in the complex background of the deep sea and improving the reliability of the identification pattern extraction.

[0029] In some embodiments, parsing the identification pattern to construct a global spatial mapping matrix of the diving bell emergency mounting interface relative to the wrist camera device includes: The centroid coordinates and axial deflection angle of the sign pattern are extracted. Specifically, the geometric centroid pixel coordinates of the sign pattern are calculated using the polygon vertex averaging method, and its three-dimensional distances in the X, Y, and Z coordinates in the camera coordinate system are calculated based on the intrinsic and extrinsic parameter matrix of the camera device to form the centroid coordinates. The axial deflection angle is obtained by analyzing the principal axis direction of the sign pattern and calculating the angle between it and the axes of the camera coordinate system. Here, the centroid coordinates refer to the spatial position coordinates of the geometric centroid of the sign pattern in the three-dimensional coordinate system of the camera; the axial deflection angle refers to the rotation angle of the principal axis direction of the sign pattern relative to the axes of the camera coordinate system.

[0030] The origin of the reference coordinate system is the focal center of the wrist-mounted camera device. Specifically, the origin of the reference coordinate system is located at the optical center of the camera device, and the coordinate axes are aligned with the imaging plane of the camera device.

[0031] Calculate the spatial translation vector of the centroid coordinates of the marker relative to the origin of the reference coordinate system. Specifically, the spatial translation vector is a 3×1 column vector containing distance components along the X, Y, and Z axes, representing the deviation of the centroid of the marker pattern relative to the optical center of the camera device; the spatial translation vector referred to here is a three-dimensional column vector characterizing the deviation of the mounting interface from the optical center of the camera device in the Cartesian coordinate system.

[0032] A three-dimensional rotation matrix is ​​constructed based on the axial deflection angle of the marker. Specifically, the axial deflection angle of the marker is converted into Euler angles about the X, Y, and Z axes, and a 3×3 three-dimensional rotation matrix R is formed by trigonometric function expansion.

[0033] Based on the three-dimensional rotation matrix, a rotation domain of a 4×4 homogeneous transformation matrix is ​​constructed, and the spatial translation vector is filled into the translation domain of the 4×4 homogeneous transformation matrix. The global spatial mapping matrix is ​​then generated through matrix multiplication. Specifically, a 3×3 rotation matrix R is filled into the rotation domain of the upper left corner of the 4×4 homogeneous transformation matrix T, and a 3×1 spatial translation vector t is filled into the translation domain of the upper right corner of the 4×4 homogeneous transformation matrix. The bottom row is filled with a constant vector (0,0,0,1).

[0034] In one alternative approach, the singular value decomposition algorithm can be used to solve the rigid body transformation relationship of the matched point cloud clusters.

[0035] In another implementation, the marking pattern can be set with at least three non-collinear feature corner points, and the global spatial mapping matrix can be obtained by solving the transformation matrix of the corresponding point pairs.

[0036] Through the above technical solution, a 4×4 homogeneous transformation matrix containing rotation and translation information is established by back-projecting the identification pattern, thereby unifying the mapping of two-dimensional pixels to a three-dimensional physical coordinate system, eliminating the depth visual illusion caused by a single two-dimensional visual image, and providing a unified spatial reference benchmark for subsequent pose calculation and capture range generation.

[0037] In some embodiments, calculating the six-degree-of-freedom relative pose sequence of the diving bell emergency docking interface in continuous time steps based on the global spatial mapping matrix includes: The global spatial mapping matrix is ​​iteratively calculated and cached at a preset image acquisition frame rate for multiple consecutive time steps to construct a temporal matrix queue. Specifically, the main control computing unit keeps synchronized with the image acquisition frame rate of the wrist camera device, for example, 25 to 50 frames per second. Within each video frame processing cycle, the latest global spatial mapping matrix is ​​iteratively solved and pushed into the first-in-first-out cache queue. The length of the cache queue can be set to 50 to 200 frames.

[0038] The relative velocity and relative angular velocity components are extracted by temporal difference calculation of the translation and rotation components of adjacent time steps in the temporal matrix queue. Specifically, the translation vectors and rotation matrices of adjacent frames in the extracted queue are separated. A first-order difference operation is performed on the translation vectors of adjacent frames, that is, the difference between the translation vectors of adjacent frames is divided by the time interval to solve for the linear velocity vector; Euler angles are extracted for the rotation matrix and a first-order difference operation is performed to solve for the angular velocity vector; the relative velocity component refers to the rate of change of the translation vector between adjacent time steps; the relative angular velocity component refers to the rate of change of the rotation matrix between adjacent time steps.

[0039] The jitter noise points in the relative velocity component and the relative angular velocity component are removed, and the motion trajectory data is extracted. Specifically, a low-pass filter or extended Kalman filter is applied to perform frequency domain filtering on the linear velocity and angular velocity vectors, attenuating spectral components higher than a preset cutoff frequency; the motion trajectory data referred to here is the data set that retains only the periodic oscillation trend component of the diving bell body after filtering out the local mechanical jitter caused by water shock waves.

[0040] Based on the motion trajectory data, a six-degree-of-freedom relative pose sequence representing the actual swing trajectory in space is reconstructed and generated. Specifically, the filtered velocity vectors are time-series integrated and reassembled into a pose sequence to obtain a six-degree-of-freedom relative pose sequence representing the actual swing trajectory of the diving bell.

[0041] In one alternative approach, a moving average window can be used to smooth the differential velocity array, with the window length set to 5 to 10 frames.

[0042] By performing time-series difference and filtering on the global spatial mapping matrix, the high-frequency visual jitter caused by underwater turbulence is removed, the true swing trend of the diving bell is preserved, and the robotic arm is prevented from outputting invalid reverse drive current due to tracking high-frequency noise.

[0043] In another implementation, curve fitting can be performed on the time-series matrix queue, using polynomial fitting or spline fitting to extract the main trend curve of the motion trajectory.

[0044] In some embodiments, generating a capture range for the emergency docking interface of a diving bell based on a six-degree-of-freedom relative pose sequence includes: Extract the spatial offset extrema of the six-DOF relative pose sequence. Specifically, traverse the temporal pose sequence and select the maximum and minimum values ​​of the coordinates in each of the X, Y, and Z directions to construct a set of spatial offset extrema.

[0045] Calculate the maximum servo response radius of the robotic arm joints of the end-effector under the current torque parameters. Specifically, read the current feedback current of each joint motor and combine it with the robotic arm link length and DH parameter matrix to calculate the spherical reachable workspace radius that can achieve a stable response under the current remaining torque margin through Jacobian matrix evaluation or kinematic solution.

[0046] The three-dimensional boundary range surrounding the emergency docking interface of the diving bell is defined based on the spatial offset extreme value. Specifically, a cuboid or sphere boundary range surrounding the emergency docking interface of the diving bell is constructed using the spatial offset extreme value as the boundary; the three-dimensional boundary range refers to the spatial region surrounding the emergency docking interface of the diving bell defined based on the spatial offset extreme value.

[0047] The capture range is generated by superimposing the maximum follow-up response radius on the three-dimensional boundary range. Specifically, the capture range is generated by superimposing the cuboid boundary range and the spherical reachable workspace in three-dimensional space. The size of the capture range should be larger than the swing range of the diving bell and smaller than the reach limit of the robotic arm.

[0048] In one alternative, the three-dimensional envelope shape of the capture interval can be a cuboid bounding box, which facilitates quick determination of whether the end effector falls into the interval.

[0049] In another implementation, the capture interval can be in the form of a directed bounding box with a gradually shrinking boundary, which controls the gain as it approaches the boundary.

[0050] By superimposing the spatial offset extreme value with the maximum follow-up response radius of the robotic arm in three dimensions, the capture range is simultaneously limited by the actual swing range of the target and the physical reach limit of the robotic arm. This prevents the control unit from issuing over-limit drive commands that could cause the robotic arm to break and jam, thus ensuring the safety of the splicing operation.

[0051] In some embodiments, triggering a pose compensation command when the end-effector enters the capture zone to control the end-effector to lock the diving bell emergency attachment interface includes: The follow-up locking process is activated when the spatial pose coordinates of the end-effector tool are detected to fall within the capture range. Specifically, the main control computing unit calculates the coordinates of the flange center point of the end-effector tool in real time and determines whether the coordinates are within the capture range. When it is determined that the tool has fallen within the capture range, the main control computing unit activates the follow-up locking process.

[0052] During the follow-up locking process, the distance deviation between the end-mounted attachment tool and the emergency attachment interface of the diving bell is monitored in real time. Specifically, the main control computing unit continuously calculates the Euclidean distance scalar between the flange center point of the end-mounted attachment tool and the interface center coordinates in the monitoring thread of this process, as the distance deviation.

[0053] When the distance deviation is lower than the engagement threshold, a closing trigger signal is output to drive the end attachment tool to engage the emergency engagement interface of the diving bell. Specifically, the distance deviation is compared with the engagement threshold preset in the memory. When the distance deviation is continuously less than the engagement threshold for a preset number of cycles, the main control computing unit operates the GPIO pin to flip, outputting a high-level closing trigger signal to the electromagnetic reversing valve or electro-hydraulic servo valve of the engagement claw, driving the end attachment tool to perform the engagement action.

[0054] In one alternative, the engagement threshold can be adjusted based on the jaw size and the inner diameter of the interface of the end-mounted tool. The greater the difference between the jaw size and the inner diameter, the larger the engagement threshold can be set.

[0055] In another implementation, before outputting the closing trigger signal, the trend of distance deviation can be determined first, and the closing signal is triggered only when the distance deviation shows a converging trend, thus avoiding false triggering during fluctuating periods.

[0056] By using the above technical solution, the distance deviation monitoring and closure triggering are strictly constrained within the follow-up locking process, and a closure signal is only issued when the deviation is continuously lower than the engagement threshold. This reduces the probability of premature malfunctions causing empty grabs or impact damage, and improves the reliability of the final engagement.

[0057] In some embodiments, after extracting an initial cluster of feature pixels by fusing depth gradient distribution features and edge contour features, the method further includes: Calculate the pixel integrity of the initial feature pixel cluster. Specifically, the ratio of the area of ​​non-zero feature pixels within the extracted identifier region to the standard value of the benchmark area library is used as the integrity scalar, with the integrity value ranging from 0 to 1. The pixel integrity here refers to the percentage ratio of the total number of effective feature pixels extracted in the current image to the total pixel area that the identifier should have under ideal unobstructed conditions; it can be understood that pixel integrity reflects the degree to which the identifier pattern is obstructed.

[0058] When the pixel integrity is less than the occlusion threshold, the two-dimensional pixel coordinates of the identified valid pattern nodes are extracted. Specifically, the occlusion threshold can be set to 0.6 to 0.8. When the integrity is lower than the threshold, it is determined that there is severe occlusion. At this time, the pixel coordinate information of the surviving key points that are not occluded is extracted from the image as valid pattern nodes.

[0059] The perspective transformation of the two-dimensional pixel coordinates is performed to obtain the homography transformation matrix. Specifically, a preset ideal three-dimensional geometric model of the pattern is called from the memory of the main control computing unit to establish the perspective mapping relationship between the current incomplete image viewpoint and the ideal model plane. The homography transformation matrix is ​​obtained by solving a homogeneous linear equation system. The homography transformation matrix mentioned here refers to a full-rank transformation matrix, which is a 3×3 matrix used in computer vision spatial geometry to describe the perspective projection mapping of the same planar target between two different camera viewpoints or between an ideal plane and the actual imaging plane.

[0060] The homography transformation matrix is ​​used to perform inverse projection calculations on the occluded area to generate a complete marker pattern. Specifically, the model coordinates of the theoretically occluded nodes are multiplied by the homography transformation matrix to calculate the pixel coordinates that they should appear in the current image, and the corresponding incomplete pattern data in the video memory is copied to complete the reconstruction of the marker pattern.

[0061] In one alternative approach, the extended Kalman filter algorithm can be used to smoothly predict the homography transformation matrix of historical frames in order to resist the distortion of single-frame pixel extraction.

[0062] In another implementation, when there are fewer than four valid pattern nodes, the location of the missing nodes can be estimated using polynomial fitting or feature point extrapolation.

[0063] By using the above technical solution, the obscured area can be reconstructed by inverse projection using the homography transformation matrix, thus restoring the complete pattern information even when the marking pattern is partially obscured, ensuring the continuous and uninterrupted calculation of the global spatial mapping relationship.

[0064] In some embodiments, perspective transformation is performed on the two-dimensional pixel coordinates to obtain the homography transformation matrix, including: Step 801: Select at least four valid pattern nodes that are not collinear. Specifically, select four or more feature corner points that are not on the same straight line from the valid feature point matrix as control points for solving the homography transformation matrix. The more control points, the better.

[0065] Construct a homogeneous linear system of equations between the two-dimensional pixel coordinates of the effective pattern node and the reference pattern coordinates.

[0066] The homography transformation matrix is ​​extracted by solving the homography system of linear equations using a singular value decomposition (SVD) algorithm. Specifically, the SVD algorithm of the underlying mathematical operation acceleration library is called to decompose the coefficient matrix A, extract the eigenvectors corresponding to the smallest singular values, and rearrange them into a 3×3 homography transformation matrix H.

[0067] The theoretical coordinates of the corresponding missing nodes are input into the homography transformation matrix to map and obtain the estimated pixel coordinates. The estimated pixel coordinates are then filled into the pattern matrix to complete the reconstruction.

[0068] In one alternative approach, when the number of effective pattern nodes exceeds four, the least squares method or the RANSAC algorithm can be used to solve the overdetermined system of equations to improve computational accuracy.

[0069] By employing the above technical solution to solve the homography transformation matrix using homogeneous linear equations and singular value decomposition, and by fully utilizing the global perspective geometric constraints of the remaining effective nodes, the original pattern can still be reconstructed even when a large area of ​​the pattern is covered, thus improving the confidence level of the reconstruction of missing patterns.

[0070] In some embodiments, when caching the global spatial mapping matrix at multiple consecutive time steps, the method further includes: Extract the inertial measurement acceleration sequence of the submersible ROV body equipped with the wrist camera device. Specifically, read the three-axis linear acceleration readings output by the MEMS six-axis or nine-axis inertial measurement unit fixed to the ROV chassis via SPI or I2C bus.

[0071] The body disturbance displacement sequence is obtained based on the inertial measurement acceleration sequence. Specifically, after eliminating the gravity component by combining gyroscope data, a second time-domain integral operation is performed on the curve of the change of the three-axis linear acceleration readings over time to obtain the relative displacement increment in three-dimensional space, which is used as the body disturbance displacement sequence; the body disturbance displacement sequence referred to here refers to the offset coordinate trajectory of the ROV body in the water body relative to the ideal static suspension state.

[0072] The extended Kalman filter algorithm is used to fuse the body perturbation displacement sequence with the translation components in the time-series matrix queue to compensate for and eliminate body sway errors. Specifically, the main control computing unit constructs the extended Kalman filter state observation equation, uses the body perturbation displacement sequence as an inertial prior estimate, uses the translation vector of the corresponding frame in the time-series matrix queue as a visual observation update, performs mathematical fusion operations, and compensates for and subtracts the axial offset value of the body wandering.

[0073] In one alternative, Kalman fusion can be performed only on the translation component, while compensation for the rotation component can be ignored as a small angular change.

[0074] In another implementation, unscented Kalman filtering or particle filtering algorithms can be used instead of extended Kalman filtering to improve the estimation accuracy of nonlinear systems.

[0075] By using the above technical solution, the swaying displacement of the ROV caused by water flow impact is separated from the relative pose by fusing the inertial measurement acceleration sequence with the visual observation Kalman filter. This makes the final calculated relative motion sequence only represent the swaying trend of the diving bell target, thus improving the accuracy of the acquisition interval generation.

[0076] In some embodiments, after the drive end attachment tool engages the emergency attachment interface of the diving bell, the method further includes: The real-time feedback hydraulic pressure value of the closing cylinder of the end-effector is obtained. Specifically, the analog voltage signal is read and converted into an actual megapascal value by a pressure sensor or A / D conversion module deployed on the hydraulic circuit of the robotic arm end.

[0077] When the real-time feedback hydraulic pressure value exceeds the no-load closing threshold and remains for a preset duration, a rigid locking feedback force is determined to be generated. Specifically, the hydraulic pressure value is compared with the upper limit of the reference pressure when the robotic arm is in no-load opening and closing, and a preset duration is introduced as a time anti-jitter window. Once the condition is met, it is determined that the mechanical end effector has made rigid contact with the diving bell hook ring.

[0078] The pose compensation command is cut off and the six-DOF robotic arm base brake is locked. Specifically, the main control logic chip immediately blocks the enable signal flow sent to each joint controller and simultaneously sends a power-on or power-off lock-up level command to the base electromagnetic brake, thereby achieving instant locking of the robotic arm base.

[0079] In another implementation, a tiered locking strategy can be adopted, which first cuts off the drive of some joints and then gradually locks all joints to avoid excessive impact load.

[0080] Through the above technical solution, when the hydraulic pressure value exceeds the no-load closing threshold and is maintained for a preset time, a rigid locking feedback force is generated. This allows the drive to be cut off and the brake to be locked in time at the moment the robotic arm and the hook ring make rigid contact, thus avoiding the risk of the robotic arm breaking due to overload or the diving bell body being overturned due to continued output of follow-up drive current.

[0081] In some embodiments, defining the three-dimensional boundary range surrounding the emergency docking interface of the diving bell based on the spatial offset extremum includes: The temporal spatial coordinate sequence is extracted from the motion trajectory data to construct a multidimensional spatial trajectory matrix. Specifically, the extracted coordinate points are unfolded sequentially and filled column by column to form a multidimensional tensor containing the current position feature column, with each row corresponding to the spatial coordinates of one time step.

[0082] Calculate the instantaneous velocity scalar and instantaneous acceleration scalar in the multidimensional spatial trajectory matrix. Specifically, perform differential differentiation on the position column vector of the tensor with respect to the time step to obtain the velocity column vector, and then differentiate again to obtain the acceleration column vector. Calculate the scalar magnitudes of the velocity and acceleration vectors.

[0083] The stagnation time window of the diving bell's kinetic energy is calculated by combining the instantaneous velocity scalar and the instantaneous acceleration scalar. Specifically, the time index interval where the instantaneous velocity scalar is minimal and the instantaneous acceleration scalar reverses direction is calculated. This interval corresponds to the period when the diving bell's kinetic energy is at its local minimum level. The stagnation time window mentioned here refers to an extremely short time-domain segment where a pendulum or suspended object reaches its farthest endpoint during its reciprocating oscillation, where its speed is slowest and the risk of collision is lowest. It can be understood that the stagnation time window is the optimal docking opportunity.

[0084] Within the swing stagnation time window, a scaling factor is calculated based on the decay gradient of the instantaneous velocity scalar. Specifically, the degree of decay of the instantaneous velocity scalar within the swing stagnation time window is monitored. When the velocity approaches zero, the decay gradient increases, and a scaling factor greater than 1 is calculated based on this decay gradient. The larger the decay gradient, the smaller the scaling factor, indicating that the capture interval can be appropriately shrunk.

[0085] The scaling factor is multiplied by the spatial offset extreme value to adjust the three-dimensional boundary range surrounding the emergency docking interface of the diving bell. Specifically, the scaling factor is algebraically multiplied by the length, width, and height scalars of the original capture interval boundary to obtain the adjusted three-dimensional boundary range.

[0086] In another implementation, the scaling factor can be second-order corrected by combining the value of the instantaneous acceleration scalar. The larger the acceleration value, the more obvious the stationary point characteristics, and the scaling factor can be appropriately increased.

[0087] Through the above technical solution, the scaling factor is calculated based on the velocity decay gradient within the swing stagnation time window to adjust the three-dimensional boundary range. The contraction of the capture interval is bound to the period when the target kinetic energy is the lowest, so that the robotic arm can dock at the stagnation point when the diving bell has the lowest kinetic energy, reducing the risk of kinetic energy impact at the moment of docking.

[0088] In some embodiments, the oscillation stationary time window of the diving bell's oscillation kinetic energy is solved by fusing instantaneous velocity scalars and instantaneous acceleration scalars, including: The moment when the instantaneous velocity scalar is below the velocity safety threshold and the instantaneous acceleration scalar reaches a local maximum is extracted as a stationary time stamp, and the corresponding instantaneous acceleration peak value is extracted. Specifically, the multidimensional spatial trajectory matrix is ​​traversed to search for time points that satisfy the condition that the instantaneous velocity scalar is below the threshold and the instantaneous acceleration scalar reaches a local maximum. The index of this time point is recorded as a stationary time stamp, and the instantaneous acceleration peak value at that moment is extracted. The stationary time stamp mentioned here refers to the time scale recorded by the main control calculation unit when the lateral motion velocity of the deep-sea diving bell instantaneously decays to near zero due to the action of gravity and buoyancy torque during its reciprocating swing cycle and is about to change its motion direction. It can be understood that the stationary time stamp is the time mark of the lowest velocity point.

[0089] The stationary timestamp is compensated for in advance based on the preset hydraulic response delay of the robotic arm to anchor the pre-trigger reference time. Specifically, the dead zone time constant from the issuance of the control signal to the complete opening of the hydraulic valve of the cylinder to establish pressure is retrieved from the parameter configuration file of the main control computing unit. The stationary timestamp is subtracted from this time constant to complete the left shift offset calculation, thus obtaining the pre-trigger reference time.

[0090] The dwell time is obtained based on the speed safety threshold and the instantaneous acceleration peak value. Specifically, the dwell time in time dimension is obtained by dividing the speed safety threshold by the instantaneous acceleration peak value. The dwell time represents the time required for the diving bell to decay from the safety threshold to zero near the dwell point.

[0091] The swinging pause time window is defined based on the dwell time, using the pre-triggered reference time as the anchor point. Specifically, a time window is constructed with the pre-triggered reference time as the anchor point, extending half a dwell time to each side, and the width of the dwell time is equal to the pre-triggered reference time. This time window serves as the swinging pause time window.

[0092] In one alternative, the width of the oscillation stagnation time window can be adaptively adjusted according to the magnitude of the instantaneous acceleration peak. The larger the acceleration peak, the sharper the stagnation feature, and the window width can be appropriately narrowed.

[0093] In another implementation, statistical analysis can be performed on the stationary timestamps of multiple consecutive periods to calculate the average stationary time interval, which can be used as the basis for determining the maximum allowable waiting time.

[0094] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0095] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0096] The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0097] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A visual guidance method for emergency attachment of an ROV-assisted diving bell, characterized in that, include: Acquire environmental images of the emergency attachment area located on the surface of the diving bell; Extract the identification pattern based on the environmental image; The identification pattern is analyzed to construct a global spatial mapping matrix of the diving bell emergency mounting interface relative to the wrist camera device; The six-degree-of-freedom relative pose sequence of the diving bell emergency mounting interface in continuous time steps is calculated based on the global spatial mapping matrix. A capture range for the emergency mounting interface of the diving bell is generated based on the six-degree-of-freedom relative pose sequence; When the end-mount tool enters the capture zone, a pose compensation command is triggered to control the end-mount tool to lock the diving bell emergency attachment interface.

2. The method according to claim 1, characterized in that, Extracting the identification pattern based on the environmental image includes: Extract the depth gradient distribution features and edge contour features from the environmental image; The initial feature pixel cluster is extracted by fusing the depth gradient distribution features with the edge contour features; The initial feature pixel cluster is compared with a preset reference geometric template to calculate the morphological matching degree. When the morphological matching degree is greater than the threshold, the corresponding initial feature pixel cluster is retained to form the identification pattern.

3. The method according to claim 2, characterized in that, The process of parsing the identification pattern to construct a global spatial mapping matrix of the diving bell emergency mounting interface relative to the wrist camera device includes: Extract the centroid coordinates and axial deflection angle of the logo pattern; The focal center of the wrist camera device is taken as the origin of the reference coordinate system; Calculate the spatial translation vector of the centroid coordinates of the marker relative to the origin of the reference coordinate system; Construct a three-dimensional rotation matrix based on the axial deflection angle of the identified axis; Based on the three-dimensional rotation matrix, a rotation domain of a 4x4 homogeneous transformation matrix is ​​constructed, and the spatial translation vector is filled into the translation domain of the 4x4 homogeneous transformation matrix. The global spatial mapping matrix is ​​then generated through matrix multiplication.

4. The method according to claim 3, characterized in that, The six-degree-of-freedom relative pose sequence of the diving bell emergency attachment interface in continuous time steps is calculated based on the global spatial mapping matrix, including: The global spatial mapping matrix at multiple consecutive time steps is iteratively calculated and cached according to a preset image acquisition frame rate to construct a temporal matrix queue. The relative velocity component and the relative angular velocity component are extracted by performing time-series difference calculation on the translation and rotation components of adjacent time steps in the time-series matrix queue. Remove jitter noise points from the relative velocity component and the relative angular velocity component, and extract motion trajectory data; Based on the motion trajectory data, a six-degree-of-freedom relative pose sequence representing the true swing trajectory in space is reconstructed and generated.

5. The method according to claim 4, characterized in that, Based on the six-degree-of-freedom relative pose sequence, a capture range is generated for the emergency mounting interface of the diving bell, including: Extract the spatial offset extrema of the six-degree-of-freedom relative pose sequence; Calculate the maximum servo response radius of the robotic arm joint of the end-effector under the current torque parameters; The three-dimensional boundary range surrounding the emergency mounting interface of the diving bell is defined based on the spatial offset extreme value; The capture range is generated by superimposing the maximum follow-up response radius on the three-dimensional boundary range.

6. The method according to claim 5, characterized in that, When the end-mount tool enters the capture zone, a pose compensation command is triggered to control the end-mount tool to lock the diving bell emergency attachment interface, including: The follow-up locking process is activated when the spatial pose coordinates of the end-mounted tool are detected to fall into the capture range. During the follow-up locking process, the distance deviation between the end attachment tool and the emergency attachment interface of the diving bell is monitored in real time; When the distance deviation is below the engagement threshold, a closing trigger electrical signal is output to drive the end attachment tool to engage the emergency attachment interface of the diving bell.

7. The method according to claim 6, characterized in that, After fusing the depth gradient distribution features and the edge contour features to extract the initial feature pixel clusters, the method further includes: Calculate the pixel integrity of the initial feature pixel cluster; When the pixel integrity is less than the occlusion threshold, extract the two-dimensional pixel coordinates of the identified valid pattern nodes; The perspective transformation of the two-dimensional pixel coordinates is performed to obtain the homography transformation matrix; The obscured area is inversely projected using the homography transformation matrix to generate the complete logo pattern.

8. The method according to claim 7, characterized in that, The perspective transformation of the two-dimensional pixel coordinates is performed to obtain the homography transformation matrix, including: Select at least four valid pattern nodes that are not collinearly distributed; Construct a homogeneous linear system of equations between the two-dimensional pixel coordinates of the effective pattern node and the reference pattern coordinates; The homography transformation matrix is ​​extracted by solving the homogeneous linear equation system using the singular value decomposition algorithm. The theoretical coordinates of the corresponding missing nodes are input into the homography transformation matrix to map and obtain the estimated pixel coordinates. The estimated pixel coordinates are then filled into the pattern matrix to complete the reconstruction.

9. The method of claim 8, wherein, When caching the global space mapping matrix at multiple consecutive time steps, the method further includes: Extract the inertial measurement acceleration sequence of the submersible ROV body equipped with the wrist camera device; The body disturbance displacement sequence is obtained based on the inertial measurement acceleration sequence; The extended Kalman filter algorithm is used to perform Kalman fusion of the body perturbation displacement sequence and the translation components in the time-series matrix queue to compensate for and eliminate the body sway error.

10. The method of claim 9, wherein, After the method involves driving the end-mounting tool to engage the emergency mounting interface of the diving bell, the method further includes: Obtain the real-time feedback hydraulic pressure value of the closing cylinder of the end attachment tool; When the real-time feedback hydraulic pressure value exceeds the no-load closing threshold and remains for a preset duration, it is determined that a rigid locking feedback force is generated. Cut off the pose compensation command and lock the brake of the six-degree-of-freedom robotic arm base.