An underwater robot autonomous inspection system based on binocular vision
By integrating IMU and binocular visual information in the underwater robot system, using the YOLOv8 model for target recognition and depth information extraction, and combining the Minimum Snap algorithm and sliding mode control strategy, the problems of insufficient positioning accuracy and unstable path planning in underwater inspection are solved, and high-precision independent inspection and depth information extraction are achieved.
Patent Information
- Application Number
- CN202510320029.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing underwater inspection technology has problems such as insufficient positioning accuracy, difficulty in extracting depth information, and unstable path planning and control, especially in complex underwater environments, which are difficult to achieve high-precision independent inspection.
The underwater robot autonomous patrol system based on binocular vision is adopted to achieve high-precision positioning by integrating IMU and binocular vision information, target recognition and depth information extraction are used using the YOLOv8 model, path planning is carried out in combination with local 3D features and Minimum Snap algorithm, and precise path tracking and attitude control are achieved through sliding mode control and PID+MPC dual-loop control strategy.
It significantly improves the autonomous inspection capabilities of underwater robots in complex environments, realizes high-precision positioning, deep information extraction and smooth path planning, and provides innovative technical solutions for the health monitoring and maintenance of underwater structures.
Smart Images

Figure CN119861739B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater robots and their inspection technologies, and particularly relates to an autonomous inspection system for underwater robots based on binocular vision. Background Art
[0002] With the development and utilization of global marine resources, the importance of underwater structures in ocean engineering has become increasingly prominent. The construction and maintenance of infrastructure such as bridge piers, offshore wind power platforms, offshore oil platforms, and submarine cables are key tasks in ocean engineering. Traditional underwater inspection methods mainly rely on human divers or remotely operated vehicles (ROVs).
[0003] When human divers conduct underwater inspections, there are many significant drawbacks. First of all, the operation cost remains high. Not only do we need to equip divers with professional and expensive diving equipment, such as high-performance diving suits, accurate depth gauges, and reliable communication devices, but also arrange professional support teams, including surface guardians, equipment maintenance personnel, etc., which undoubtedly greatly increases the labor and material costs.
[0004] At the same time, diving operations face extremely high risks. The underwater environment is complex and changeable, and there may be undercurrents and vortices at any time. If not careful, divers will be carried away by the water flow, endangering their lives. Moreover, divers also face health risks such as decompression sickness. Every dive is like a game with danger.
[0005] Not only that, the operation is also greatly restricted by factors such as water depth, water flow, and light. As the water depth increases, the water pressure surges, posing a huge test to the physical endurance of divers; the turbulent water flow will interfere with the actions of divers, making it difficult for them to stay stable at the target position for operation; and in deep sea areas or extreme environments, the light is weak or even almost dark, seriously affecting the visibility of divers and making it difficult to accurately observe and judge the underwater situation. Therefore, it is difficult for human divers to carry out inspection work in the deep sea or extreme environments for a long time and efficiently.
[0006] Although the use of ROVs can solve the limitations of manual inspection to a certain extent, its operation still relies on manual control, and there are also limitations such as operation depth and communication bandwidth. Especially in complex water quality conditions, the image recognition and positioning accuracy of the ROV system often cannot meet the high-precision inspection requirements.
[0007] In addition, the particularity of the underwater environment, such as low light, strong water flow, high turbidity, etc., poses a severe challenge to the existing vision processing systems. Due to the scattering and absorption of light, underwater images usually show characteristics such as low contrast, blurring, and distortion, resulting in difficulty for monocular vision systems to accurately obtain effective depth information and target recognition results. Although multiocular vision or other sensors can partially make up for this deficiency, their accuracy and real-time performance are still affected by complex environmental conditions.
[0008] Meanwhile, the problems faced by the positioning and navigation of underwater robots cannot be ignored. The dynamic changes in the underwater environment pose great challenges to traditional positioning algorithms in terms of error accumulation and drift, and conventional positioning sensors cannot be used in the underwater environment due to physical property limitations (such as GPS), which forms a technical bottleneck for underwater robots to achieve high-precision autonomous inspection.
[0009] Especially in complex underwater environments, the robot needs to have sufficient autonomy to complete tasks, which usually depends on the precise planning and control of the robot's motion trajectory. Therefore, how to achieve high-precision positioning, depth information extraction, path planning, and autonomous inspection of underwater robots in complex environments through multi-sensor fusion, image processing, and intelligent control technologies has become a technical hot spot and research difficulty in the field of underwater robots. Summary of the Invention
[0010] The present invention provides an autonomous inspection system for an underwater robot based on binocular vision, which can achieve high-precision positioning, depth information extraction, path planning, and autonomous inspection of the underwater robot in a complex environment.
[0011] The present invention provides an autonomous inspection system for an underwater robot based on binocular vision, comprising:
[0012] A positioning module, including a front-end tracking unit and a back-end optimization unit. The front-end tracking unit is used to obtain a predicted initial pose through an IMU, track feature points of the left camera image by the LK optical flow method, then obtain the current frame visual pose through the PnP algorithm, correct the predicted initial pose through the current frame visual pose to obtain key frames, construct a local map of a sliding window with map points obtained by binocular triangulation of the key frames and the right camera image, and map points obtained by binocular triangulation of the key frames and their front and back frames. The back-end optimization unit is used to optimize the key frames, their corresponding map points, and camera poses within the sliding window to obtain the actual pose and speed;
[0013] A perception module, used to achieve target object recognition by semantic segmentation of the left camera image through the YOLOv8 method, match the left camera image and the right camera image using a stereo matching algorithm to obtain a depth map, obtain the depth information of the target based on the depth map and the recognized target object, obtain local point cloud data through the camera internal parameters based on the depth information of the target, calculate the point cloud normal vector and local surface curvature based on the local point cloud data, obtain the geometric features of the target object based on the local surface curvature, and construct the local 3D features of the target object based on the geometric features of the target object;
[0014] A planning module, which is used to estimate an initial search trajectory through a depth gauge and an ultra-short baseline system, screen out trajectory control points based on the actual pose and the local 3D features of the target object, allocate the time consumption of each section of the trajectory, and optimize the initial search trajectory based on the trajectory control points and the time consumed by each section of the trajectory through the Minimum Snap algorithm to obtain the target pose and target speed;
[0015] A control module, which is used to construct attitude, position, and speed errors through the actual pose and speed and the corresponding target pose and target speed, control the attitude of the autonomous cruise system based on the attitude error using a sliding mode control method, and control the position and speed of the autonomous cruise system based on the position and speed errors using a double-loop control structure, so as to realize the autonomous cruise of the autonomous cruise system for underwater structures.
[0016] Preferably, the perception module includes target recognition, target extraction, and feature calculation;
[0017] Among them, the target recognition is used to perform real-time semantic segmentation and target detection on the left camera image using the YOLOv8 model, so as to identify the category and position of the target object;
[0018] The target extraction is used to calculate the depth information of each pixel point based on the left camera image and the right camera image through a stereo matching algorithm, so as to obtain a depth map, and obtain the depth of the target object through the depth map and the category and position of the identified target object;
[0019] The feature calculation is used to obtain local point cloud data based on the depth information of the target object through the camera internal parameters, calculate the point cloud normal vector by least squares fitting based on the local point cloud data, calculate the local surface curvature through the point cloud normal vector, obtain the geometric features of the target object based on the local surface curvature, and construct the local 3D features of the target object based on the geometric features of the target object.
[0020] Preferably, the target extraction is used to calculate the depth information of each pixel point based on the left camera image and the right camera image through a stereo matching algorithm, including:
[0021] After matching the left camera image and the right camera image, the depth values of each point in the image are obtained using the principle of triangulation, so as to construct a three-dimensional point cloud of the underwater environment.
[0022] Preferably, estimating the initial trajectory search based on a depth gauge and an ultra-short baseline system includes:
[0023] Obtain preliminary positioning information through a depth gauge and an ultra-short baseline system, and perform a global path search for the target object based on the positioning information to form an initial search trajectory.
[0024] Preferably, screening out trajectory control points based on the actual pose and the local 3D features of the target object includes:
[0025] During the initial search trajectory process, geometric information of the target object is extracted through surface fitting and feature extraction techniques based on the actual pose and local 3D features of the target object. The geometric information of the target object is analyzed to obtain key structural regions, and trajectory control points passed by the inspection are screened based on the key structural regions.
[0026] Preferably, the initial search trajectory is optimized based on the trajectory control points and the time between the trajectory control points through the Minimum Snap algorithm to obtain the target pose and target speed, including:
[0027] Time is allocated to the path between the trajectory control points. Based on the trajectory control points and the time allocated to each interval between the trajectory control points, the initial search trajectory is optimized through the Minimum Snap algorithm by minimizing the high-order derivatives of the trajectory to obtain the target pose and target speed.
[0028] Preferably, the control module includes a sliding mode controller, a position controller, and a thrust distribution unit;
[0029] Among them, the sliding mode controller is used to obtain the attitude control law through the sliding mode control method based on the attitude error between the current actual attitude and the current target attitude;
[0030] The position controller is used to minimize the position error and speed error at time t + k by adopting a double-loop control structure to obtain the position control law;
[0031] The thrust distribution unit is used to distribute the thrust based on the attitude control law and the position control law to facilitate path tracking and attitude control.
[0032] Preferably, the position controller includes a PID controller and an MPC controller;
[0033] Among them, the PID controller is used to obtain the initial position control law based on the error between the current actual position and the target position;
[0034] The MPC controller is used to optimize the initial position control law by minimizing the position error and speed error at future times to obtain the position control law.
[0035] Preferably, it further includes a reconstruction module. The reconstruction module is used to obtain local point cloud data, key frames, and corresponding actual poses, align the key frames and corresponding actual poses with the local point cloud data, and then optimize the aligned point cloud data through the TSDF algorithm to obtain a three-dimensional model point cloud. Based on the three-dimensional model point cloud, a target object model is rendered through 3DGS rendering technology.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] Through multi-sensor fusion, deep learning, path optimization, and precise control technologies, the present invention significantly enhances the autonomous inspection ability of underwater robots in complex environments, achieving high-precision positioning, depth information extraction, and smooth path planning, providing an innovative technical solution for the health monitoring and maintenance of underwater structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Schematic diagram of autonomous inspection of an underwater robot with an underwater pile foundation as the target structure provided by a specific embodiment of the present invention;
[0039] Figure 2 Block diagram of the autonomous inspection system of an underwater robot provided by a specific embodiment of the present invention;
[0040] Figure 3 Schematic diagram of the structure of the underwater binocular vision module provided by a specific embodiment of the present invention;
[0041] Figure 4 Schematic diagram of the tight coupling optimization of the positioning module provided by a specific embodiment of the present invention;
[0042] Figure 5 Schematic diagram of local trajectory optimization provided by a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0044] The present invention proposes an autonomous inspection system for underwater robots based on binocular vision, aiming to solve problems such as insufficient positioning accuracy, difficulty in extracting depth information, and unstable path planning and control in existing underwater inspection technologies. The system realizes high-precision positioning and state estimation by fusing inertial measurement unit (IMU) and binocular vision information, uses the YOLOv8 model for real-time target recognition and depth information extraction of underwater structures, and estimates the morphological characteristics of structures by combining local point cloud analysis. For the optimization of the underwater robot's motion trajectory, the system generates a smooth trajectory based on the 3D feature information of underwater structures using the MinimumSnap algorithm, and realizes precise path tracking and attitude stability through a sliding mode controller and a PID+MPC control strategy. In addition, the system optimizes the point cloud data using the TSDF algorithm by fusing positioning and perception data, and performs high-precision 3D reconstruction in combination with 3DGS rendering technology.
[0045] Please refer toFigure 1 , Figure 1 This is an exemplary embodiment of the present invention, namely, a schematic diagram of the autonomous inspection of an underwater robot with an underwater pile foundation as the target structure. The process uses an underwater binocular vision module as shown in Figure 3 .
[0046] A specific embodiment of the present invention provides an autonomous inspection system for an underwater robot based on binocular vision. As shown in Figure 2 , it includes a positioning module, a perception module, a planning module, a control module, and a reconstruction module.
[0047] Among them, the positioning module provided by the specific embodiment of the present invention fuses the IMU and binocular vision information through tight coupling (tight coupling: jointly optimizing the observations and state estimates of multiple sensors; the loose coupling method processes the data of each sensor separately and then combines and optimizes the results) to achieve high-precision positioning of the underwater robot. The positioning module includes a front-end tracking unit and a back-end optimization unit. The front-end tracking unit uses the IMU to predict the initial pose through the pre-integration method to obtain the predicted initial pose, that is, the preliminary motion estimation. The purpose of pre-integration is to integrate the acceleration and angular velocity information measured by the IMU between two key frames into the motion change amount within a period of time, and then calculate the position, velocity, and rotation information of the system. This process can greatly reduce the computational amount, effectively reduce the errors caused by IMU data noise and bias, and provide a more accurate initial estimate for the subsequent fusion of vision and IMU information.
[0048] In a specific embodiment, the specific steps for obtaining the predicted initial pose include: bringing the acceleration and angular velocity measured by the IMU between the i-th and j-th moments into the pre-integration model to obtain the position, velocity, and rotation information of the underwater robot, which is expressed as follows: , and defining the pre-integration quantities of position, velocity, and angle as follows: , where q represents the quaternion rotation, represents the rotation angle, a represents the acceleration, represents the gravitational acceleration in the world coordinate system, t is the current moment, p represents the position, v represents the velocity, and q represents the rotation.
[0049] At the same time, the specific embodiment of the present invention uses the LK optical flow method to track the feature points of the left camera image, uses the PnP algorithm to solve the current frame pose, and fuses the visual pose solved by the PnP algorithm and the predicted initial pose predicted by the IMU using the Kalman filter to optimize the pose while screening out the key frames.
[0050] In a specific embodiment, the method for screening key frames includes: 1. Parallax: When the parallax between the current frame and the previous key frame is large, it is selected as a key frame; 2. Number of feature points: When the number of feature points tracked in the current frame is small, it is selected as a key frame; 3. Time interval: When the time interval from the previous key frame is long, it is selected as a key frame.
[0051] In the specific embodiment of the present invention, the map points obtained by binocular triangulation using the key frames and the right camera images, and the map points obtained by binocular triangulation of the key frames with their front and rear frames are used to construct the local map of the sliding window. The selected key frames are triangulated with their corresponding right images. At the same time, the key frames are also triangulated with their front and rear frame images. Through the triangulation process, the corresponding 3D map points can be obtained from the 2D feature points in the key frames, and these map points provide a preliminary representation of the 3D structure of the local environment. To improve the positioning accuracy and the quality of the map, we use the sliding window technique to optimize the states of these map points and key frames. Specifically, the size of the sliding window is set to 10 frames, that is, the system only focuses on the latest 10 frames of images and the relevant map information. As new frames are added, the sliding window will remove the earliest frame from the window, and the latest frame and map points will be added to the optimization process. Within the sliding window, all the key frames and their corresponding map points and camera poses jointly participate in the optimization. By minimizing objective functions such as the visual reprojection error and the IMU residuals, the data within the sliding window is continuously optimized, thereby improving the accuracy of the pose estimation and map points. The information such as the position, velocity, acceleration, angular velocity, and pre-integration quantities of the IMU within the sliding window and the information such as the position, pose, 2D feature points, and 3D map points of the camera are jointly modeled to form an overall optimization problem. The L-M optimization method is used to perform BA (Bundle Adjustment) tight-coupling optimization on the above system, and the tight-coupling strategy is as Figure 4 shown. The constructed least squares optimization objective function is as follows: , where is the IMU residual, is the camera reprojection error. The IMU residual refers to the error between the pose estimated by the IMU pre-integration and the pose solved visually; the reprojection error is the error between the point obtained by projecting the map point in the 3D space onto the camera image plane and the corresponding feature point actually observed in the image.
[0052] After the back-end optimization provided by the specific embodiment of the present invention, the position, pose, and velocity of the underwater robot are obtained, and the position and pose are returned to the IMU pre-integration process as the initial values to predict the pose of the next stage.
[0053] The perception module provided by the specific embodiment of the present invention is used to extract the depth information of underwater structures and identify targets through the YOLOv8 method and the stereo matching algorithm. The YOLOv8 method is used to perform semantic segmentation on the left camera image to identify target objects. Through this identification method, it can effectively cope with challenges such as low contrast, blurring, and insufficient lighting in complex underwater environments, ensuring the efficient identification and classification of structures.
[0054] In one embodiment, the convolutional neural network (CNN) of YOLOv8 is sampled to extract features from the input underwater image, and output the class label, bounding box, and its position of the structure, so as to achieve accurate target positioning.
[0055] The specific embodiment of the present invention uses the stereo matching algorithm to match the left camera image and the right camera image to obtain the depth information of each pixel point, thereby obtaining a depth map. In one specific embodiment, by matching the left and right images, the depth values of each point in the image are obtained using the principle of triangulation, so as to construct a three-dimensional point cloud of the underwater environment. The specific depth information extraction can be achieved through the following formula: , where Z is the depth of the pixel point, f is the focal length of the camera, B is the baseline distance between the two cameras, and d is the disparity of the corresponding point in the image. By obtaining accurate depth information of underwater structures, the accuracy of underwater visual perception is further enhanced.
[0056] After obtaining the local point cloud data of the structure, the specific embodiment of the present invention further calculates the point cloud normal vector and surface curvature to estimate the morphological characteristics of the underwater structure.
[0057] Specifically, the normal vector calculation is based on the neighborhood points of the set local point cloud, and the normal vector of the point is calculated by the least squares fitting: where, , is the point cloud normal vector, is the coordinate of the local neighborhood point, is the weight of each point, is the center of the point cloud. Then, the curvature of the local surface is calculated through the normal vector information to further evaluate the geometric characteristics and surface state of the underwater structure. Through this series of depth information extraction, point cloud calculation, and geometric feature estimation, the present invention can accurately obtain the three-dimensional morphological characteristics of underwater structures, providing strong support for subsequent underwater robot trajectory planning and target structure reconstruction.
[0058] The planning module provided by the specific embodiment of the present invention is used to search for and optimize the inspection path based on the local 3D features of the underwater structure: This planning module is used to estimate the initial search trajectory through a depth gauge and an ultra-short baseline system (USBL), and based on the actual pose and the actual pose of the local 3D features of the target object, optimize the initial search trajectory through the Minimum Snap algorithm based on the trajectory control points and the time between the trajectory control points to obtain the target pose and the target speed.
[0059] In a specific embodiment, the embodiment of the present invention estimates the initial trajectory search based on a depth gauge and an ultra-short baseline system, including:
[0060] The embodiment of the present invention obtains preliminary positioning information through a depth gauge and an ultra-short baseline system, and conducts a global path search for the target object based on the positioning information to form an initial search trajectory.
[0061] In a specific embodiment, the embodiment of the present invention screens out trajectory control points based on the actual pose and the local 3D features of the target object, including:
[0062] In the process of the initial search trajectory of the embodiment of the present invention, the geometric information of the target object is extracted through surface fitting and feature extraction techniques based on the local 3D features of the target object. The geometric information includes extracting the surface shape, curvature, edges, etc. of the target object from the local point cloud information. By analyzing the geometric information of the target object, the underwater robot can obtain the key structural areas, and based on the key structural areas, the trajectory control points passed by the inspection are screened out.
[0063] The underwater robot provided by the specific embodiment of the present invention uses a depth gauge and an Ultra-Short Baseline System (USBL) to obtain preliminary positioning information. Based on this, a global path search is carried out for the underwater structure to be inspected, and a preliminary inspection trajectory is formed. The global path planning process takes into account the complexity of the underwater environment and the relative positions of the underwater structures to ensure that the robot can effectively cover the key areas of the target structure. Subsequently, combined with the local point cloud information of the target structure, 3D features are extracted to screen out trajectory control points. Specifically, through surface fitting and feature extraction techniques, the local point cloud obtains geometric information such as the shape, curvature, and edges of the surface of the underwater structure. Based on these features, areas with high curvature, complex shapes, or edge features are selected as trajectory control points for inspection. During the screening process, in addition to considering geometric features, the safety distance between the robot and the target object also needs to be considered to ensure that the robot does not collide with the target object during the inspection process. Therefore, the selection of control points needs to cover important areas and ensure safety, thereby optimizing the inspection path and improving the inspection efficiency and accuracy. After screening out the control points, time allocation needs to be carried out for each sub-trajectory formed by separating these control points. The time allocation strategy is as follows: 1. Allocate time based on regional importance: Allocate more time for detailed inspection of areas with high curvature, complex surfaces, or potential failure areas (such as cracks, joints, etc.); allocate less time for relatively flat and simple areas; 2. Allocate time based on the length of the trajectory segment: Allocate more time to longer sub-segments of the trajectory and less time to shorter sub-segments to maintain the balance of the overall path; 3. Allocate time based on the safety distance and speed limit: When the robot approaches a complex area of the target object, in order to ensure safety, reduce the speed and appropriately increase the time to ensure the smooth operation of the trajectory; 4. Allocate time based on the maximum acceleration and deceleration limits: According to the maximum acceleration and deceleration limits of the robot, reasonably arrange the time for the acceleration and deceleration processes. Sharp acceleration and deceleration will affect the stability and accuracy of the robot, so more time is needed for a smooth transition.
[0064] The embodiment of the present invention takes into account the dynamic constraints of the underwater robot during the movement process and uses the Minimum Snap algorithm to optimize the local path. The optimization of the local trajectory is shown in Figure 5As shown. The Minimum Snap algorithm generates a smooth and dynamically constrained trajectory by minimizing the high-order derivatives of the trajectory, thus avoiding drastic changes and instabilities in the robot's motion. This algorithm can reduce energy consumption and mechanical load during motion while maintaining motion smoothness, effectively enhancing the mobility and stability of the robot in complex underwater environments. The optimized trajectory not only satisfies the dynamic constraints of the robot but also enables the underwater robot to perform inspection tasks more smoothly in narrow underwater spaces through smoothing. The finally generated smooth trajectory provides a feasible and efficient inspection path for the underwater robot, ensuring that the robot can complete a comprehensive inspection of underwater structures with a stable attitude and precise control.
[0065] The control module provided by the specific embodiment of the present invention realizes the precise path tracking control of the underwater robot based on the planned local target trajectory and the positioning feedback. The planning module calculates the expected position, attitude, and velocity information at the current moment according to the local target trajectory. These expected values constitute the goals for the control module to perform path tracking, guiding the robot to perform inspection tasks along the predetermined trajectory during actual operation. At the same time, with the help of the positioning module mentioned in the first step, the robot can obtain the current actual position, attitude, and velocity information in real time, so as to compare with the expected values to achieve precise trajectory tracking. In the control system, the errors between the actual position, attitude, and velocity information and the expected values are transmitted to the control module as input signals.
[0066] The control module provided by the specific embodiment of the present invention consists of two parts, including an attitude controller and a position controller. The attitude controller adopts a sliding mode control method to achieve the attitude stability of the robot. The sliding mode controller can effectively cope with model uncertainties and external disturbances in the system, ensuring that the underwater robot can accurately maintain its attitude and remain stable in complex water flow environments.
[0067] The goal of the sliding mode control provided by the specific embodiment of the present invention is to make the system state tend to the sliding mode surface and stay on this surface. The control law can be expressed as: , , where is the output of the sliding mode controller, the forces required for the movement of each axis of the underwater robot; is the sliding mode surface function, which is a function of the system state error; is the attitude error, is the change rate of the error, and are the design parameters of the sliding mode surface function; k is the control gain; is the sign function, which is used to push the system state into the sliding mode surface. Through this control law, the sliding mode controller can make the underwater robot quickly and accurately adjust the attitude error to approach zero.
[0068] The position controller provided by the specific embodiment of the present invention adopts a double-loop control structure, where the inner loop is a PID controller based on position error, and the outer loop is a model predictive control (MPC) controller based on speed. The PID controller adjusts the control input by calculating the error between the current position and the desired position. The initial position control law can be expressed as: , where is the output of the PID controller, that is, the initial position control rate, which is specifically an initial control signal of the MPC controller as the inner loop of the position controller; is the current position error; , , are the proportional, integral, and differential gains respectively. The PID controller reduces the position error and smoothly guides the robot to the target position by adjusting the control input.
[0069] The MPC controller provided by the specific embodiment of the present invention optimizes the future state of the robot according to the prediction model and outputs the optimal control input. The goal of MPC is to obtain the best control strategy by minimizing the future prediction error and constraint conditions. Its optimization problem can be expressed as: , where is the predicted value of the position and speed state of the robot at time , that is, the future time, which is obtained by substituting the current position and speed state and control input of the underwater robot into the kinematic and dynamic models of the robot for iteration; is the target state at the future time; by continuously optimizing and iteratively solving this sequence, the output of the MPC controller is finally obtained, which is the force required for the movement of each axis for the underwater robot specifically. Among them, the output of the PID controller is used as the initial sequence before optimization iteration; is the regularization factor of the control input; N is the length of the prediction horizon. By optimizing this objective function, MPC can generate the optimal control input according to the current state and future target of the robot at the current moment, ensuring the minimum trajectory tracking error of the robot.
[0070] Through the collaborative work of the above two controllers in the specific embodiment of the present invention, the underwater robot can achieve precise path tracking and attitude control. The sliding mode controller ensures the rapid stability of the robot's attitude, while the combination of the PID and MPC controllers guarantees the accuracy and smoothness of position tracking. The combination of the sliding mode control and the PID+MPC double-loop control strategy effectively reduces the error accumulation and dynamic instability of the robot in the complex underwater environment, thereby improving the stability and efficiency of the inspection task.
[0071] The reconstruction module provided by the specific embodiment of the present invention fuses the relevant data of the positioning module and the perception module, and performs high-quality three-dimensional reconstruction on the underwater target structure through 3DGS rendering technology.
[0072] The perception module provided by the specific embodiment of the present invention obtains the three-dimensional depth data of the underwater structure in real time through depth information extraction technology and generates local point clouds. At the same time, the positioning module records the key frames and the corresponding camera pose information, including the rotation matrix R and the translation vector t of the camera, which helps to accurately align the point cloud and image data under different perspectives.
[0073] In order to achieve accurate three-dimensional reconstruction in the specific embodiment of the present invention, the corresponding relationship between point clouds is first established through point cloud matching and key frame pose alignment. In this process, standard point cloud registration technology is adopted to ensure the accurate alignment of point clouds from different perspectives by minimizing the reprojection error. Next, the Truncated Signed Distance Function (TSDF) algorithm is used to optimize the point cloud. The TSDF method maps the surface of each point cloud to a distance field and updates it weighted according to the distance from the point to the surface, obtaining a dense depth field representation. Specifically, for each three-dimensional point P, the TSDF value d(P) is obtained by calculating the distance from this point to the nearest surface For each update of the distance function, it can be weighted by the following formula: , is the newly added point cloud data, is the value in the current distance field, is the weighting factor. By repeatedly updating the distance values of all points, the TSDF algorithm can gradually eliminate noise and optimize the point cloud, generating a high-quality underwater three-dimensional model point cloud.
[0074] Finally, the 3DGS (3D Graphics System) rendering technology is used to perform three-dimensional reconstruction on the optimized point cloud to form an accurate visualization model of the underwater structure. The 3DGS technology can map the optimized three-dimensional point cloud data to the view through per-pixel rendering calculations, and perform realistic rendering in combination with the underwater lighting model, material model, etc. The light intensity of each rendered pixel can be calculated by the following formula: , where, I is the light intensity of the pixel, is the lighting model, P is the position of the point in three-dimensional space, L is the light source direction, V is the viewing direction, Tis the distance of light propagation. Through precise lighting calculations and efficient rendering techniques, 3DGS can generate high-quality three-dimensional models of underwater structures, providing clear visual effects to assist in the in-depth analysis and evaluation of structures during underwater inspection tasks.
[0075] In summary, the present invention provides an autonomous inspection system and method for an underwater robot based on binocular vision, aiming to solve problems such as insufficient accuracy, low efficiency, high difficulty of manual operation, and difficult guarantee of safety in traditional underwater inspection methods.
[0076] First, by fusing IMU and binocular vision information and adopting a tightly coupled method to achieve high-precision positioning of the underwater robot, and improving the positioning accuracy through an optimization algorithm to provide reliable pose information for subsequent tasks.
[0077] Second, based on YOLOv8 semantic segmentation and stereo matching techniques, the system can perform real-time target recognition and depth information extraction of underwater structures, accurately estimating the morphological characteristics of the structures.
[0078] In the third step, the system precisely plans the inspection path by combining local 3D features and an optimization algorithm (such as the Minimum Snap algorithm) to generate a smooth trajectory that meets dynamic constraints, ensuring that the robot can efficiently and stably complete the inspection task.
[0079] In the fourth step, a sliding mode control and a PID+MPC double-loop control strategy are adopted to achieve path tracking and attitude control of the underwater robot, ensuring the stability and accuracy of the underwater robot in a complex environment.
[0080] Finally, the system fuses the data of the positioning and perception modules, optimizes the point cloud through the TSDF algorithm, and realizes high-quality three-dimensional reconstruction of the underwater target structure by means of 3DGS rendering technology, providing accurate three-dimensional data support for subsequent defect detection and health assessment.
[0081] The above content is only the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the claims of the present invention.
Claims
1. An autonomous inspection system for underwater robots based on binocular vision, characterized in that: include: A positioning module includes a front-end tracking unit and a back-end optimization unit, wherein the front-end tracking unit is used to obtain a predicted initial posture through an IMU, track feature points of the left camera image through an LK optical flow method, and then obtain the visual posture of the current frame through a PnP algorithm, obtain a key frame by correcting the predicted initial posture through the visual posture of the current frame, and construct a local map of the sliding window through map points obtained by binocular triangulation of the key frame and the right camera image, and map points obtained by binocular triangulation of the key frame and its previous and subsequent frames, and the back-end optimization unit is used to optimize the key frame in the sliding window and its corresponding map points and camera posture to obtain the actual posture and speed; A perception module is used to perform semantic segmentation on the camera left image through the YOLOv8 method to realize target object recognition, match the camera left image and the camera right image using a stereo matching algorithm to obtain a depth map, obtain the depth information of the target based on the depth map and the recognized target, obtain local point cloud data through the camera intrinsic parameter based on the depth information of the target, calculate the point cloud normal vector and local surface curvature based on the local point cloud data, obtain the geometric features of the target based on the local surface curvature, and construct the local 3D features of the target based on the geometric features of the target; The planning module is used to estimate the initial search trajectory through the depth meter and the ultra-short baseline system, select the trajectory control points based on the actual posture and local 3D features of the target object, and allocate the time consumption of each trajectory segment. The initial search trajectory is optimized based on the trajectory control points and the time consumed by each trajectory segment through the Minimum Snap algorithm to obtain the target posture and target speed; The control module is used to construct attitude, position and velocity errors through actual posture and velocity and corresponding target posture and velocity, control the attitude of the autonomous cruise system by using a sliding mode control method based on the attitude error, and control the position and velocity of the autonomous cruise system by using a double-loop control structure based on the position and velocity errors, thereby realizing autonomous cruising of the autonomous cruise system on underwater structures.
2. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: The perception module includes target recognition, target extraction and feature calculation; The target recognition is used to perform real-time semantic segmentation and target detection on the left camera image using the YOLOv8 model, so as to identify the category and position of the target object; The target extraction is used to calculate the depth information of each pixel through a stereo matching algorithm based on the camera left image and the camera right image, thereby obtaining a depth map, and obtaining the depth of the target object through the depth map and the category and position of the identified target object; The feature calculation is used to obtain local point cloud data through camera intrinsic parameters based on the depth information of the target object, calculate the point cloud normal vector through least squares fitting based on the local point cloud data, obtain the local surface curvature through the point cloud normal vector calculation, obtain the geometric features of the target object based on the local surface curvature, and construct the local 3D features of the target object based on the geometric features of the target object.
3. The underwater robot autonomous inspection system based on binocular vision according to claim 2 is characterized in that: The target extraction is used to calculate the depth information of each pixel point based on the camera left image and the camera right image through a stereo matching algorithm, including: After matching the camera left image with the camera right image, the depth value of each point in the image is obtained using the triangulation principle to construct a three-dimensional point cloud of the underwater environment.
4. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: Initial trajectory search based on depth meter and ultra-short baseline system estimation, including: Preliminary positioning information is obtained through the depth meter and ultra-short baseline system, and a global path search is performed for the target object based on the positioning information to form an initial search trajectory.
5. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: The trajectory control points are selected based on the actual posture and local 3D features of the target object, including: In the initial search trajectory process, the geometric information of the target object is extracted through surface fitting and feature extraction technology based on the actual posture and local 3D features of the target object. The geometric information of the target object is analyzed to obtain the key structural area, and the trajectory control points passed by the inspection are screened out based on the key structural area.
6. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: The Minimum Snap algorithm is used to optimize the initial search trajectory based on the trajectory control points and the time between the trajectory control points to obtain the target pose and target speed, including: Time is allocated to the path between trajectory control points. Based on the trajectory control points and the time allocated to each trajectory control point, the initial search trajectory is optimized by minimizing the high-order derivatives of the trajectory through the Minimum Snap algorithm to obtain the target posture and target velocity.
7. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: The control module includes a sliding mode controller, a position controller and a thrust distribution unit; Wherein, the sliding mode controller is used to obtain the attitude control law through the sliding mode control method based on the attitude error between the current actual attitude and the current target attitude; The position controller is used to obtain a position control law by minimizing the position error and speed error at time t+k using a dual-loop control structure; The thrust distribution unit is used to distribute the thrust based on the attitude control law and the position control law, so as to realize path tracking and attitude control.
8. The underwater robot autonomous inspection system based on binocular vision according to claim 7 is characterized in that: The position controller includes a PID controller and an MPC controller; Wherein, the PID controller is used to obtain an initial position control law based on the error between the current actual position and the target position; The MPC controller is used to minimize the position error and speed error at future moments to optimize the initial position control law to obtain the position control law.
9. The underwater robot autonomous inspection system based on binocular vision according to claim 1 is characterized in that: It also includes a reconstruction module, which is used to obtain local point cloud data, key frames and corresponding actual postures, align the key frames and the corresponding actual postures with the local point cloud data, and then use the TSDF algorithm to optimize the aligned point cloud data to obtain a three-dimensional model point cloud, and render the target object model based on the three-dimensional model point cloud using 3DGS rendering technology.
Citation Information
Patent Citations
Binocular vision navigation system and method based on inspection robot in transformer substation
CN103400392A
Building inspection robot navigation method and system based on non-prior map
CN118758287A