Visual positioning method, device and equipment for underwater robot and storage medium
By combining binocular cameras and neural network models, low-cost and precise positioning of underwater robots is achieved through visual positioning, which solves the problems of expensive sonar positioning and sparse data in underwater environments and improves the accuracy and efficiency of positioning.
Patent Information
- Application Number
- CN202511221279.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing underwater robot positioning systems rely on sonar positioning, which is expensive and data-sparse, making it difficult to achieve cost-effective and precise positioning in underwater environments.
A binocular camera is used to acquire original images, and a dense point cloud map is constructed through depth estimation and dynamic object detection models. The pose estimation of the underwater robot is achieved by combining simultaneous positioning and mapping algorithms. A neural network model is used to process underwater images, remove moving targets, and improve positioning accuracy.
It achieves precise underwater close-range positioning at low cost, improves the accuracy of visual positioning, reduces the interference of underwater fish dynamic objects on positioning, and enhances image quality and positioning accuracy.
Smart Images

Figure CN120707613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a visual positioning method, device, equipment and storage medium for an underwater robot. Background Art
[0002] With the growing demand for ocean exploration and underwater operations, underwater robotics technology is developing rapidly. Existing technologies primarily rely on sonar and other methods for positioning and tracking. However, these systems are expensive, have sparse data, and are infrequently updated. Visual positioning systems offer high cost-effectiveness and precision at close range. However, given the complex underwater environment, implementing visual positioning that adapts to specific underwater scenarios is a pressing challenge. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for underwater robot visual positioning, which can achieve accurate underwater visual positioning at low cost. The specific scheme is as follows:
[0004] In a first aspect, the present application discloses a visual positioning method for an underwater robot, comprising: Obtaining an original image captured by a camera, and preprocessing the original image to obtain a processed image; Determining a target depth map corresponding to the processed image using depth estimation; Locating and removing moving objects from the processed image using a dynamic object detection model, thereby obtaining a static scene image; the dynamic object detection model is a neural network model pre-trained using underwater images; Based on the target depth map and the static scene image, the pose of the underwater robot is estimated in real time through a synchronous positioning and mapping algorithm, a dense point cloud map is constructed, and the pose map of the underwater robot is updated.
[0005] Optionally, the camera is a binocular camera, and the original image includes a binocular image; Determining a target depth map corresponding to the processed image using depth estimation includes: Input any monocular image in the original image into a monocular depth estimation system to obtain a monocular depth map corresponding to the original image; Inputting the binocular image into a binocular depth estimation model to obtain a binocular depth map corresponding to the original image; the monocular depth estimation system and the binocular depth estimation model are both pre-trained using an underwater image dataset; The monocular depth map and the binocular depth map are compared to determine the target depth map.
[0006] Optionally, the monocular depth estimation system includes a monocular relative depth estimation model and a monocular absolute depth estimation model in sequence, and the input of the monocular absolute depth estimation model is the output of the monocular relative depth estimation model; The monocular relative depth estimation model is used to output a corresponding relative depth map according to the input monocular image, and the monocular absolute depth estimation model is used to output a corresponding absolute depth map according to the input relative depth map.
[0007] Optionally, comparing the monocular depth map and the binocular depth map to determine the target depth map includes: The monocular depth map and the binocular depth map are subjected to quality analysis according to a depth map evaluation index, and a depth map with high quality is used as the target depth map.
[0008] Optionally, the preprocessing includes image resizing, image enhancement and image correction. Optionally, the image correction process for the original image includes: The radial distortion coefficient and the tangential distortion coefficient obtained by camera calibration are used to perform distortion correction on the original image.
[0009] Optionally, estimating the pose of the underwater robot in real time by a simultaneous positioning and mapping algorithm based on the target depth map and the static scene image, constructing a dense point cloud map, and updating the pose graph of the underwater robot includes: Extracting feature points from the current static scene image and performing pose estimation based on the target depth map and the pose of the previous static scene image to obtain the pose of the current static scene image; Generate a single-frame point cloud according to the target depth map, filter the single-frame point cloud and fuse it into a dense point cloud map; Determining whether the current static scene image is a key frame, and if so, performing local bundling adjustment optimization based on the current static scene image;
[0010] Detect whether a loop occurs according to the current static scene image, and if so, perform pose graph optimization according to the current static scene image.
[0011] Optionally, estimating the position and pose of the underwater robot in real time by a simultaneous positioning and mapping algorithm based on the target depth map and the static scene image includes: Based on the target depth map and the static scene image, as well as the sensor data collected by the target sensor, the position and posture of the underwater robot are estimated in real time through a synchronous positioning and mapping algorithm; the target sensor includes an inertial measurement unit and / or sonar.
[0012] In another aspect, the present application discloses a visual positioning device for an underwater robot, comprising: An image preprocessing module is used to obtain the original image captured by the camera and preprocess the original image to obtain a processed image; a depth estimation module, configured to determine a target depth map corresponding to the processed image using depth estimation; A dynamic object detection module, configured to locate and remove moving objects from the processed image using a dynamic object detection model to obtain a static scene image; the dynamic object detection model is a neural network model pre-trained using underwater images; The visual positioning module is used to estimate the position and posture of the underwater robot in real time based on the target depth map and the static scene image through a synchronous positioning and mapping algorithm, construct a dense point cloud map, and update the position and posture graph of the underwater robot.
[0013] On the other hand, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program implements the aforementioned underwater robot visual positioning method when executed by a processor.
[0014] On the other hand, the present application discloses a computer program product, including a computer program, which implements the aforementioned underwater robot visual positioning method when executed by a processor.
[0015] In another aspect, the present application discloses an underwater robot visual positioning system, comprising: Binocular digital camera, used to collect original images; A computing unit in communication with the binocular digital camera, used to implement the aforementioned underwater robot visual positioning method;
[0016] An unmanned remotely operated underwater vehicle host computer is communicatively connected to the computing unit.
[0017] Optionally, the computing unit includes: an AI acceleration processor, an image processing unit, and an underwater power supply; The AI acceleration processor is connected to the host of the unmanned remotely operated underwater vehicle via Ethernet, and the AI acceleration processor is also connected to the left camera and the right camera of the binocular digital camera respectively, for receiving images captured by the left camera and the right camera; The image processing unit is connected to the AI acceleration processor, the left camera, and the right camera respectively, and is used to receive a shooting instruction issued by the AI acceleration processor, and synchronously trigger the left camera and the right camera through an IO hard trigger according to the shooting instruction; The underwater power supply is connected to the AI acceleration processor for supplying power to the AI acceleration processor.
[0018] In another aspect, the present application discloses an underwater robot, comprising a sealed pressure chamber and the aforementioned underwater robot visual positioning system; the underwater robot visual positioning system comprises a binocular digital camera and a computing unit communicatively connected to the binocular digital camera; The underwater robot visual positioning system is inside the sealed pressure chamber.
[0019] Optionally, the left camera of the binocular digital camera is inside the first circular sealed pressure-bearing cavity, the right camera of the binocular digital camera is inside the second circular sealed pressure-bearing cavity, and the computing unit is inside the third circular sealed pressure-bearing cavity.
[0020] In this application, the original image captured by the camera is obtained, and the original image is pre-processed to obtain a processed image; depth estimation is used to determine the target depth map corresponding to the processed image; a dynamic object detection model is used to locate and remove moving targets from the processed image to obtain a static scene image; the dynamic object detection model is a neural network model pre-trained using underwater images; based on the target depth map and the static scene image, the underwater robot's posture is estimated in real time using a synchronous positioning and mapping algorithm, and a dense point cloud map is constructed to update the underwater robot's posture map. It can be seen that by performing visual positioning through a synchronous positioning and mapping algorithm, accurate underwater close-range positioning can be achieved at a low cost; and the collected original image is pre-processed before subsequent calculations are performed to improve image quality and lay the foundation for subsequent visual positioning; in addition, for special underwater environments, moving targets in the image are deleted before visual positioning to avoid interference from dynamic objects such as underwater fish that affect the positioning effect, thereby improving the accuracy of visual positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work. Figure 1 A flow chart of a visual positioning method for an underwater robot provided in this application; Figure 2 A schematic diagram of a specific underwater robot depth estimation provided in this application; Figure 3 A specific underwater robot visual positioning flow chart provided for this application; Figure 4 A schematic diagram of a specific underwater robot visual positioning system provided in this application; Figure 5A schematic diagram of a circular sealed pressure-bearing cavity corresponding to a specific binocular camera provided in this application; Figure 6 A schematic diagram of the interior of a specific circular sealed pressure-bearing chamber provided in this application; Figure 7 A schematic diagram of a circular sealed pressure-bearing cavity corresponding to a specific computer provided in this application; Figure 8 This is a schematic diagram of the structure of an underwater robot visual positioning device provided in this application. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0023] Existing technologies primarily rely on sonar and other methods for positioning and tracking. However, sonar-based positioning systems are expensive, have sparse data, and are infrequently updated. Visual positioning systems offer high cost-effectiveness and precision for close-range positioning. However, given the complex underwater environment, implementing visual positioning tailored to specific underwater scenarios is a pressing challenge. To overcome these technical challenges, this application proposes a visual positioning method for underwater robots that achieves accurate underwater visual positioning at a low cost.
[0024] The present application discloses a method for visual positioning of an underwater robot. Figure 1 As shown, the method may include the following steps:
[0025] Step S11: obtaining an original image captured by a camera, and preprocessing the original image to obtain a processed image.
[0026] The underwater robot's camera captures raw underwater images, which are then preprocessed to produce processed images. This preprocessing includes image resizing, image enhancement, and image correction. Understandably, in underwater environments, low contrast makes it difficult to extract features from images, necessitating image preprocessing. This primarily involves image enhancement, image correction, and underwater dynamic target detection. Image enhancement primarily enhances image contrast and brightness, improving image quality for subsequent processing. Image correction primarily corrects image distortion, such as camera distortion, to improve reconstruction accuracy.
[0027] In a specific implementation, the image correction process of the original image may include: using the radial distortion coefficient and the tangential distortion coefficient obtained by camera calibration to perform distortion correction on the original image, that is, in practice, the camera is affected by the lens and causes imaging distortion.
[0028] In order to remove distortion, radial distortion and tangential distortion are used to correct camera distortion. Radial distortion is caused by the shape of the lens and manifests as distortion at the edge of the image (such as barrel distortion and pincushion distortion). Tangential distortion is caused by the non-parallelism between the lens and the imaging plane and manifests as tilted deformation of the image. The distortion model formula is: ; in, are the horizontal and vertical coordinates of the undistorted image (original image), is the horizontal and vertical coordinates of the distorted image, and the polar coordinates are expressed as Where k1, k2, and k3 are radial distortion coefficients, and p1 and p2 are tangential distortion coefficients. Distortion parameters are obtained by capturing an image using a checkerboard calibration plate. Based on the positions of the known geometric feature points on the checkerboard calibration plate in the image and their coordinate relationships in the real world, the distortion model parameters are solved using the distortion model formula. Using the known distortion model and parameters, the correct position of each pixel in the dedistorted image captured by the corresponding camera can be calculated.
[0029] Step S12: Determine a target depth map corresponding to the processed image using depth estimation.
[0030] After obtaining the processed image, depth estimation is performed on the processed image to obtain a target depth map corresponding to the processed image.
[0031] In some embodiments, the camera is a binocular camera, and the original image includes a binocular image; using depth estimation to determine the target depth map corresponding to the processed image can include: inputting any monocular image in the original image into a monocular depth estimation system to obtain a monocular depth map corresponding to the original image; inputting the binocular image into a binocular depth estimation model to obtain a binocular depth map corresponding to the original image; the monocular depth estimation system and the binocular depth estimation model are both pre-trained using an underwater image dataset; and comparing the monocular depth map with the binocular depth map to determine the target depth map. It can be seen that this embodiment simultaneously performs monocular depth estimation and binocular depth estimation to obtain a monocular depth map and a binocular depth map, and then compares the monocular depth map and the binocular depth map to select the depth map with better quality as the target depth map. Alternatively, any monocular image in the original image can be selected to calculate its monocular depth map, or two monocular images in the original image can be simultaneously calculated, and then the left depth map, the right depth map, and the binocular depth map are compared to select the target depth map.
[0032] In some embodiments, the monocular depth estimation system includes a monocular relative depth estimation model and a monocular absolute depth estimation model in sequence, wherein the input of the monocular absolute depth estimation model is the output of the monocular relative depth estimation model; the monocular relative depth estimation model is used to output a corresponding relative depth map according to the input monocular image, and the monocular absolute depth estimation model is used to output a corresponding absolute depth map according to the input relative depth map. For example Figure 2 As shown, depth estimation is divided into monocular depth estimation and binocular depth estimation. Both depth estimation methods first train models using underwater image datasets to obtain the corresponding offline depth estimation models. The depth estimation models are then used to perform online depth estimation on the camera's real-time images. The monocular depth estimation process involves training a monocular relative depth estimation model using the underwater image dataset to obtain a relative depth map. Furthermore, a monocular absolute depth estimation model is trained to obtain a monocular depth map. Model training is used to obtain both the monocular relative depth estimation model and the monocular absolute depth estimation model. The binocular depth estimation process involves training a binocular depth estimation model using the underwater image dataset to obtain a depth map. This process of obtaining the binocular depth estimation model is performed offline. The real-time depth of a single camera is obtained using the monocular relative depth estimation model and the monocular absolute depth estimation model, respectively. The real-time depth of the binocular camera is then estimated using the binocular depth estimation model on the camera's real-time images.
[0033] Specifically, the monocular depth map and the binocular depth map are compared to determine the target depth map, including: performing quality analysis on the monocular depth map and the binocular depth map according to depth map evaluation indicators, and using the depth map with higher quality as the target depth map. Depth map evaluation indicators include but are not limited to geometric accuracy indicators, edge / boundary accuracy indicators, robustness to occlusion and weak texture areas, edge alignment, noise and artifacts, etc. That is, the depth map effects of the two depth estimation methods are compared to determine the output camera real-time depth map. This process of obtaining real-time depth is performed online in real time.
[0034] Step S13: removing moving objects from the processed image through dynamic object detection to obtain a static scene image.
[0035] While estimating depth, a dynamic object detection model is used to locate and remove moving targets from the processed image to obtain a static scene image. The dynamic object detection model is a neural network model pre-trained using underwater images. Dynamic object detection mainly identifies and deletes dynamic targets in the image, such as underwater organisms, to reduce the impact on positioning and reconstruction.
[0036] Dynamic object detection identifies and segments moving objects in video or image sequences while eliminating static background interference. It analyzes pixel or region changes over time, combining motion signatures with semantic information for accurate detection. It also integrates deep learning object detection with motion cues and incorporates camera motion compensation technology to mitigate the effects of device motion.
[0037] The training process of the dynamic object detection model includes: collecting underwater video sequences of different types of waters and different lighting conditions; labeling the objects in the images in the underwater video sequences; performing data enhancement processing on the labeled images to obtain an underwater image training set; and iteratively training the initial neural network using the underwater image training set until the conditions are met to obtain a dynamic object detection model.
[0038] Specifically, a training set of underwater images can be generated using cable-controlled underwater robots, autonomous underwater vehicles, or fixed cameras to capture video sequences of different water types and lighting conditions. These water types include, but are not limited to, turbid, clear, deep, and shallow; and lighting conditions include natural and artificial light. The images in the underwater image training set are labeled with object types, distinguishing between dynamic and static objects. For example, static labels are added to corals, while dynamic labels are added to moving objects such as fish, divers, and underwater robots. Furthermore, semantic segmentation and annotation of static backgrounds can be performed for subsequent data enhancement or background suppression. To increase data diversity, data augmentation can be performed, such as by adding scattered noise, color shifting, and simulating varying turbidity levels to simulate diverse underwater environments. Linear or nonlinear motion blur can also be applied to dynamic objects to simulate rapid swimming. Optical flow can also be used to generate interpolated data between frames to improve the model's understanding of motion continuity. Neural network models can use lightweight models as their backbone to balance the computational limitations of underwater equipment and incorporate multi-scale feature fusion modules to address small object detection.
[0039] Step S14: Based on the target depth map and the static scene image, the pose of the underwater robot is estimated in real time through a synchronous positioning and mapping algorithm, and a dense point cloud map is constructed to update the pose map of the underwater robot.
[0040] Finally, based on the obtained target depth map and static scene image, the underwater robot's pose is estimated in real time through a simultaneous localization and mapping algorithm, and a dense point cloud map is constructed to update the underwater robot's pose graph. Figure 3 As shown in the figure, the left and right eye images are acquired through the binocular camera, and the image size is adjusted, the image is rectified, and the image is enhanced. Then, depth estimation is performed to obtain a depth map, dynamic object detection is performed, and then visual SLAM (Simultaneous Localization and Mapping) is performed to generate an image pose graph and a dense point cloud map.
[0041] Specifically, based on the target depth map and the static scene image, the pose of the underwater robot is estimated in real time through a synchronous positioning and mapping algorithm, and a dense point cloud map is constructed to update the pose graph of the underwater robot, including: extracting feature points from the current static scene image and performing pose estimation according to the target depth map and the pose of the previous frame of the static scene image to obtain the pose of the current static scene image; generating a single-frame point cloud according to the target depth map, filtering the single-frame point cloud and fusing it into a dense point cloud map; judging whether the current static scene image is a key frame, and if so, performing local bundling adjustment optimization based on the current static scene image; detecting whether there is a loop according to the current static scene image, and if so, performing pose graph optimization based on the current static scene image.
[0042] The tracking thread is responsible for real-time tracking and localization. It extracts feature points from the current image and estimates the initial pose of the current image using the pose information from the previous image. A depth estimation model is used to estimate depth from keyframes, obtaining depth information and generating a single-frame point cloud. After filtering the point cloud, the point cloud data is published. The local mapping thread builds and optimizes local maps, selecting representative keyframes from the captured images. It then performs local bundle adjustment (BA) optimization to minimize reprojection error, significantly improving map accuracy. The loop closure and map fusion thread builds and optimizes the global map to ensure map integrity and consistency. Specifically, it identifies whether the robot has revisited previously visited areas, reducing redundant information and improving map accuracy. After detecting a loop, the loop closure and map fusion thread further identifies the current position and matches it to the known map. Finally, to further improve the accuracy of the global map, the loop closure and map fusion thread optimizes the global map, significantly improving its accuracy and more accurately reflecting the true structure of the environment.
[0043] In some embodiments, estimating the underwater robot's pose in real time using a simultaneous localization and mapping algorithm based on the target depth map and the static scene image includes: estimating the underwater robot's pose in real time using a simultaneous localization and mapping algorithm based on the target depth map and the static scene image, as well as sensor data collected by a target sensor; the target sensor includes an inertial measurement unit and / or sonar. That is, in addition to visual localization based on the target depth map and the static scene image, sensors such as an inertial measurement unit and sonar may also be used.
[0044] As can be seen from the above, in this embodiment, the original image captured by the camera is obtained, and the original image is pre-processed to obtain a processed image; depth estimation is used to determine the target depth map corresponding to the processed image; a dynamic object detection model is used to locate and remove moving targets from the processed image to obtain a static scene image; the dynamic object detection model is a neural network model pre-trained using underwater images; based on the target depth map and the static scene image, the underwater robot's position and posture are estimated in real time using a synchronous positioning and mapping algorithm, and a dense point cloud map is constructed to update the underwater robot's position and posture map. It can be seen that by performing visual positioning through a synchronous positioning and mapping algorithm, accurate underwater close-range positioning can be achieved at a low cost; and the collected original image is pre-processed before subsequent calculations are performed to improve image quality and lay the foundation for subsequent visual positioning; in addition, for special underwater environments, moving targets in the image are deleted before visual positioning to avoid interference from dynamic objects such as underwater fish that affect the positioning effect, thereby improving the accuracy of visual positioning.
[0045] Based on the above embodiments, the present application also discloses an underwater robot visual positioning system, which includes: a binocular digital camera for collecting original images; a computing unit communicatively connected to the binocular digital camera, for implementing the aforementioned underwater robot visual positioning method; and an unmanned remote-controlled submersible host communicatively connected to the computing unit.
[0046] The underwater robot's visual positioning system consists of camera hardware and software. The hardware system comprises a binocular camera and a computing unit. After acquiring image data from the camera, the computing unit performs image enhancement, correction, and dynamic object detection, laying the foundation for visual positioning and 3D reconstruction.
[0047] Among them, the above-mentioned computing unit includes: an AI acceleration processor, an image processing unit, and an underwater power supply; the AI acceleration processor is connected to the unmanned remote controlled underwater vehicle host through Ethernet, and the AI acceleration processor is also connected to the left camera and the right camera of the binocular digital camera respectively, for receiving images captured by the left camera and the right camera; the image processing unit is connected to the AI acceleration processor, the left camera and the right camera respectively, for receiving the shooting instructions issued by the AI acceleration processor, and synchronously triggering the left camera and the right camera through IO hard triggering according to the shooting instructions; the underwater power supply is connected to the AI acceleration processor for powering the AI acceleration processor.
[0048] For example Figure 4As shown, the system's binocular cameras use digital cameras. Digital cameras offer stable signals and high resolution, and can transmit digital image information to a computer via USB or Gigabit Ethernet, making the system simpler than analog cameras. The binocular cameras use a USB 3.0 data interface for high-resolution image transmission. This interface connects to the AI accelerator processor (embedded AI application platform) via USB, simplifying the system. This system uses a USB interface. To ensure simultaneous capture of both binocular cameras, cameras with hard triggering are required. Hard triggering utilizes the same image processing unit to trigger both cameras. External triggering triggers the cameras, achieving higher synchronization accuracy. The IO ports of cameras 1 and 2 are connected to the IO ports of the image processing unit. IO hard triggering synchronizes the two cameras, ensuring high synchronization accuracy. Cameras 1 and 2 are connected to the embedded AI application platform via USB 3.0. (The embedded AI application platform can feature options such as 16GB of memory, 1TSSD, and up to 100TOPS of AI computing power.) The embedded AI platform is connected to an Ethernet switch on the remotely underwater operated vehicle (ROV) via Ethernet, enabling communication between the binocular camera and the ROV host. The AI application platform is powered by a DC-DC power supply. The binocular camera system is connected to the ROV via a cable.
[0049] Based on the above embodiments, the present application also discloses an underwater robot, including a sealed pressure chamber and the aforementioned underwater robot visual positioning system; the underwater robot visual positioning system includes a binocular digital camera and a computing unit communicatively connected to the binocular digital camera; the underwater robot visual positioning system is inside the sealed pressure chamber.
[0050] The left camera of the binocular digital camera is located inside the first circular sealed pressure chamber, the right camera of the binocular digital camera is located inside the second circular sealed pressure chamber, and the computing unit is located inside the third circular sealed pressure chamber.
[0051] It is understandable that since the binocular camera works in an underwater environment, it needs to solve the pressure and sealing problems. This application designs a circular cavity to evenly withstand high water pressure, and the camera lens, control platform, etc. are built into the circular sealed cavity to adapt to the high-pressure working environment of 6000 meters deep water. Figure 5 The figure shows the sealed pressure-bearing cavity of the binocular camera, including the first circular sealed pressure-bearing cavity and the second circular sealed pressure-bearing cavity. Figure 6 Shown is a schematic diagram of the interior of the first circular sealed pressure-bearing chamber / the second circular sealed pressure-bearing chamber, as well as a side view of the circular sealed pressure-bearing chamber. Figure 7 The figure shows the computing unit inside the third circular sealed pressure-bearing chamber.
[0052] Correspondingly, the present application also discloses a visual positioning device for an underwater robot, see Figure 8 As shown, the device includes: The image preprocessing module 11 is used to obtain the original image captured by the camera and preprocess the original image to obtain a processed image; a depth estimation module 12, configured to determine a target depth map corresponding to the processed image using depth estimation; A dynamic object detection module 13 is configured to locate and remove moving objects from the processed image using a dynamic object detection model to obtain a static scene image; the dynamic object detection model is a neural network model pre-trained using underwater images; The visual positioning module 14 is used to estimate the position and posture of the underwater robot in real time based on the target depth map and the static scene image through a synchronous positioning and mapping algorithm, and to construct a dense point cloud map to update the position and posture graph of the underwater robot.
[0053] As can be seen from the above, in this embodiment, visual positioning is performed through synchronous positioning and mapping algorithms, achieving close-range precise underwater positioning at a low cost. In addition, the collected original image is preprocessed before subsequent calculations to improve image quality and lay the foundation for subsequent visual positioning. In addition, for the special underwater environment, moving targets in the image are deleted before visual positioning to avoid interference from dynamic objects such as underwater fish that affect the positioning effect, thereby improving the accuracy of visual positioning.
[0054] In some specific embodiments, the camera is a binocular camera, and the original image includes a binocular image; the depth estimation module 12 includes: a monocular depth map determining unit, configured to input any monocular image in the original image into a monocular depth estimation system to obtain a monocular depth map corresponding to the original image; A binocular depth map determination unit is configured to input the binocular image into a binocular depth estimation model to obtain a binocular depth map corresponding to the original image; the monocular depth estimation system and the binocular depth estimation model are both pre-trained using an underwater image dataset; A comparison unit is configured to compare the monocular depth map with the binocular depth map to determine the target depth map.
[0055] In some specific embodiments, the monocular depth estimation system includes a monocular relative depth estimation model and a monocular absolute depth estimation model in sequence, and the input of the monocular absolute depth estimation model is the output of the monocular relative depth estimation model; The monocular relative depth estimation model is used to output a corresponding relative depth map according to the input monocular image, and the monocular absolute depth estimation model is used to output a corresponding absolute depth map according to the input relative depth map.
[0056] In some specific embodiments, the comparison unit may be specifically configured to perform quality analysis on the monocular depth map and the binocular depth map according to a depth map evaluation index, and use a depth map with higher quality as the target depth map.
[0057] In some specific embodiments, the preprocessing includes image resizing, image enhancement, and image rectification.
[0058] In some specific embodiments, the image preprocessing module 11 may specifically include: The image correction unit is used to perform distortion correction on the original image using the radial distortion coefficient and the tangential distortion coefficient obtained by camera calibration.
[0059] In some specific embodiments, the visual positioning module 14 may specifically include: A pose estimation unit is used to extract feature points from the current static scene image and perform pose estimation based on the target depth map and the pose of the previous static scene image to obtain the pose of the current static scene image; a point cloud fusion unit, configured to generate a single-frame point cloud according to the target depth map, and filter the single-frame point cloud and fuse it into a dense point cloud map; a local bundling adjustment and optimization unit, configured to determine whether the current static scene image is a key frame, and if so, perform local bundling adjustment and optimization based on the current static scene image; A pose graph optimization unit is used to detect whether a loop occurs according to the current static scene image, and if so, perform pose graph optimization according to the current static scene image.
[0060] In some specific embodiments, the visual positioning module 14 may specifically include: A visual positioning unit is used to estimate the position and posture of the underwater robot in real time through a synchronous positioning and mapping algorithm based on the target depth map and the static scene image, as well as the sensor data collected by the target sensor; the target sensor includes an inertial measurement unit and / or a sonar.
[0061] Furthermore, an embodiment of the present application also discloses a computer storage medium, in which computer executable instructions are stored. When the computer executable instructions are loaded and executed by a processor, the steps of the underwater robot visual positioning method disclosed in any of the aforementioned embodiments are implemented.
[0062] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program, which, when executed by a processor, implements the steps of the underwater robot visual positioning method disclosed in any of the aforementioned embodiments.
[0063] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0064] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0065] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0066] The above is a detailed introduction to the underwater robot visual positioning method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A visual positioning method for an underwater robot, characterized in that: include: Obtaining an original image captured by a camera, and preprocessing the original image to obtain a processed image; Determining a target depth map corresponding to the processed image using depth estimation; Locating and removing moving objects from the processed image using a dynamic object detection model to obtain a static scene image; The dynamic object detection model is a neural network model pre-trained using underwater images; Based on the target depth map and the static scene image, the pose of the underwater robot is estimated in real time through a synchronous positioning and mapping algorithm, a dense point cloud map is constructed, and the pose map of the underwater robot is updated.
2. The underwater robot visual positioning method according to claim 1, characterized in that: The camera is a binocular camera, and the original image includes a binocular image; Determining a target depth map corresponding to the processed image using depth estimation includes: Input any monocular image in the original image into a monocular depth estimation system to obtain a monocular depth map corresponding to the original image; Inputting the binocular image into a binocular depth estimation model to obtain a binocular depth map corresponding to the original image; the monocular depth estimation system and the binocular depth estimation model are both pre-trained using an underwater image dataset; The monocular depth map and the binocular depth map are compared to determine the target depth map.
3. The underwater robot visual positioning method according to claim 2, characterized in that: The monocular depth estimation system includes a monocular relative depth estimation model and a monocular absolute depth estimation model in sequence, wherein the input of the monocular absolute depth estimation model is the output of the monocular relative depth estimation model; The monocular relative depth estimation model is used to output a corresponding relative depth map according to the input monocular image, and the monocular absolute depth estimation model is used to output a corresponding absolute depth map according to the input relative depth map.
4. The underwater robot visual positioning method according to claim 2, characterized in that: Comparing the monocular depth map with the binocular depth map to determine the target depth map includes: The monocular depth map and the binocular depth map are subjected to quality analysis according to a depth map evaluation index, and a depth map with high quality is used as the target depth map.
5. The underwater robot visual positioning method according to claim 1, characterized in that: The preprocessing includes image resizing, image enhancement and image correction.
6. The underwater robot visual positioning method according to claim 5, characterized in that: The image correction process for the original image includes: The radial distortion coefficient and the tangential distortion coefficient obtained by camera calibration are used to perform distortion correction on the original image.
7. The underwater robot visual positioning method according to claim 1, characterized in that: Based on the target depth map and the static scene image, the pose of the underwater robot is estimated in real time by a synchronous positioning and mapping algorithm, a dense point cloud map is constructed, and the pose graph of the underwater robot is updated, including: Extracting feature points from the current static scene image and performing pose estimation based on the target depth map and the pose of the previous static scene image to obtain the pose of the current static scene image; Generate a single-frame point cloud according to the target depth map, filter the single-frame point cloud and fuse it into a dense point cloud map; Determining whether the current static scene image is a key frame, and if so, performing local bundling adjustment optimization based on the current static scene image; Detect whether a loop occurs according to the current static scene image, and if so, perform pose graph optimization according to the current static scene image.
8. The underwater robot visual positioning method according to any one of claims 1 to 7, characterized in that: The method estimates the position and posture of the underwater robot in real time based on the target depth map and the static scene image by using a synchronous positioning and mapping algorithm, including: Based on the target depth map and the static scene image, as well as the sensor data collected by the target sensor, the position and posture of the underwater robot are estimated in real time through a synchronous positioning and mapping algorithm; the target sensor includes an inertial measurement unit and / or sonar.
9. A visual positioning device for an underwater robot, characterized in that: include: An image preprocessing module is used to obtain the original image captured by the camera and preprocess the original image to obtain a processed image; a depth estimation module, configured to determine a target depth map corresponding to the processed image using depth estimation; A dynamic object detection module, configured to locate and remove moving objects from the processed image using a dynamic object detection model to obtain a static scene image; The dynamic object detection model is a neural network model pre-trained using underwater images; The visual positioning module is used to estimate the position and posture of the underwater robot in real time based on the target depth map and the static scene image through a synchronous positioning and mapping algorithm, construct a dense point cloud map, and update the position and posture graph of the underwater robot.
10. A computer-readable storage medium, characterized in that Used to store computer programs; wherein when the computer programs are executed by the processor, the underwater robot visual positioning method according to any one of claims 1 to 8 is implemented.
11. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the underwater robot visual positioning method according to any one of claims 1 to 8.
12. An underwater robot visual positioning system, characterized in that: include: Binocular digital camera, used to collect original images; A computing unit in communication with the binocular digital camera, configured to implement the underwater robot visual positioning method according to any one of claims 1 to 8; An unmanned remotely operated underwater vehicle host computer is communicatively connected to the computing unit.
13. The underwater robot visual positioning system according to claim 12, characterized in that: The computing unit includes: an AI acceleration processor, an image processing unit, and an underwater power supply; The AI acceleration processor is connected to the host of the unmanned remotely operated underwater vehicle via Ethernet, and the AI acceleration processor is also connected to the left camera and the right camera of the binocular digital camera respectively, for receiving images captured by the left camera and the right camera; The image processing unit is connected to the AI acceleration processor, the left camera, and the right camera respectively, and is used to receive a shooting instruction issued by the AI acceleration processor, and synchronously trigger the left camera and the right camera through an IO hard trigger according to the shooting instruction; The underwater power supply is connected to the AI acceleration processor for supplying power to the AI acceleration processor.
14. An underwater robot, characterized in that: It comprises a sealed pressure-bearing chamber and the underwater robot visual positioning system according to claim 12 or 13; the underwater robot visual positioning system comprises a binocular digital camera and a computing unit in communication with the binocular digital camera; The underwater robot visual positioning system is inside the sealed pressure chamber.
15. The underwater robot according to claim 14, characterized in that: The left camera of the binocular digital camera is inside the first circular sealed pressure-bearing cavity, the right camera of the binocular digital camera is inside the second circular sealed pressure-bearing cavity, and the computing unit is inside the third circular sealed pressure-bearing cavity.
Citation Information
Patent Citations
An indoor positioning method and device based on SLAM
CN109671119A
Method and device for training image processing network and image processing
CN112862877A
Visual positioning and static map construction method and system in dynamic environment
CN112991447A
Dynamic scene multi-semantic map construction method and device based on visual SLAM
CN115937451A
Monocular vision SLAM algorithm based on deep learning
CN116481540A