An underwater visual positioning method based on deep learning visual inertial odometry
By introducing deep learning visual inertial odometers into the underwater visual positioning system, the stability problem of feature point extraction and matching in complex underwater environments is solved, and stable positioning is achieved under dynamic light and weak texture environments is achieved, which meets the real-time positioning needs of underwater robots.
Patent Information
- Application Number
- CN202510295185.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing underwater visual positioning technology is difficult to achieve stable feature point extraction and matching in complex underwater environments, especially in dynamic light and weak texture scenarios, resulting in inaccurate positioning and lack of real-time and universality.
Using a deep learning visual inertial odometer method, by constructing an underwater monocular visual inertial odometer system, combining the deep learning visual odometer model and a nonlinear optimization backend, the graphics processor and the central processor are used to operate feature extraction and matching separately to enhance the reliability and stability of feature points.
The number of reliable matching feature points is significantly increased in dynamic lighting and weak texture environments, the accuracy and stability of the position and posture estimation of the underwater camera is improved, the autonomous positioning needs of underwater robots are met, and real-time positioning is achieved on the embedded platform.
Smart Images

Figure CN119810491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underwater visual positioning method, which relates to the fields of computer vision, underwater visual positioning technology and marine technology, and particularly relates to an underwater visual positioning method based on deep learning visual inertial odometry. Background Art
[0002] The underwater world contains rich energy, minerals and biological resources, which have great economic and strategic value. The exploration and acquisition of underwater resources are inseparable from advanced marine equipment. As a new type of marine equipment, underwater robots can carry various sensor devices and operation tools to perform underwater tasks for a long time, over a large range and at low cost instead of operators under complex sea conditions, minimizing the risk of casualties. When an underwater robot performs an underwater operation task, its own positioning and attitude estimation are the basis, which gives rise to tasks such as autonomous navigation obstacle avoidance and path tracking. Higher-level marine environment detection, seabed resource exploration, maintenance of marine pipelines and cables, etc. all rely on accurate and reliable underwater positioning technology support. However, due to the lack of a global positioning system similar to GPS (Global Positioning System) underwater, related tasks are difficult to carry out. The real-time positioning technology of underwater mobile robots has become an important technology for venturing into the deep blue.
[0003] Although traditional underwater acoustic positioning technologies such as long baseline, short baseline and ultra-short baseline can obtain the position information of the robot relatively accurately, they also have defects such as the need to deploy base stations in advance, high costs for equipment installation and maintenance, difficulty in obtaining the motion attitude of the underwater robot, and lack of texture information of the detected target. Underwater visual positioning technology is an important supplement to underwater acoustic positioning. Different from the underwater acoustic positioning method, the advantages of underwater visual positioning are intuitive and detailed imaging (the camera can obtain more abundant information), faster optical detection speed (the transmission speed of light in water is on the order of 109 m / s), suitable for small-range precise detection, not actively emitting signals outward, strong concealment, low cost of the camera, etc. Visual SLAM (Simultaneous Localization and Mapping) in visual positioning technology (using visual methods for simultaneous localization and mapping) can measure data through sensors such as cameras installed on mobile devices to achieve self-positioning in an unknown environment. The research of visual odometry is the key in this visual positioning technology. The underwater visual SLAM algorithm is an important part of underwater visual positioning technology, and the research of the underwater visual odometry algorithm is also the key content of underwater visual positioning technology.
[0004] At present, the visual odometers using the Scale-Invariant Feature Transform (SIFT) feature extraction algorithm and the Speeded-Up Robust Features (SURF) feature extraction algorithm have been tested in the sea area by using the Zeno Autonomous Underwater Vehicle (AUV). In simple trajectory tracking experiments, the SIFT feature extraction method is more consistent with the real trajectory. The complex underwater lighting environment and the turbidity of seawater make the underwater grayscale images generally have problems of uneven exposure and low contrast, which makes it more difficult to extract image feature points compared with the atmospheric environment. Therefore, it is undoubtedly a solution idea to first perform underwater image enhancement and restoration and then send it into the feature extraction algorithm. Some scholars first use the histogram equalization algorithm to enhance the underwater images of the AQYACOL underwater dataset, and then send the enhanced underwater images into the visual odometer, and use the Oriented FAST and Rotated BRIEF (ORB) algorithm for feature extraction, increasing the number of feature points obtained, making the initialization and pose estimation of the visual odometer more stable; some scholars also take the means of underwater image enhancement. Aiming at the problems of low resolution and color distortion of underwater images, an image enhancement Retinex method based on the guided filter in the Hue Saturation Value (HSV) color space is proposed, and the AUV is used to collect images in the laboratory pool for verification. Similarly, the enhanced images are also used for feature extraction by the ORB algorithm, significantly increasing the number of feature points.
[0005] The above-mentioned underwater visual SLAM algorithms based on image enhancement strategies can, to a certain extent, increase the feature points that can be extracted from underwater images with low resolution and color distortion, providing a good basis for subsequent optimization and loop detection. However, these methods are only verified in a single water area environment, and the universality in different sea areas is not verified. In addition, none of the image enhancement means gives the running time on the hardware platform, and the real-time performance cannot be guaranteed. Summary of the Invention
[0006] To solve the problems existing in the background technology, the present invention provides an underwater visual positioning method based on deep learning visual inertial odometer. The present invention is a positioning technology that uses an underwater mobile device as a platform, fuses underwater image data and Inertial Measurement Unit (IMU) data to obtain the position and attitude information of the camera mounted on the underwater mobile device, so as to realize the positioning of the underwater mobile device.
[0007] The technical solution adopted by the present invention is as follows:
[0008] The underwater visual positioning method based on deep learning visual inertial odometer of the present invention includes:
[0009] S1: Construct an underwater monocular visual inertial odometer system, which includes a deep learning visual odometer model based on a descriptor loss function introducing a scale scaling factor and a non-linear optimization backend.
[0010] S2: Obtain a number of underwater videos through an underwater robot equipped with a monocular camera and extract consecutive underwater video frames. Form an image pair with every two consecutive underwater video frames, and construct each group of image pairs into a training set; use the training set to train the deep learning visual odometer model to obtain a trained underwater deep learning visual odometer model.
[0011] S3: Obtain the underwater video to be positioned through an underwater robot equipped with a monocular camera and extract consecutive underwater video frames, then input them into the underwater deep learning visual odometer model. After processing, output the information of successfully associated feature points and publish them through the communication interface of the Robot Operating System (ROS).
[0012] S4: The non-linear optimization backend subscribes to the node data of the ROS communication interface of the robot operating system, processes it, and outputs the pose of the monocular camera to achieve underwater visual positioning.
[0013] In the step S1 described above, the deep learning visual odometer model includes a feature extraction layer and a feature matching layer connected in sequence.
[0014] The deep learning model of the present invention is applied to the underwater environment with unstable lighting and weak texture, which can significantly increase the number of reliable matching feature points, making the position and pose estimation of the underwater camera more accurate and stable.
[0015] In the step S1 described above, the descriptor loss function introducing a scale scaling factor L d is as follows:
[0016] L d (D, D ' , S ) = (∑ h,w ∑ h',w' ( l d ( d hw , d' h'w' ; s hwh'w'))) / ( H c W c ) 2
[0017] l d ( d hw , d' h'w' ; s hwh'w' )= λ d × s hwh'w' × max (0, m p -( d hw T d' h'w' ) / ( d k ) 1 / 2 )+(1- s hwh'w' )× max (0,( d hw T d' h'w' ) / ( d k ) 1 / 2 - m n )
[0018] h , h' =1, 2, …, H c
[0019] w , w' =1, 2, …, W c
[0020] Among them, D and D ' respectively represent the original image of the decoder in the feature extraction layer of the input deep learning visual odometry model and its output image at the same position after being transformed by the homography relationship matrix S transformed, d hw and d' h'w' respectively represent the feature vectors in the input original image D and the transformed output image D ' transformed, s hwh'w'Represents the elements in the homography relationship matrix S ; l d ( ) represents the descriptor loss; T Represents the transpose; H c and W c respectively represent the height and width of the input original image D and the output image D ' ; λ d 、 m p and m n respectively represent the preset first, second, and third hyperparameters of the descriptor loss l d ; d k Represents the scale factor.
[0021] In the step S2 described above, for each image pair, the two underwater video frames in the image pair are input into the feature extraction layer for processing, and the feature extraction layer extracts each feature point and its descriptor in the two underwater video frames respectively.
[0022] In the step S2 described above, for each image pair, the descriptors of each feature point in the two underwater video frames output by the feature extraction layer are input into the feature matching layer for processing. The feature matching layer matches each descriptor one by one, thereby associating each feature point between the two underwater video frames in the image pair, and publishes the information of the successfully associated feature points through the robot operating system ROS communication interface.
[0023] In the step S2 described above, the deep learning visual odometry model is trained, that is, the feature extraction layer is trained. First, the underwater video frames are randomly affine-transformed and then input into the feature extraction layer, so as to perform transfer learning on the feature extraction layer to generate pseudo ground truth labels, obtain the trained underwater feature extraction layer, and finally obtain the trained underwater deep learning visual odometry model.
[0024] In the step S4 described above, an inertial measurement unit IMU is also installed on the underwater robot. After the nonlinear optimization backend subscribes to the node data of the robot operating system ROS communication interface, it is processed by the sliding window optimization algorithm to minimize the weighted sum of the reprojection error of the feature points and the pre-integration error of the inertial measurement unit IMU, and finally outputs the monocular camera pose of the camera frequency and the inertial measurement unit IMU frequency.
[0025] In step S4 described above, the deep learning visual odometry model of the underwater monocular visual inertial odometry system runs on a Graphics Processing Unit (GPU) device, and the non-linear optimization backend runs on a Central Processing Unit (CPU) device.
[0026] The method of the present invention introduces deep learning methods into the image feature extraction and matching part of underwater visual odometry, decouples this part from the non-linear optimization backend, and runs them on a Graphics Processing Unit (GPU) device and a Central Processing Unit (CPU) device respectively, making full use of the hardware computing resources.
[0027] The electronic device of the present invention includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method as described above.
[0028] The computer-readable storage medium of the present invention stores program data thereon, and when the program data is executed by a processor, the method as described above is implemented.
[0029] The method of the present invention aims at the task requirement of an underwater robot to achieve autonomous underwater visual positioning in a real working environment. Even in a harsh underwater visual environment such as a dynamic light and weak texture scene, stable and reliable visual positioning can still be achieved.
[0030] The beneficial effects of the present invention are:
[0031] Compared with other visual inertial odometers, by introducing deep learning methods in feature point extraction and matching, the present invention enables the visual inertial odometer to obtain sufficient matching feature point position information in an underwater weak texture and unstable light environment, provides a good initial value for the non-linear optimization backend, enables the non-linear optimization loss function to converge more quickly, and thus obtains more accurate camera pose information, achieving better visual positioning ability for an underwater mobile robot. Description of the Drawings
[0032] Figure 1 is a working mode diagram of the method of the present invention;
[0033] Figure 2 is a comparison diagram of an underwater unstable light environment image matching experiment, where Figure 2 (a) is a result diagram of feature point extraction and matching of an underwater unstable light environment image using the sparse optical flow method KLT (Kanade-Lucas-Tomasi), Figure 2 and (b) is a result diagram of feature point extraction and matching of an underwater unstable light environment image using the feature extraction ORB method. Figure 2Figure (c) shows the result of feature point extraction and matching for underwater unstable illumination environment images using the method of the present invention;
[0034] Figure 3 Figure is a comparison graph for the matching experiment of underwater weak texture environment images. Among them, Figure 3 Figure (a) shows the result of feature point extraction and matching for underwater weak texture environment images using the KLT sparse optical flow method; Figure 3 Figure (b) shows the result of feature point extraction and matching for underwater weak texture environment images using the ORB feature extraction method; Figure 3 Figure (c) shows the result of feature point extraction and matching for underwater weak texture environment images using the method of the present invention;
[0035] Figure 4 Figure is the model training flow chart of the transfer learning of the present invention;
[0036] Figure 5 Figure is a schematic diagram of an image extended by Homographic Adaptation;
[0037] Figure 6 Figure is a precision and recall curve graph of model training records. Among them, Figure 6 Figure (a) is the precision curve graph of model training records; Figure 6 Figure (b) is the recall curve graph of model training records. Detailed implementation manners
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] The underwater visual positioning method based on deep learning visual inertial odometer of the present invention is specifically as follows:
[0040] As shown in Figure 1 , first, an underwater monocular visual inertial odometer system is constructed. The underwater monocular visual inertial odometer system includes a deep learning visual odometer model based on a descriptor loss function introducing a scale scaling factor and a non-linear optimization backend; the deep learning visual odometer model includes a feature extraction layer and a feature matching layer connected in sequence. The feature extraction layer adopts a pre-trained MagicPoint model for feature point extraction, and the visual geometry VGG structure in the MagicPoint model for feature point extraction is replaced with the structure of the first three layers of the ResNet50 model; the feature matching layer adopts the LightGlue (Sparse Lightweight Graph Matching with Uncertainty Estimation) model. The non-linear optimization backend adopts a non-linear optimization backend based on Simultaneous Localization and Mapping (SLAM).
[0041] Then, an underwater robot equipped with a monocular camera is used to obtain several underwater videos and extract underwater video frames of consecutive frames. Each pair of consecutive underwater video frames is formed into an image pair, and each group of image pairs is constructed into a training set. The training set is used to train the deep learning visual odometry model, that is, to train the feature extraction layer. First, the underwater video frames are randomly affine-transformed and then input into the feature extraction layer, so as to perform transfer learning on the feature extraction layer to generate pseudo-ground truth labels, obtain the trained underwater feature extraction layer, and finally obtain the trained underwater deep learning visual odometry model. Specifically, during implementation, the parameters of the first three layers of the ResNet50 model in the SuperPoint model remain unchanged, and the parameters of other structures are optimized during training. During the training process, the feature extraction workflow in the deep learning visual odometry model is as follows:
[0042] First, the first grayscale image with a width of W and a height of H is input into the original SuperPoint model. It first goes through several encoder modules composed of convolutional layers, pooling layers, and activation layers, reducing the width and height of the image to 1 / 8 of the original, that is, the width in the first grayscale image becomes W c = W / 8, and the height becomes H c = H / 8. At the same time, the number of output feature layer channels is increased to obtain high-dimensional feature information of the image. The final feature layer is used as the input for the subsequent feature point position decoder and feature point descriptor decoder. However, the encoder of the original model uses the Visual Geometry Group (VGG) network structure, and during the process of continuously increasing the number of network layers, low-dimensional features will gradually be lost and it will be difficult to update the network weights. Therefore, the VGG structure of the original model is replaced with the structure of the first three layers of the ResNet50 model, which not only retains the 1 / 8 ratio of the output feature layer of the original backbone network to the length and width of the input image, but also introduces residual connections, greatly alleviating the defects of the original model's encoder. The feature point position decoder is used to generate the probability of whether each pixel point position on the image is a feature point. The input feature layer size of this module is W c × H c ×65, where X is obtained by adjusting the width, height, and channel dimensions of the feature layer after passing through convolutional layers, pooling layers, and activation functions in the decoder output feature layer, and is the input feature layer of the feature point position decoder. After Softmax and Pixel shuffle operations, it outputs H × WA feature layer of size ×1, where each feature vector corresponds one-to-one with each pixel on the image, representing the probability that the pixel point is a feature point. The pixel point positions that meet the threshold requirements are filtered out, and finally output m A position matrix of size ×2 p A 。The feature point descriptor decoder is used to generate a descriptor with a vector length of 256 for each pixel point on the image. The input feature layer D of this module has a size of W c × H c ×256. After bilinear interpolation, the width and height of the feature layer are enlarged to be the same as those of the original image, and then L2 normalization is performed channel by channel to ensure that the norm length of each descriptor is 1. After the operation, output H × W A feature layer of size ×256, where each feature vector corresponds one-to-one with each pixel on the image, representing the descriptor of the pixel point. The descriptors at the same positions as those of the feature point position encoder are filtered out, and finally output m A first descriptor matrix of size ×256 d A 。The second image also undergoes the above operations to obtain a second position matrix p B and a second descriptor matrix d B 。In addition, the training and validation data sets used in the original model are the COCO (Common Objects in Context) data set of common objects in the context. The pictures in this data set are taken in land scenes. Due to the absorption and scattering of light by water bodies, underwater optical images are prone to low contrast, fogging, and color distortion, which also greatly limits the extraction success rate of feature points of the original SuperPoint model and further affects the subsequent feature point matching effect. To solve the above problems, the present invention randomly extracts some images from four publicly available underwater image data sets and uses such as Figure 4The two-stage training strategy shown above fine-tunes the model parameters. First, it uses a MagicPoint feature point extraction model pre-trained on the context common object COCO dataset, and retrains it using the underwater images processed by image transformation to generate pseudo ground truth point labels for transfer learning of the SuperPoint model. These labels not only contain the original image information, but more importantly, provide reliable feature point position information. Then, the loss function is minimized according to the valid feature point labels to obtain the underwater MagicPoint model trained on underwater images. During the training of the underwater MagicPoint model, in order to improve the diversity and representativeness of the training samples, the Homographic Adaptation technology is adopted. This technology applies a series of random affine transformations to the input images. As Figure 5 shown, if the position error between the feature points detected in the original image and the feature points detected in the transformed image is within the threshold range after the same transformation, they are regarded as valid feature points. This data augmentation strategy significantly expands the number of training samples and improves the adaptability of the model to different perspectives at the same time. The model is trained on a training set containing 3685 images and a validation set containing 932 images constructed from four public underwater datasets, and the Adaptive Moment Estimation (Adam) optimizer is used. The initial learning rate is set to 0.0001, and the training is carried out for 600000 iterations, and a validation is performed on the validation dataset every 5000 iterations.
[0043] In order to make the training results of the model converge more quickly, the present invention does not use the loss function of the descriptor in the SuperPoint method, but introduces a scale factor d k (the number of channels of the input feature layer D), and the descriptor loss function L d is as follows:
[0044] L d (D, D ' , S ) = (∑ h,w ∑ h',w' ( l d ( d hw , d' h'w' ; s hwh'w' ))) / ( H cW c ) 2
[0045] h , h' = 1, 2, …, H c
[0046] w , w' = 1, 2, …, W c
[0047] wherein, D and D ' respectively represent the original image of the feature point descriptor decoder in the feature extraction layer of the input model and its output image at the same position after being transformed by the homography relationship matrix S transformed, d hw and d' h'w' respectively represent the feature vectors in the input original image D and the output image D ' ; s hwh'w' represents the element in the homography relationship matrix S If the Euclidean distance between the feature point coordinates extracted from the input original image D and the corresponding feature point coordinates after inverse transformation of the feature points extracted from the output image D ' after homography adaptation transformation is less than 8, then s hwh'w' = 1, otherwise s hwh'w' = 0; l d () represents the descriptor loss; T represents transpose; H c and W c respectively represent the height and width of the input original image D and the output image D ' ;
[0048] The present invention introduces the scale scaling factor d k into the descriptor loss l d as follows:
[0049] l d ( d hw , d' h'w' ; s hwh'w' ) = λd × s hwh'w' × max (0, m p -( d hw T d' h'w' ) / ( d k ) 1 / 2 ) + (1 - s hwh'w' ) × max (0, ( d hw T d' h'w' ) / ( d k ) 1 / 2 - m n )
[0050] Wherein, λ d 、 m p and m n respectively represent the preset first, second, and third hyperparameters of the descriptor loss l d .
[0051] Introducing the scale factor d k can prevent d hw T and d' h'w' from having too large or too small inner product of these two feature vectors. When it is too large, it will cause the curve of the loss function value to oscillate during the training process. When it is too small, the model parameters cannot be effectively updated. Both of these situations will cause the model parameters to fail to converge quickly or even lead to training failure.
[0052] The precision and recall curves of the model training records are as shown in (a) of Figure 6 and (b) of Figure 6 . The precision and recall of the model show an upward trend during the model training process, indicating that the extraction results of the feature points by the training network are satisfactory.
[0053] To further demonstrate the superiority of the optimized underwater feature extraction layer in extracting feature points in extremely underwater lighting environments, in the specific implementation of the present invention, extreme scenario images in two sequences from two publicly available datasets were respectively selected for feature point extraction experiments. The extreme scenario images were underwater weak texture scenario images and underwater dynamic lighting scenario images. The original SuperPoint model and the optimized model were used to conduct feature point extraction experiments on these images. The total number of feature points extracted was recorded for each series, and the percentage obtained by dividing the number of images in which the optimized model extracted more feature points than the original SuperPoint on a single image by the total number of images was denoted as BR (Better Rate). As shown in Table 1, the results indicate that whether it is the total number of feature points in each sequence or the BR value of the performance improvement rate, the model after transfer learning fine-tuning is superior to the original model.
[0054] Table 1 Feature Extraction Performance under Different Underwater Conditions
[0055]
[0056] Among them, AQ represents the AQUALOC dataset, FLS represents the FLSea_VI dataset, WT represents weak texture, DL represents dynamic lighting, and FC represents the number of feature points.
[0057] Then the feature matching workflow is as follows:
[0058] Take the first position matrix p A , the first descriptor matrix d A , the second position matrix p B and the second descriptor matrix d B as the input of the attention layer for feature matching. The first position matrix p A is first subjected to relative position rotation position encoding and then used together with the first descriptor matrix d A as the input of the self-attention module (Self-attention block) to calculate the self-attention score a ij , specifically as follows:
[0059] a ij = q i T R ( p j - pi ) k j
[0060] Among them, q i and k j are the key vector and query vector respectively obtained by decomposing the first descriptor matrix d A after different linear transformations. R () is the relative position rotation position encoding between feature points in the d-dimensional space. R () ∈ R d×d , p i 、 p j ∈ p A . In this way, the influence weight coefficients of each feature point on each other in the first image are obtained. Then, the output of the self-attention module and the first descriptor matrix d A are jointly used as the input of the multi-layer perceptron MLP (Multilayer Perceptron) to adjust the influence weight coefficients of each feature point on each other. The multi-layer perceptron includes a hidden layer, a layer normalization layer, an activation function GELU layer, and a convolutional layer. The second position matrix p B and the second descriptor matrix d B perform the above operations in the same way. After the above operations, the outputs of the first multi-layer perceptrons of the two are jointly used as the input of the cross-attention block to calculate the cross-attention score b ij , specifically as follows:
[0061] b ij = K i AT K j B
[0062] Among them, K i A is the key vector from the feature point i in the first image, K j B is the feature point from the second image jThe key vectors are used to obtain the influence weight coefficients between different image feature points; the output of the cross-attention module and the output of the first multi-layer perceptron are jointly used as the input of the second multi-layer perceptron to adjust the descriptors of their respective image feature points. Assume the output of the second multi-layer perceptron is x i , calculate the confidence of the adjusted descriptor c i , specifically as follows:
[0063] c i = Sigmoid ( x i ) ∈ [0, 1]
[0064] Sigmoid ( x ) = 1 / (1 + e -x )
[0065] When the confidence meets the threshold condition, exit the inference of the upcoming next attention layer and enter the similarity matrix module for similarity calculation. If the condition is not met, the pruning strategy will be executed to eliminate the feature point information with low confidence to reduce the computational load of the next attention layer. The input of the similarity matrix module is the descriptors of the first image adjusted by the influence weight coefficients x 1 A 、 x 2 A 、…、 x i A 、…、 x m A and the descriptors of the second image adjusted by the influence weight coefficients x 1 B 、 x 2 B 、…、 x j B 、…、 x n B , each descriptor x i A ( i ∈ [1, m ) are all calculated with the descriptor x j B ( j ∈ [1, n ) to form a similarity matrix, and select the top SAnd find the feature points of the two images they correspond to, and then you can get the successfully matched feature points.
[0066] Table 2 Feature extraction and matching performance under different underwater conditions
[0067]
[0068] Among them, the total number is the number of matching feature points after initial matching; the correct number is the number of internal points after RANSAC filtering; the accuracy is the percentage of correct matches (correct number / total number × 100%).
[0069] As shown in Table 2, the quantitative indicators of four feature point extraction and matching methods in two images with similar time in the presence of a common view area are shown. The four methods are the combination of the GFTT (Good Features To Track) feature point extraction method represented by the monocular visual-inertial system VINS-Mono (Visual-Inertial System Monocular) and the visual-inertial multi-sensor fusion system VINS-Fusion (Visual-Inertial System Fusion) and the sparse optical flow method KLT matching method, the ORB feature point extraction method represented by the ORB-SLAM series and the Fast Library for Approximate Nearest Neighbor Fast Search Library FLANN (Fast Library for Approximate Nearest Neighbor Fast Search Library FLANN) and the combination of the GFTT (Good Features To Track) feature point extraction method represented by the ORB-SLAM series and the sparse optical flow method KLT matching method, ...GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching method, and the GFTT (Good Features To Track) feature point matching Neighbors) feature point matching method, and a combination of the SuperPoint feature point extraction method and the LightGlue matching method. The feature point extraction and matching method combination used in the present invention, the evaluation indicators are respectively the total number of matching feature points, the total number of correctly matched feature points obtained after matching feature points screening using the random sampling consistency method, and the matching accuracy obtained by dividing the total number of correctly matched feature points by the total number of matching feature points. The present invention is tested on two public underwater data sets under representative natural environment conditions. The deep-sea underwater data set is the AQUALOC data set. The first sequence of the underwater archaeological remains data set in the data set is selected. The feature of this sequence is that there are many weak-texture scenes, but the illumination is provided by the underwater submersible and is therefore relatively stable. The shallow-sea underwater data set is the FLSea_VI data set. The tiny_canyon sequence in the data set is selected. The feature of this sequence is that the texture is rich but there is an underwater optical environment with drastic changes in light intensity. It can be seen from the table that in the deep-sea environment with stable illumination and weak texture represented by the AQUALOC data set, such as Figure 3The feature extraction and matching results shown in the black-and-white image. The visual front-end composed of the high-quality traceable feature point GFTT feature point extraction method and the sparse optical flow method KLT feature point matching method has advantages in the deep-sea underwater optical environment with stable lighting and weak texture. However, in the shallow-sea underwater optical environment with drastic changes in light intensity, a large number of mis-matched feature points and a low matching accuracy rate occur; the ORB feature point extraction method and the approximate nearest neighbor fast search library FLANN feature point matching method perform poorly both in the deep-sea underwater optical environment with stable lighting but weak texture and in the shallow-sea underwater optical environment with rich texture but drastic changes in light intensity, proving that the visual odometer under this combination is difficult to adapt to the underwater optical environment; the SuperPoint feature point extraction method and the LightGlue matching method also perform poorly in the above two environments, proving that simply splicing deep learning models and using the pre-trained weights of the models cannot demonstrate excellent performance in this specific underwater environment; the method in the present invention not only uses deep learning models for both feature point extraction and matching, but also deeply optimizes the models for the underwater environment. The matching accuracy rates of the two extreme visual scenes in the table demonstrate the strong robustness and accuracy of the front-end of the underwater visual inertial odometer proposed in the present invention in the underwater optical environment, proving that it exhibits excellent performance in both weak texture and dynamic lighting environments.
[0070] After training, an underwater deep learning visual odometer model is obtained. Applying the underwater deep learning visual odometer model to the underwater environment with unstable lighting and weak texture can significantly increase the number of reliable matching feature points, making the position and attitude estimation of the underwater camera more accurate and stable.
[0071] Then, an underwater robot equipped with a monocular camera acquires the underwater video to be localized and extracts consecutive frames of underwater video frames, which are input into the underwater deep learning visual odometry model. For each image pair, the two underwater video frames in the image pair are input into the feature extraction layer for processing. The feature extraction layer extracts each feature point and its descriptor in the two underwater video frames respectively. The descriptors of each feature point in the two underwater video frames output by the feature extraction layer are input into the feature matching layer for processing. The feature matching layer matches each descriptor one by one, thereby associating each feature point between the two underwater video frames in the image pair, and publishes the information of the successfully associated feature points through the Robot Operating System (ROS) communication interface. Then, the non-linear optimization backend subscribes to the node data of the ROS communication interface. An Inertial Measurement Unit (IMU) is also installed on the underwater robot. After the non-linear optimization backend subscribes to the node data of the ROS communication interface, it is processed by the sliding window optimization algorithm to minimize the weighted sum of the feature point reprojection error and the IMU pre-integration error, and finally outputs the pose of the monocular camera at the camera frequency and the IMU frequency. The specific workflow of the non-linear optimization backend is as follows:
[0072] The data of the Inertial Measurement Unit (IMU) consists of three-axis angular velocity and three-axis acceleration. The data of the IMU is pre-integrated and an initialization operation is completed with the feature points matched on the image. The initialization only needs to be performed once, that is, the matched feature points of two images are triangulated to obtain the depth information of each feature point from the two-dimensional information on the image, and the depth information is calibrated by integrating the IMU data. At the same time, through the epipolar geometry constraint algorithm, the rotation matrix and translation vector of the cameras of the two images can be obtained from the position information of the matched feature points, and the deviation in the angular velocity measured by the IMU is calibrated by this rotation matrix. At the same time, the initial world coordinate system is determined during this process. After the initialization is completed, the pre-integration information of the IMU and the position information of the matched feature points are maintained through the sliding window algorithm. The camera pose and the positions of the three-dimensional points in space corresponding to the feature points are optimized simultaneously through bundle adjustment, and the pose forward propagation based on the IMU is performed. Finally, the pose estimates of the camera at two frequencies are output, namely the pose estimate of the camera with the same frequency as the IMU and the pose estimate of the camera with the same frequency as the camera frame rate.
[0073] As Figure 2 shown in (a) of Figure 2 shown in (b) of Figure 2As shown in (c), in an underwater scene with unstable illumination, the feature extraction and matching method of the underwater visual-inertial odometer based on the deep learning feature extraction and matching method can match a larger number of feature points accurately compared with the traditional sparse optical flow method KLT and ORB feature point extraction method; as Figure 3 in (a), Figure 3 in (b) and Figure 3 as shown in (c), in an underwater scene with stable illumination but weak texture, the feature extraction and matching method of the underwater visual-inertial odometer based on the deep learning feature extraction and matching method can also obtain sufficient position information of matching feature points.
[0074] The present invention has carried out a quantitative comparison with the visual-inertial odometers in visual-inertial SLAM such as the monocular vision-inertial system VINS-Mono and the vision-inertial multi-sensor fusion system VINS-Fusion in sequences 3, 4, 5, 6, 8, 9 of the underwater dataset AQUALOC. The root mean square error of the absolute trajectory error between the three-dimensional camera motion trajectory generated by the visual-inertial odometer and the accurate three-dimensional camera motion trajectory provided by the official of this dataset is used as the comparison index. As shown in Table 3, it can be seen that the method of the present invention performs better, and the smaller the number in the table, the better.
[0075] Table 3 Quantitative comparison of visual-inertial odometers
[0076]
[0077] Among them, F represents system crash.
[0078] As shown in Table 3, the present invention exhibits consistent and reliable performance in all test sequences and is the only method that successfully maintains stable operation throughout the evaluation process. The system achieved very good results in Sequence 3, Sequence 6, and Sequence 8, with positioning accuracies of 0.5 m, 0.36 m, and 0.60 m, respectively. Most notably, in challenging scenarios such as Sequence 6, the monocular vision-inertial system VINS-Mono, the vision-inertial multi-sensor fusion system VINS-Fusion, and the lightweight sparse mapping visual inertial odometer LSM-VIO (Lightweight Sparse-Mapping Visual Interial Odometry) combined with the unoptimized original SuperPoint and LightGlue were unable to maintain operation (denoted as system crash F), while the underwater monocular vision inertial odometer system LSM-VIO of the present invention continued to operate robustly and provided an excellent accuracy of 0.36 m. More importantly, the underwater monocular vision inertial odometer system LSM-VIO achieved the best overall performance, with an average positioning error of 0.63 m, which significantly improved by 50.8% and 58.8% respectively compared to the monocular vision-inertial system VINS-Mono (1.28 m) and the vision-inertial multi-sensor fusion system VINS Fusion (1.53 m). This comprehensive performance advantage, combined with its consistent operational stability under various underwater conditions, further demonstrates the excellent robustness and practical applicability of the underwater monocular vision inertial odometer system LSM-VIO in practical underwater applications.
[0079] To ensure the stable operation of the underwater monocular visual inertial odometer system LSM-VIO on a resource-constrained underwater vehicle platform and achieve a pose output speed of more than 10 frames per second, key optimization strategies at the system implementation level were carried out. The optimization method solves two problems of computing resource allocation and model acceleration to jointly ensure real-time performance. In terms of computing resource allocation, the present invention adopts a decoupling strategy between the deep learning front end and the optimization back end. In order to improve the operation efficiency of the system and make better use of hardware resources, the system is split into a visual front end and a non-linear optimization back end and run on two threads respectively. Considering that deep learning inference is suitable for parallel processing, while optimization calculations require sequential operations, and visual odometry involves image feature extraction and matching, which involves a large number of repetitive operations. Therefore, while the system is running the back-end optimization module, the visual front-end calculation is mainly executed on the graphics processing unit (GPU) device, and the feature point information after matching is published through the Robot Operating System (ROS) topic communication method. The non-linear optimization back end involves a large number of complex operations and data access operations, so it runs on the central processing unit (CPU) device to perform the pose optimization process of the monocular camera and the IMU. Similarly, it subscribes to the feature point information through the ROS topic communication method. After decoupling, this method can make the most of computing resources and accelerate the operation efficiency. The two modules communicate through the ROS topic mechanism: the front end continuously publishes the feature point matching information, and the back end subscribes in real time and performs state estimation. This heterogeneous computing design avoids resource contention between modules while ensuring processing performance. In addition, thanks to the modular design, each functional module maintains an independent parameter adjustment interface, which is convenient for targeted optimization for different application scenarios. At the level of deep learning model acceleration, we carried out the following multi-layer optimization strategies. First, a basic acceleration environment was established by converting the PyTorch model into an engine format and using TensorRT as the inference framework. Second, aiming at the characteristics of the underwater visual positioning task, the operator fusion optimization target attention mechanism was realized. Specifically, in the LightGlue model, the attention mechanism module contains multiple computationally intensive operations such as matrix multiplication and softmax calculation, and there is a large amount of intermediate result transmission between these operations. The traditional method needs to repeatedly read and write these intermediate data in the GPU device memory, which will cause substantial memory access overhead. By analyzing the data flow dependence of the calculations in the attention module, the key and values matrices were cached to accelerate the model memory access. In addition, the grouped-query attention (GQA) method was used to infer the attention mechanism module, achieving a balance between the speed of multi-query attention (MQA) and the accuracy of multi-head attention (MHA).These optimization strategies reduce the storage and access frequency of intermediate results, not only reducing the memory bandwidth pressure, but also increasing the computational density, thus significantly improving the model inference efficiency.
[0080] The method of the present invention introduces deep learning methods into the image feature extraction and matching part of underwater visual odometry, decouples this part from the non-linear optimization backend, and runs them on a Graphics Processing Unit (GPU) device and a Central Processing Unit (CPU) device respectively, making full use of the hardware computing resources.
[0081] The real-time performance of the underwater visual inertial odometer proposed by the present invention has been evaluated on the NVIDIA Jetson Orin NX embedded development platform. NVIDIA Jetson Orin NX is a low-power embedded platform equipped with a 6-core NVIDIA ARM Cortex A78AE v8.2 64-bit Central Processing Unit (CPU) and a 1024-core NVIDIA Ampere Graphics Processing Unit (GPU) with 32 Tensor cores. The input consists of color images with a resolution of 968×608 pixels. For fair comparison, 256 feature points were extracted from each image before and after model optimization. The evaluation used 1012 consecutive images captured at intervals of 0.05 seconds to ensure sufficient co-visual areas for feature point extraction and matching in the images.
[0082] Table 4 Comparison of Feature Extraction and Matching Time Optimization
[0083]
[0084] As shown in Table 4, the computational efficiency of the method of the present invention is significantly improved after optimization, making it very suitable for embedded platforms. The original implementation required 0.175 seconds for each feature processing cycle. After applying the optimization method proposed by the present invention, the processing time was significantly shortened to 0.048 s, and the speed increased by 3.6 times. Detailed analysis shows that substantial improvements have been made in all processing stages. The feature extraction and matching times were reduced by 74.6% and 69.4% respectively, resulting in a 72.6% increase in overall efficiency. This optimization enables the underwater monocular visual inertial odometer system LSM-VIO to operate at a speed of approximately 15 Hz, meeting the real-time requirements of underwater visual positioning tasks.
[0085] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. The present application is described according to the flowcharts of the methods, systems and computer program products of the embodiments of the present application.
[0086] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the present invention is intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0087] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the equivalent technology of the present invention, the present application is also intended to include these changes and modifications.
Claims
1. An underwater vision positioning method based on deep learning visual inertial odometry, characterized in that Including: S1: Construct an underwater monocular visual inertial odometry system, which includes a deep learning visual odometry model based on a descriptor loss function introducing a scale scaling factor and a non-linear optimization backend; S2: Use an underwater robot equipped with a monocular camera to obtain several underwater videos and extract consecutive frames of underwater video frames. Combine every two consecutive underwater video frames to form an image pair, and construct each group of image pairs into a training set. Use the training set to train the deep learning visual odometry model to obtain a trained underwater deep learning visual odometry model; S3: Use an underwater robot equipped with a monocular camera to obtain an underwater video to be located and extract consecutive frames of underwater video frames, then input them into the underwater deep learning visual odometry model. After processing, output the successfully associated feature point information and publish it through the Robot Operating System (ROS) communication interface; S4: The non-linear optimization backend subscribes to the node data of the ROS communication interface and processes it to output the pose of the monocular camera, realizing underwater visual positioning; In the step S1, the deep learning visual odometry model includes a feature extraction layer and a feature matching layer connected in sequence; In the described step S1, a descriptor loss function introducing a scale scaling factor L d is as follows: L d (D,D ' , S )=(∑ h,w ∑ h',w' ( l d ( d hw , d' h'w' ; s hwh'w' ))) / ( H c W c ) 2 l d ( d hw , d' h'w' ; s hwh'w' )= λ d × s hwh'w' × max (0, m p -( d hw T d' h'w' ) / ( d k ) 1 / 2 )+(1- s hwh'w' )× max (0,( d hw T d' h'w' ) / ( d k ) 1 / 2 - m n ) h , h' =1、2、…、 H c w , w' =1、2、…、 W c Among them, D and D ' respectively represent the original image of the decoder in the feature extraction layer of the input deep learning visual odometry model and its output image after being transformed by the homography relationship matrix S ; d hw and d' h'w' respectively represent the feature vectors in the input original image D and the transformed output image D ' ; s hwh'w' represents the element in the homography relationship matrix S ; l d ( ) represents the descriptor loss; T represents transpose; H c and W c respectively represent the input original image D and the output image D ' 's height and width; λ d 、 m p and m n respectively represent the descriptor loss l d 's preset first, second, and third hyperparameters; d k represents the scale factor; In the step S2, when training the deep learning visual odometry model, that is, training the feature extraction layer. First, perform random affine transformation on the underwater video frames and then input them into the feature extraction layer, so as to perform transfer learning on the feature extraction layer to generate pseudo ground truth labels, obtain a trained underwater feature extraction layer, and finally obtain a trained underwater deep learning visual odometry model; The feature extraction module uses an improved feature point extraction SuperPoint model, that is, an improved feature point extraction SuperPoint_uw model. In the improved feature point extraction model SuperPoint_uw, the residual layer structure in the ResNet model is used to replace the Visual Geometry Group (VGG) structure in the feature point extraction SuperPoint model. The feature matching module uses the LightGlue model to pair the feature points output by the improved feature point extraction SuperPoint_uw model.
2. The underwater vision positioning method based on deep learning visual inertial odometer according to claim 1, wherein: In the step S2, for each image pair, input the two frames of underwater video frames in the image pair into the feature extraction layer for processing, and the feature extraction layer extracts each feature point and its descriptor in the two frames of underwater video frames respectively.
3. The underwater vision positioning method based on deep learning visual inertial odometer according to claim 2, wherein: In the step S2, for each image pair, input the descriptors of each feature point in the two frames of underwater video frames output by the feature extraction layer into the feature matching layer for processing. The feature matching layer matches each descriptor one by one, so as to associate each feature point between the two frames of underwater video frames in the image pair, and publish the successfully associated feature point information through the ROS communication interface.
4. The underwater vision positioning method based on deep learning visual inertial odometer according to claim 1, characterized in that: In the described step S2, the deep learning visual odometry model is trained, that is, the feature extraction layer is trained. First, the underwater video frames are subjected to random affine transformation processing and then input into the feature extraction layer, so as to perform transfer learning on the feature extraction layer to generate pseudo ground truth labels, obtain the trained underwater feature extraction layer, and finally obtain the trained underwater deep learning visual odometry model.
5. The underwater vision positioning method based on deep learning visual inertial odometer according to claim 1, characterized in that: In the described step S4, an inertial measurement unit IMU is also installed on the underwater robot. After the non-linear optimization backend subscribes to the node data of the robot operating system ROS communication interface, it is processed through the sliding window optimization algorithm to minimize the weighted sum of the feature point reprojection error and the inertial measurement unit IMU pre-integration error, and finally outputs the monocular camera pose at the camera frequency and the inertial measurement unit IMU frequency.
6. The underwater vision positioning method based on deep learning visual inertial odometer according to claim 1, characterized in that: In the described step S4, the deep learning visual odometry model of the underwater monocular visual-inertial odometry system runs on a graphics processing unit GPU device, and the non-linear optimization backend runs on a central processing unit CPU device.
7. An electronic device, characterized in that, Including: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1-6.
8. A computer-readable storage medium having program data stored thereon, characterized in that, When the program data is executed by the processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Monocular vision odometer method, device and system and storage medium
CN116182894A
Robot controller application-oriented YOLO-V5 model lightweight method and system
CN118505952A