Efficient six-degree-of-freedom pose estimation method for unmanned underwater vehicle
By constructing an underwater pose estimation model based on EfficientNet-B4 and GFPN, and using the six-degree of freedom interleaving and ratio loss function for training, the problems of low pose estimation accuracy and insufficient generalization performance in the underwater environment are solved, and efficient and robust pose estimation is achieved.
Patent Information
- Application Number
- CN202510319778.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
Smart Images

Figure CN120219494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned underwater vehicle (UUV) pose estimation, and particularly to an efficient six-degree-of-freedom pose estimation method for UUVs. Background Art
[0002] Accurate 6-DoF pose estimation of an unmanned underwater vehicle (UUV) during tracking and mapping tasks can significantly improve task execution efficiency and safety. Since lidar and sonar cannot obtain accurate positioning information in the underwater environment, this task is mostly performed using underwater cameras. However, due to the limitations of complex task environments, there is a lack of open-source datasets for UUV pose estimation, and at the same time, existing related algorithms cannot accurately and efficiently perform the task. In addition, current related work in the field all discusses pose estimation for a single type of UUV, and the methods and application scenarios are not generalizable. Therefore, researching an efficient, real-time, and lightweight UUV six-degree-of-freedom pose estimation method to solve problems such as low detection accuracy, insufficient generalization performance, and lack of effective training data in existing methods has important research value and significance for ensuring the efficient operation of unmanned underwater vehicles.
[0003] Deep convolutional neural networks can learn to obtain the ability to represent high-level features of images, and can efficiently utilize the semantic and spatial position information hidden in images, thereby realizing the estimation of the pose of rigid bodies in images. Six-degree-of-freedom pose estimation based on deep learning methods has become the current mainstream method. Existing deep learning rigid body pose estimation methods are mainly divided into correspondence-based methods, template-based methods, and voting-based methods. Correspondence-based methods include Single-Shot-Pose, YOLOv5-6D, etc. Although such methods are relatively real-time, their positioning accuracy is poor. Voting-based methods include PVNet, PVN3D, etc. Such methods have high positioning accuracy, but since they usually need to regress pixel-level vectors pointing to key points, the computational complexity is intensive, so their real-time performance is low and the computational amount is large. Template-based methods are not conducive to applications in underwater scenarios because they require template information and usually need to collect complete and clear images of the target object. In addition, due to technical limitations, there are limited literatures on UUV attitude estimation in underwater environments, and the number of open-source datasets is scarce, resulting in problems such as low accuracy and low domain adaptability in existing methods. Therefore, there is an urgent need for an efficient underwater UUV pose estimation method to solve the current problems and thus face the actual application of unmanned underwater vehicles. Summary of the Invention
[0004] The present invention provides an efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle to overcome the technical problems of low positioning accuracy and insufficient domain adaptation ability of existing deep learning object detection methods due to difficult factors such as small domain data volume and low algorithm positioning accuracy.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] An efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle, the specific steps include:
[0007] S1: Construct an unmanned underwater vehicle dataset, the unmanned underwater vehicle dataset includes a synthetic pose dataset and an unmanned underwater vehicle image dataset;
[0008] S2: Construct an underwater pose estimation model, and train the underwater pose estimation model based on the unmanned underwater vehicle dataset;
[0009] S3: Design the loss function of the underwater pose estimation model as a six-degree-of-freedom intersection over union loss function. When the loss function converges, obtain the trained underwater pose estimation model;
[0010] S4: Use the trained underwater pose estimation model to predict the prediction results of the key points of the actual unmanned underwater vehicle;
[0011] S5: Calculate the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the prediction results of the key points of the actual unmanned underwater vehicle.
[0012] Further, the underwater pose estimation model includes: a backbone network module, a neck network module, and a head network module;
[0013] The backbone network module is established based on the EfficientNet-B4 network architecture, and is used to extract the multi-scale pose features of the unmanned underwater vehicle in the unmanned underwater vehicle dataset and transmit them to the neck network module;
[0014] The neck network module is built based on the GFPN network, and is used to fuse the multi-scale pose features of the unmanned underwater vehicle to obtain the fused multi-scale features and transmit them to the head network module;
[0015] The head network module is used to perform key point regression processing on the fused multi-scale features to obtain the prediction results of the key points of the unmanned underwater vehicle.
[0016] Further, the expression of the intersection over union loss function under six degrees of freedom is:
[0017] Loss = λ pt L pt + λ conf Lconf +λ Bbox L Bbox (1)
[0018]
[0019] Among them, L pt 、L conf and L Bbox respectively represent coordinates, confidence, and bounding box loss; λ pt 、λ conf and λ Bbox respectively represent the weight coefficients of coordinates, confidence, and bounding box loss; smooth L1 is the smooth L1 loss function; IOU proj is the intersection over union result calculated from the projection of the bounding box on the plane; eps is the regularization term; I and J are geometric parameters; i and j are defined point pairs.
[0020] Furthermore, the steps of constructing the backbone network module include:
[0021] Introduce deformable convolutional network (DCN) into the C2f modules located in the fourth, sixth, and eighth layers of the EfficientNet-B4 network, so as to replace the C2f modules with DCN-C2f modules. After replacement, the backbone network module includes a first downsampling module, a first DCN-C2f module, a second downsampling module, a second DCN-C2f module, a third downsampling module, and a third DCN-C2f module;
[0022] Among them, the first downsampling module is used to downsample the input unmanned underwater vehicle dataset to obtain a first downsampled feature map and transmit it to the first DCN-C2f module;
[0023] The second DCN-C2f module is used to extract the multi-scale feature map of the first downsampled feature map, enhance the sampling frequency of the key position features through deformable convolutional operations, splice the extracted multi-scale feature map with the first downsampled feature map to obtain a first multi-scale pose feature map, and transmit it to the second downsampling module;
[0024] The second downsampling module is used to downsample the first multi-scale pose feature map to obtain a second downsampled feature map and transmit it to the second DCN-C2f module;
[0025] The second DCN-C2f module is used to extract the multi-scale feature map of the second downsampled feature map, enhance the sampling frequency of the key position features through deformable convolutional operations, splice the extracted multi-scale feature map with the second downsampled feature map to obtain a second multi-scale pose feature map, and transmit it to the third downsampling module;
[0026] The third downsampling module is used to downsample the second multi-scale pose feature map to obtain a third downsampled feature map, and transmit it to the third DCN-C2f module;
[0027] The third DCN-C2f module is used to extract the multi-scale feature map of the third downsampled feature map, enhance the sampling frequency of the key position features through deformable convolution operation, and splice the extracted multi-scale feature map with the third downsampled feature map to obtain a third multi-scale pose feature map.
[0028] Further, in S1, the steps of constructing the unmanned underwater vehicle dataset include:
[0029] S11: Establish an unmanned underwater vehicle model based on CAD software, and set the key points of the unmanned underwater vehicle model. The key points include the eight endpoints of the minimum bounding box and a center of gravity point;
[0030] S12: Preset the six-degree-of-freedom pose of the unmanned underwater vehicle model based on the Blender Proc engine to obtain a synthetic pose dataset, which includes the key point data of the unmanned underwater vehicle and the pose data containing various morphological features;
[0031] S13: Obtain the UIEB dataset, extract underwater images from the UIEB dataset, and use the underwater images as the background of the pose data to obtain the synthetic images and depth maps of the unmanned underwater vehicle simulating the underwater environment;
[0032] S14: Screen out several unmanned underwater vehicle images from the synthetic images of the unmanned underwater vehicle according to actual requirements and based on the surface shape complexity, texture, and whether it is centrosymmetric;
[0033] S15: Set the transmittance of several unmanned underwater vehicle images, and use the existing underwater scattering model combined with the depth map to calculate the unmanned underwater vehicle image dataset simulating underwater light and shadow.
[0034] Further, the neck network module includes a first Fusion Block module, a second Fusion Block module, a third Fusion Block module, a fourth Fusion Block module, a fifth Fusion Block module, a sixth Fusion Block module, a first splicing module, and a second splicing module;
[0035] The first Fusion Block module is used to perform feature fusion on the first multi-scale pose feature map and the second multi-scale pose feature map to obtain a first fusion feature map, and transmit it to the first splicing module;
[0036] The second Fusion Block module is used to perform feature fusion on the second multi-scale pose feature map and the third multi-scale pose feature map to obtain a second fusion feature map, and transmit it to the first splicing module and the second splicing module;
[0037] The third Fusion Block module is used to extract pose features from the third multi-scale pose feature map to obtain a fourth multi-scale pose feature map, and transmit it to the second splicing module, the fifth Fusion Block module and the sixth Fusion Block module;
[0038] The first splicing module is used to splice the first fusion feature map and the second fusion feature map along the channel dimension to obtain a first spliced feature map, and transmit it to the fourth Fusion Block module;
[0039] The second splicing module is used to splice the second fusion feature map and the fourth multi-scale pose feature map along the channel dimension to obtain a second spliced feature map, and transmit it to the fourth Fusion Block module and the fifth Fusion Block module;
[0040] The sixth Fusion Block module is used to extract pose features from the fourth multi-scale pose feature map to obtain a fifth multi-scale pose feature map, and transmit it to the fifth Fusion Block module and the head network module;
[0041] The fifth Fusion Block module is used to perform feature fusion on the second spliced feature map, the fourth multi-scale pose feature map and the fifth multi-scale pose feature map to obtain a third fusion feature map, and transmit it to the fourth Fusion Block module and the head network module;
[0042] The fourth Fusion Block module is used to perform feature fusion on the first spliced feature map, the second spliced feature map and the third fusion feature map to obtain a fused multi-scale feature map, and transmit it to the head network module.
[0043] Further, each Fusion Block module includes a 1×1 convolutional layer, a BasicBlock module and a Spatial Pyramid Pooling (SPP) module;
[0044] The 1×1 convolutional layer is used to perform dimensionality reduction on the input feature map to obtain a first dimensionality-reduced feature map, and transmit it to the Basic Block module;
[0045] The Basic Block module is used to extract features from the first dimensionality-reduced feature map, obtain feature maps of different scales, and transmit them to the Spatial Pyramid Pooling (SPP) module;
[0046] The Spatial Pyramid Pooling (SPP) module is used to divide the input feature map into grids of different scales, perform max-pooling operations on the features within each grid, and splice the pooling results along the channel dimension.
[0047] Furthermore, the head network module includes a first Pose-head branch, a second Pose-head branch, and a third Pose-head branch, and each output branch includes a 1×1 convolutional layer, an anchor box prediction module, and a key point coordinate prediction module;
[0048] The head network module is used to perform key point regression processing on the fused multi-scale features, output three tensors, and the three tensors respectively correspond to the prediction results of key points at different scales.
[0049] Furthermore, in S5, calculating the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the prediction results of the key points of the actual unmanned underwater vehicle includes:
[0050] S51: Using the non-maximum suppression method to filter the prediction results of the key points of the actual unmanned underwater vehicle to obtain the filtered data;
[0051] S52: Inputting the filtered data into the following formula:
[0052] c(x) = (2δ(x) - 0.5) + c offset (4)
[0053] where x represents the network output value, δ(x) represents the Sigmoid function, and its output range is [0, 1], c offset is the offset of the predicted key point to the upper left corner of a specific grid cell, and c(x) represents the coordinate value of the key point in the pixel coordinate system;
[0054] S53: Obtain the prior coordinates of the key points in the water body coordinate system of the unmanned underwater vehicle and the camera internal parameters, and calculate the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the data results of S52 using the solve PnP algorithm.
[0055] Beneficial effects: The present invention proposes an underwater pose estimation model applied to underwater scenarios, synthesizes a large-scale high-quality unmanned underwater vehicle dataset, and trains the underwater pose estimation model. The trained underwater pose estimation model can achieve high robustness to underwater scene changes and improve the six-degree-of-freedom pose estimation accuracy. The present invention also designs a six-degree-of-freedom intersection over union loss function to establish the constraint relationship between the relative positions of the bounding box corner points, thereby increasing the gradient and adding constraints to further improve the six-degree-of-freedom pose estimation accuracy. Brief Description of the Drawings
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a flowchart of an efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle in the present invention;
[0058] Figure 2 It is a schematic structural diagram of the underwater pose estimation model in the embodiment of the present invention;
[0059] Figure 3 It is a schematic structural diagram of the structure adopting the deformable convolution (DCN) module in the embodiment of the present invention;
[0060] Figure 4 It is a schematic structural diagram of the C2f module reconstructed by using the deformable convolution in the embodiment of the present invention;
[0061] Figure 5 It is an architecture diagram of the Fusion Block module in the embodiment of the present invention;
[0062] Figure 6 It is a schematic diagram of the calculation process of the six-degree-of-freedom intersection over union loss function proposed in the embodiment of the present invention;
[0063] Figure 7 It is a schematic diagram of the synthesized pose dataset constructed in the embodiment of the present invention;
[0064] Figure 8 It is a schematic diagram of the visualization of the pose labels of the synthesized pose dataset constructed in the embodiment of the present invention (the bounding box corner points in the figure are the key points). Detailed Embodiments
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] This embodiment provides an efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle, as Figure 1 shown, and the specific steps include:
[0067] S1: Construct an unmanned underwater vehicle (UUV) dataset, where the unmanned underwater vehicle dataset includes a synthetic pose dataset and an unmanned underwater vehicle image dataset;
[0068] In a specific embodiment, in S1, the steps of constructing the unmanned underwater vehicle dataset include:
[0069] S11: Establish an unmanned underwater vehicle model based on CAD software, and set the key points of the unmanned underwater vehicle model, where the key points include eight endpoints of the minimum bounding box and a center of gravity point;
[0070] S12: Preset the six-degree-of-freedom pose of the unmanned underwater vehicle model based on the Blender Proc engine to obtain a synthetic pose dataset, where the synthetic pose dataset includes unmanned underwater vehicle key point data and pose data containing various morphological features;
[0071] S13: Obtain an underwater image enhancement benchmark dataset (UIEB), extract underwater images from the UIEB dataset, use the underwater images as the background of the pose data to obtain an unmanned underwater vehicle synthetic image and a depth map simulating the underwater environment;
[0072] S14: Screen out a number of unmanned underwater vehicle images from the unmanned underwater vehicle synthetic images according to actual requirements and based on the criteria of surface shape complexity, texture, and whether it is centrosymmetric;
[0073] Specifically, in this embodiment, unmanned underwater vehicle images that meet any one of the conditions of high surface shape complexity, low texture, or centrosymmetry are screened out from the unmanned underwater vehicle synthetic images.
[0074] S15: Set the transmittance of a number of unmanned underwater vehicle images, and use the existing underwater scattering model combined with the depth map to calculate an unmanned underwater vehicle image dataset simulating underwater light and shadow.
[0075] Specifically, in this embodiment, the Blender Proc engine is used in combination with the CAD model of the unmanned underwater vehicle to synthesize (training / validation) data with various morphological features, so as to verify the applicability and domain adaptation capabilities of the model on various UUV morphologies.
[0076] Specifically, as Figure 2 shown in the left half of Figure 7 and Figure 8 , by presetting the 6D pose and rendering the UUV model, and selecting underwater images from the UIEB dataset as the background of the synthetic samples, UUV synthetic images simulating the underwater environment can be obtained. The examples included in the synthesized UUV-6D are as Figure 7 (03), the low-texture UUV Figure 7 (01), and the centrosymmetric UUV Figure 7 (06) and Figure 7 (07); UUV-6D contains 13,500 pictures. By using the stratified sampling method, 12,150 pictures are selected as the training set and 1,350 pictures are selected as the test set. Finally, the obtained dataset is used to verify the applicability of the proposed algorithm to UUVs with different morphologies, and it is used as a reasonable index to measure the robustness of the algorithm.
[0077] S2: Construct an underwater pose estimation model, and train the underwater pose estimation model based on the unmanned underwater vehicle dataset;
[0078] S3: Design the loss function of the underwater pose estimation model as the six-degree-of-freedom intersection over union loss function. When the loss function converges, the trained underwater pose estimation model is obtained;
[0079] Specifically, in this embodiment, the weight model pre-trained on the COCO dataset is used to initialize the backbone part of the network. During the training process, the SGD optimizer with warm-up and cosine learning rate scheduler is used. In each training iteration, the data is input into the model for forward propagation, the loss function value is calculated, then the gradient is calculated by backpropagation, and finally the model parameters are updated according to the gradient. This process continues for multiple training epochs until the model converges. During the training process, the validation set is used to verify the model, check the performance of the model on unseen data, and obtain the best weights according to the performance of the model on the validation set. When the performance of the model on the validation set no longer improves or begins to decline, the training can be stopped and the corresponding model parameters can be selected as the best weights. Deploy it to the high-performance edge computing device Jetson Xavier, and carry the unmanned underwater vehicle to perform real-time pose estimation on other underwater vehicles.
[0080] S4: Use the trained underwater pose estimation model to predict the prediction results of the key points of the actual unmanned submersible;
[0081] S5: Calculate the six-degree-of-freedom pose estimation result of the unmanned submersible based on the prediction results of the key points of the actual unmanned submersible.
[0082] In a specific embodiment, the underwater pose estimation model includes: a backbone network module, a neck network module, and a head network module;
[0083] The backbone network module is established based on the EfficientNet-B4 network architecture, and is used to extract the multi-scale pose features of the unmanned submersible in the unmanned submersible dataset and transmit them to the neck network module;
[0084] Specifically, in this embodiment, EfficientNet-B4 is used as the backbone network (backbone). As Figure 4 shown, a lightweight deformable convolution DCN (Deformable Convolutional Networks) is introduced into the C2f of the backbone network to form a DCN-C2f module. The deformable amount is modulated to make the sampling position of the convolution flexible and variable, so that the model has the ability of geometric transformation and realizes multiple samplings of the key features for positioning.
[0085] The neck network module is built based on the GFPN network and is used to fuse the multi-scale pose features of the unmanned submersible to obtain the fused multi-scale features and transmit them to the head network module;
[0086] Specifically, in this embodiment, GFPN (Generalized-FPN) is adopted in the neck network, which is composed of Fusion Block modules. The special feature fusion method of the Fusion Block module can further make the network notice the local key features sampled by the DCN, thereby further enhancing the key point positioning accuracy.
[0087] The head network module is used to perform key point regression processing on the fused multi-scale features to obtain the prediction results of the key points of the unmanned submersible.
[0088] In a specific embodiment, the present invention designs a six-degree-of-freedom intersection over union loss function 3D-Bbox Loss. The process is as Figure 6 shown. First, map the output result of the network back to the three-dimensional space, then calculate the Intersection over Union (IOU) result between the preset point pairs, and use these results to calculate the Loss value, thereby establishing the constraint relationship between the relative positions of the bounding box corner points, thereby increasing the gradient and adding constraints. The specific design process includes:
[0089] Define \(O\) as the set formed by the planar projections of the key points, and \(P\) as the set formed by the \(Z\)-axis coordinates of each key point in the ground-truth. The specific process is described as follows:
[0090] For \((u, v)\in O\), \(Z\) p \(\in P\):
[0091]
[0092] Repeat the above mapping process for each point in the point set \(O\), then the mapped point set \(P\) of the network output result in the three-dimensional space can be obtained. pr . The obtained point set and the ground-truth point set \(P\) gt are in the same reference frame. Next, a new set of point pairs is obtained as follows:
[0093] Define \(p\) a , \(p\) b \(\in P\) gt , define \(\{p\) a , \(p\) b}\) as a point pair. For all point pairs that satisfy the condition that the line connecting \(p\) a and \(p\) b is the diagonal of the bounding box face, the point pair set \(A\) is given by formula (2). gt Define:
[0094] \(A\) gt \(=\{X|X = \{p\) a , \(p\) b \}\} (2)
[0095] Similarly, the point pair set \(A\) pr is obtained. As shown in Figure 6 (b), the obtained point set \(A\) gt and \(A\) pr tightly enclose the OBB of the target, and the lines connecting them pairwise form the diagonals of each face of the OBB.
[0096] Define the function \(IOU\) proj (i, j). As shown in Figure 6 (c), there are point pairs \(i = (\alpha, \beta)\) and \(j = (\eta, \varepsilon)\). Similar to the previous definition, the line connecting the two points in the point pair forms the diagonal of the bounding box face. Let the rectangle where the point pair \(i\) is located be projected onto the coordinate plane, and the geometric parameters of the projected rectangle are \(I=(x\) α , \(y\) α , \(w\) (α,β) , \(h\) (α,β) ). Similarly, \(J=(x\) η , \(y\) η , \(w\) (η,ε) , \(h\) (η,ε)) Among them, x and y are the horizontal and vertical coordinates of the upper left corner point of the rectangle, w and h are the width and height of the rectangle respectively. Let eps be the regularization term, and the function expression is:
[0097]
[0098] For I pr ∈L pr , I gt ∈L gt , L Bbox is expressed as:
[0099]
[0100] Sum the 12 intersection over union calculation results to obtain formula (4). The number of face diagonals of the entire bounding box is 12. The above operation process utilizes the prior shape feature of the UUV to introduce a constraint on the relative position between corner points for the calculation process of the 3D - Bbox Loss.
[0101] Finally, the expression of the six - degree - of - freedom intersection over union loss function is obtained as:
[0102] Loss = λ pt L pt + λ conf L conf + λ Bbox L Bbox (5)
[0103] Among them, L pt , L conf and L Bbox respectively represent the coordinate, confidence, and bounding box losses; λ pt , λ conf and λ Bbox respectively represent the weight coefficients of the coordinate, confidence, and bounding box losses; smooth L1 is the smooth L1 loss function; IOU proj is the intersection over union result calculated from the projection of the bounding box on the plane; eps is the regularization term; I and J are geometric parameters; i = (α, β), j = (η, ε); where α, β, η, ε respectively correspond to Figure 6 the lower left and upper right of the projection plane in (c).
[0104] In a specific embodiment, the steps of constructing the backbone network module include:
[0105] Deformable Convolution (DCN) is introduced into the C2f modules located in the fourth, sixth, and eighth layers of the EfficientNet-B4 network, thereby replacing the C2f modules with DCN-C2f modules. After replacement, the backbone network module includes a first downsampling module, a first DCN-C2f module, a second downsampling module, a second DCN-C2f module, a third downsampling module, and a third DCN-C2f module;
[0106] Among them, the first downsampling module is used to downsample the input unmanned underwater vehicle dataset to obtain a first downsampled feature map and transmit it to the first DCN-C2f module;
[0107] The second DCN-C2f module is used to extract the multi-scale feature map of the first downsampled feature map, enhance the sampling frequency of key position features through deformable convolution operations, splice the extracted multi-scale feature map with the first downsampled feature map to obtain a first multi-scale pose feature map, and transmit it to the second downsampling module respectively;
[0108] The second downsampling module is used to downsample the first multi-scale pose feature map to obtain a second downsampled feature map and transmit it to the second DCN-C2f module;
[0109] The second DCN-C2f module is used to extract the multi-scale feature map of the second downsampled feature map, enhance the sampling frequency of key position features through deformable convolution operations, splice the extracted multi-scale feature map with the second downsampled feature map to obtain a second multi-scale pose feature map, and transmit it to the third downsampling module respectively;
[0110] The third downsampling module is used to downsample the second multi-scale pose feature map to obtain a third downsampled feature map and transmit it to the third DCN-C2f module;
[0111] The third DCN-C2f module is used to extract the multi-scale feature map of the third downsampled feature map, enhance the sampling frequency of key position features through deformable convolution operations, and splice the extracted multi-scale feature map with the third downsampled feature map to obtain a third multi-scale pose feature map.
[0112] Specifically, as Figure 3As shown, the DCN has a variable receptive field. The DCN adds an offset to the sampling operation in the standard convolution operation. The offset is obtained based on the (Offset field) after the input feature map passes through the convolution operation (Conv), and can adaptively select the sampling position according to the feature distribution in the feature map. The feature map (2N dim) generated by the DCN models the offsets (Offsets) of the sampling points in each channel of the input feature map (N dim). Subsequently, the offset is added to the standard sampling position to obtain a new sampling position. This process is described by formula (6) as follows:
[0113]
[0114] Specifically, in Figure 3 , K = 9 and p k ∈{(-1, -1), (-1, 0),..., (1, 1)} define a 3×3 convolution kernel. p k enumerates the standard sampling positions in K, while Δp k is the offset at the corresponding position. Among them, x(p) is implemented by bilinear interpolation and is described by formula (7) as follows:
[0115]
[0116] Among them, q is the value at each position on the feature map x; G(q, p) is a two-dimensional bilinear interpolation kernel and is described by formula (8) as follows:
[0117] G(q, p) = g(q x , p x )·g(q y , p y ) (8)
[0118] Among them, g(a, b) = max(0, 1 - |a - b|).
[0119] To solve the problem of redundant sampling regions in the DCN, a weight coefficient Δm k is added to the sampling calculation for each sampling point and can be expressed as:
[0120]
[0121] Among them, Δm k is the weight coefficient, within the range of [0, 1], and is the N-dimensional feature map obtained by the convolution operation during the actual inference process. At this time, the total dimension of the feature map output by the convolution module is 3N, including the x-direction offset (N dim), the y-direction offset (N dim), and the weight coefficient (N dim).
[0122] In a specific embodiment, the neck network module includes a first Fusion Block module, a second Fusion Block module, a third Fusion Block module, a fourth Fusion Block module, a fifth Fusion Block module, a sixth Fusion Block module, a first splicing module, and a second splicing module;
[0123] The first Fusion Block module is used to perform feature fusion on the first multi-scale pose feature map and the second multi-scale pose feature map to obtain a first fusion feature map and transmit it to the first splicing module;
[0124] The second Fusion Block module is used to perform feature fusion on the second multi-scale pose feature map and the third multi-scale pose feature map to obtain a second fusion feature map and transmit it to the first splicing module and the second splicing module;
[0125] The third Fusion Block module is used to extract pose features from the third multi-scale pose feature map to obtain a fourth multi-scale pose feature map and transmit it to the second splicing module, the fifth Fusion Block module, and the sixth Fusion Block module;
[0126] The first splicing module is used to splice the first fusion feature map and the second fusion feature map along the channel dimension to obtain a first spliced feature map and transmit it to the fourth Fusion Block module;
[0127] The second splicing module is used to splice the second fusion feature map and the fourth multi-scale pose feature map along the channel dimension to obtain a second spliced feature map and transmit it to the fourth Fusion Block module and the fifth Fusion Block module;
[0128] The sixth Fusion Block module is used to extract pose features from the fourth multi-scale pose feature map to obtain a fifth multi-scale pose feature map and transmit it to the fifth Fusion Block module and the head network module;
[0129] The fifth Fusion Block module is used to perform feature fusion on the second spliced feature map, the fourth multi-scale pose feature map, and the fifth multi-scale pose feature map to obtain a third fusion feature map and transmit it to the fourth Fusion Block module and the head network module;
[0130] The fourth Fusion Block module is used to perform feature fusion on the first spliced feature map, the second spliced feature map, and the third fused feature map to obtain a fused multi-scale feature map, and transmit it to the head network module.
[0131] Specifically, after being processed by the Fusion Block module multiple times, the neck network module can effectively fuse the features of different scales of the unmanned underwater vehicle to obtain fused multi-scale features, thereby enhancing the robustness and adaptability of the model to different shape features.
[0132] In a specific embodiment, each Fusion Block module includes a 1×1 convolutional layer, a BasicBlock module, and a spatial pyramid pooling SPP module. The calculation process of the Fusion Block module is as follows:
[0133] p ′ =Conv 1×1 (Concat(Conv 1×1 (p),BasicBlock ×N (p),SPP(p))) (10)
[0134] The 1×1 convolutional layer is used to perform dimensionality reduction on the input feature map to obtain a first dimensionality-reduced feature map, and transmit it to the Basic Block module;
[0135] The Basic Block module is used to extract features from the first dimensionality-reduced feature map to obtain feature maps of different scales, and transmit them to the spatial pyramid pooling SPP module. The calculation process of the Basic Block module is as follows:
[0136] BasicBlock(p)=p+Conv 3×3 (Repconv(p)) (11)
[0137] The spatial pyramid pooling SPP module is used to divide the input feature map into grids of different scales, perform max-pooling operations on the features within each grid, that is, perform pooling operations of different scales on the input feature map, and finally splice the pooling results along the channel dimension. The calculation process of the spatial pyramid pooling SPP module is as follows:
[0138] SPP(p)=Conv 1×1 (Concat(p,maxpool(p), maxpool(p), maxpool(p))) (12)
[0139] Specifically, the Fusion Block is used to form the Neck network, and the organization method of the feature flow graph refers to the GFPN network. In the Fusion Block, its special feature fusion method can further enable the network to notice the local key features sampled by the DCN. The attention received by the located key features will be further superimposed, while the irrelevant features are suppressed in contrast. This indicates that the synergistic effect of the two structures can theoretically bring better pose estimation ability to the network.
[0140] Specifically, the Fusion Block is as Figure 5 shown, adopting the CSP structure, and the core part is composed of Repconv using the reparameterization technique.
[0141] In a specific embodiment, the head network module includes a first Pose-head branch, a second Pose-head branch, and a third Pose-head branch, and each output branch includes a 1×1 convolutional layer, an anchor box prediction module, and a key point coordinate prediction module;
[0142] The head network module is used to perform key point regression processing on the fused multi-scale features, output three tensors, and the three tensors respectively correspond to the prediction results of the key points at different scales.
[0143] Further, the tensors at the three scales respectively generate a grid of size W×H×(L×n a ), where n a is the preset number of anchor boxes. In this embodiment, n a =3, that is, 3 anchor boxes are generated for each grid cell on each grid at each scale. For an input image of 640×640, H×W corresponds to 80×80, 40×40, and 20×20 at the three scales respectively. That is, the network will generate (80×80×3 + 40×40×3 + 20×20×3) = 20880 prediction candidate boxes, and each prediction candidate box includes L elements, where L=(2×9 + 1 + n cls ), corresponding to the x and y coordinate values of 9 key points, a confidence value, and n cls class probability values respectively.
[0144] Among them, the confidence value describes the probability that there is a UUV target in the current grid, and the class probability value reflects the probability of the UUV class existing in the grid cell. In the inference stage, first, the grid cell corresponding to the maximum confidence value is screened out, and then the values contained in this grid cell are used to model the key point coordinates. During the training process, the confidence function proposed by Tekin et al. (2018) is used to standardize the network to learn how to judge the probability that there is a target in a certain grid area, which is expressed as:
[0145]
[0146] Among them, D T (x) represents the distance from the model at the position to the detected object, and d th is a threshold value, indicating that when the distance is less than this threshold value, the detection result is credible; α is an adjustable parameter used to control the growth rate of the confidence, and p(x) is the confidence value (probability value) that there is a target in the grid cell.
[0147] In a specific embodiment, in S5, the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle calculated based on the prediction result of the key points of the actual unmanned underwater vehicle includes:
[0148] S51: Use the non-maximum suppression method to filter the prediction result of the key points of the actual unmanned underwater vehicle to obtain the filtered data;
[0149] S52: Input the filtered data into the following formula:
[0150] c(x) = (2δ(x) - 0.5) + c offset (14)
[0151] Among them, x represents the network output value, δ(x) represents the Sigmoid function, and its output range is [0, 1], and c offset is the offset of the predicted key point to the upper left corner of the specific grid cell, and c(x) represents the coordinate value of the key point in the pixel coordinate system;
[0152] S53: Obtain the prior coordinates of the key points in the water body coordinate system of the unmanned underwater vehicle and the camera internal parameters, and calculate the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the data result of S52 and using the solve PnP algorithm.
[0153] Specifically, in this embodiment, the homogeneous transformation matrix of the UUV, that is, the six-degree-of-freedom pose estimation result, is calculated through the slove PnP algorithm.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle, characterized in that: The specific steps include: S1: constructing an unmanned underwater vehicle dataset, wherein the unmanned underwater vehicle dataset includes a synthetic pose dataset and an unmanned underwater vehicle image dataset; S2: constructing an underwater pose estimation model, and training the underwater pose estimation model based on the unmanned underwater vehicle dataset; S3: The loss function of the underwater pose estimation model is designed to be a six-degree-of-freedom intersection-and-combination loss function. When the loss function converges, the trained underwater pose estimation model is obtained. S4: Use the trained underwater pose estimation model to predict the key points of the actual unmanned underwater vehicle; S5: Calculate the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the prediction results of the key points of the actual unmanned underwater vehicle.
2. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 1, characterized in that: The underwater posture estimation model includes: a backbone network module, a neck network module and a head network module; The backbone network module is established based on the EfficientNet-B4 network architecture, and is used to extract multi-scale posture features of the unmanned underwater vehicle in the unmanned underwater vehicle data set, and transmit them to the neck network module; The neck network module is built based on the GFPN network and is used to fuse the multi-scale posture features of the unmanned underwater vehicle to obtain fused multi-scale features and transmit them to the head network module; The head network module is used to perform key point regression processing on the fused multi-scale features to obtain prediction results of the key points of the unmanned underwater vehicle.
3. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 1, characterized in that: The expression of the intersection-over-union loss function under the six degrees of freedom is: Loss=λ pt L pt +λ conf L conf +λ Bbox L Bbox (1) Among them, L pt , L conf With L Bbox Represents coordinates, confidence, and bounding box loss respectively; λ pt , conf With λ Bbox Respectively represent the weight coefficients of coordinates, confidence, and bounding box loss; smooth L1 is the smooth L1 loss function; IOU proj is the intersection-and-union ratio result calculated by the projection of the bounding box on the surface; eps is the regularization term; I and J are geometric parameters; i and j are defined point pairs.
4. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 2, characterized in that: The steps of constructing the backbone network module include: Introducing deformable convolution DCN in the C2f modules located in the fourth, sixth and eighth layers of the EfficientNet-B4 network, thereby replacing the C2f modules with DCN-C2f modules. After the replacement, the backbone network module includes a first downsampling module, a first DCN-C2f module, a second downsampling module, a second DCN-C2f module, a third downsampling module and a third DCN-C2f module; The first downsampling module is used to downsample the input unmanned underwater vehicle data set to obtain a first downsampling feature map, and transmit it to the first DCN-C2f module; The second DCN-C2f module is used to extract a multi-scale feature map of the first down-sampling feature map, and enhance the sampling frequency of key position features through a deformable convolution operation, splice the extracted multi-scale feature map with the first down-sampling feature map to obtain a first multi-scale posture feature map, and transmit it to the second down-sampling module; The second downsampling module is used to downsample the first multi-scale pose feature map to obtain a second downsampled feature map, and transmit it to the second DCN-C2f module; The second DCN-C2f module is used to extract a multi-scale feature map of the second down-sampled feature map, and enhance the sampling frequency of key position features through a deformable convolution operation, and concatenate the extracted multi-scale feature map with the second down-sampled feature map to obtain a second multi-scale pose feature map, and transmit it to the third down-sampling module; The third downsampling module is used to downsample the second multi-scale pose feature map to obtain a third downsampled feature map, and transmit it to the third DCN-C2f module; The third DCN-C2f module is used to extract a multi-scale feature map of the third down-sampled feature map, and enhance the sampling frequency of key position features through a deformable convolution operation, and splice the extracted multi-scale feature map with the third down-sampled feature map to obtain a third multi-scale posture feature map.
5. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 1, characterized in that: In S1, the steps to construct the unmanned underwater vehicle dataset include: S11: establishing an unmanned underwater vehicle model based on CAD software, and setting key points of the unmanned underwater vehicle model, wherein the key points include eight endpoints and a center of gravity of a minimum tight bounding box; S12: Preset the six-degree-of-freedom pose of the unmanned underwater vehicle model based on the Blender Proc engine to obtain a synthetic pose data set, wherein the synthetic pose data set includes key point data of the unmanned underwater vehicle and pose data including multiple morphological features; S13: Acquire a UIEB data set, extract an underwater image from the UIEB data set, use the underwater image as the background of the pose data, and obtain a synthetic image of the unmanned underwater vehicle simulating an underwater environment and a depth map; S14: selecting a plurality of unmanned underwater vehicle images from the unmanned underwater vehicle synthetic image according to actual needs and based on the surface shape complexity, texture and whether the center is symmetrical; S15: setting the transmittance of several unmanned underwater vehicle images, and using an existing underwater scattering model in combination with the depth map to calculate an unmanned underwater vehicle image dataset that simulates underwater light and shadow.
6. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 4, characterized in that: The neck network module includes a first Fusion Block module, a second Fusion Block module, a third FusionBlock module, a fourth Fusion Block module, a fifth Fusion Block module, a sixth Fusion Block module, a first splicing module, and a second splicing module; The first Fusion Block module is used to perform feature fusion on the first multi-scale posture feature map and the second multi-scale posture feature map to obtain a first fused feature map, and transmit it to the first splicing module; The second Fusion Block module is used to perform feature fusion on the second multi-scale pose feature map and the third multi-scale pose feature map to obtain a second fused feature map, and transmit the second fused feature map to the first splicing module and the second splicing module; The third Fusion Block module is used to extract posture features from the third multi-scale posture feature map to obtain a fourth multi-scale posture feature map, and transmit it to the second splicing module, the fifth Fusion Block module and the sixth FusionBlock module; The first splicing module is used to splice the first fusion feature map and the second fusion feature map according to the channel dimension to obtain a first splicing feature map, and transmit it to the fourth Fusion Block module; The second splicing module is used to splice the second fusion feature map and the fourth multi-scale pose feature map according to the channel dimension to obtain a second splicing feature map, and transmit it to the fourth Fusion Block module and the fifth Fusion Block module; The sixth Fusion Block module is used to extract posture features from the fourth multi-scale posture feature map to obtain a fifth multi-scale posture feature map, and transmit it to the fifth Fusion Block module and the head network module; The fifth Fusion Block module is used to perform feature fusion on the second splicing feature map, the fourth multi-scale pose feature map and the fifth multi-scale pose feature map to obtain a third fused feature map, and transmit it to the fourth Fusion Block module and the head network module; The fourth Fusion Block module is used to perform feature fusion on the first spliced feature map, the second spliced feature map and the third fused feature map to obtain a fused multi-scale feature map, and transmit it to the head network module.
7. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 6, characterized in that: Each Fusion Block module includes a 1×1 convolution layer, a BasicBlock module, and a spatial pyramid pooling SPP module; The 1×1 convolution layer is used to perform dimensionality reduction processing on the input feature map to obtain a first dimensionality reduction feature map, and transmit it to the Basic Block module; The Basic Block module is used to extract features from the first dimension reduction feature map to obtain feature maps of different scales, and transmit them to the spatial pyramid pooling SPP module; The spatial pyramid pooling SPP module is used to divide the input feature map into grids of different scales, perform a maximum pooling operation on the features in each grid, and splice the pooling results along the channel dimension.
8. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 6, characterized in that: The head network module includes a first Pose-head branch, a second Pose-head branch and a third Pose-head branch, and each output branch includes a 1×1 convolution layer, an anchor box prediction module and a key point coordinate prediction module; The head network module is used to perform key point regression processing on the fused multi-scale features and output three tensors, and the three tensors correspond to the prediction results of the key points at different scales.
9. The efficient six-degree-of-freedom pose estimation method for an unmanned underwater vehicle according to claim 1, characterized in that: In S5, the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle is calculated based on the prediction result of the key points of the actual unmanned underwater vehicle, including: S51: filtering the prediction results of the key points of the actual unmanned underwater vehicle using a non-maximum suppression method to obtain filtered data; S52: Input the filtered data into the following formula: c(x)=(2δ(x) - 0.5)+ c offset (4) Among them, x represents the network output value, δ(x) represents the Sigmoid function, whose output range is [0,1], and c offset is the offset from the predicted key point to the upper left corner of a specific grid cell, and c(x) represents the coordinate value of the key point in the pixel coordinate system; S53: Obtain the prior coordinates of the key points in the water coordinate system of the unmanned underwater vehicle and the camera internal parameters, and calculate the six-degree-of-freedom pose estimation result of the unmanned underwater vehicle based on the data results of S52 and using the solve PnP algorithm.