A method and system for six-dimensional pose estimation of an object applicable to harsh environments
By using traditional and deep learning methods to enhance images in pose estimation and combining self-coded fusion grids for image fusion, the problem of insufficient robustness of pose estimation in harsh environments is solved, and a six-dimensional pose estimation with high accuracy and reliability in harsh environments is achieved.
Patent Information
- Application Number
- CN202210962731.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-11
AI Technical Summary
The existing six-dimensional pose estimation methods for objects are insufficient in operation efficiency and adaptability in harsh environments (such as foggy days and low-light conditions), resulting in weak robustness in pose estimation.
The image is enhanced based on traditional and deep learning methods, combined with self-coded fusion grids for image fusion, and six-dimensional pose estimation is performed through feature extraction, semantic segmentation, key point prediction and regression pose.
The accuracy and reliability of six-dimensional pose estimation are significantly improved in harsh environments, ensuring the normal progress of autonomous driving and robot grasping tasks under night and foggy conditions.
Smart Images

Figure CN115294433B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing and computer vision, and particularly relates to a method and system for six-dimensional pose estimation of an object applicable to harsh environments. Background Art
[0002] Six-dimensional pose (three degrees of freedom of displacement and three degrees of freedom of rotation) is a relative concept, referring to the displacement and rotation transformation between two coordinate systems. For the six-dimensional pose estimation of an object, the rotation and translation transformation of the object from the world coordinate system to the camera coordinate system is usually used. Six-dimensional pose estimation is an important part in many real-world applications, such as augmented reality, autonomous driving, and robotic grasping. However, in harsh environments (such as foggy days and low-light conditions), the image details are not obvious, and optical imaging faces problems such as poor visibility and a lot of noise, which pose great challenges to pose estimation.
[0003] Existing methods for six-dimensional pose estimation of an object can generally be divided into three categories: methods based on point cloud matching, methods based on template matching, and methods based on deep learning. In harsh environments such as foggy days or low light, due to the influence of image noise, there will be large errors in key point matching for these methods, so the robustness of pose estimation in harsh environments such as foggy days and low-light conditions is weak. Therefore, it is of great significance to adopt a six-dimensional pose estimation method that adapts to harsh environments. Summary of the Invention
[0004] Aiming at the deficiencies in the operating efficiency and adaptability of existing six-dimensional pose estimation methods in harsh environments, the present invention provides a six-dimensional pose estimation method and system that can adapt to harsh environments.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for six-dimensional pose estimation of an object applicable to harsh environments, including the following steps:
[0007] Step 1, enhancing the image by using both traditional and deep learning methods;
[0008] Step 2, performing image fusion by using an auto-encoding fusion grid;
[0009] Step 3, performing six-dimensional pose estimation through feature extraction, semantic segmentation, key point prediction, and pose regression.
[0010] Further, in the above Step 1, enhancing the image by using the traditional method is to use an image enhancement sub-module composed of several differentiable filters and a small convolutional neural network for predicting filter hyperparameters, and the image enhancement sub-module includes a sharpening filter and a defogging filter;
[0011] In the defogging filter, a fog image formation model described by the following equation is adopted:
[0012] I(x) = J(x)t(x) + A(1 - t(x)) (1)
[0013] Wherein, I(x) is the input image, J(x) is the defogged image output, A is the global atmospheric light component, and t(x) is the transmittance;
[0014] According to the formula, an approximate value of t(x) can be obtained:
[0015]
[0016] Wherein, C represents the RGB three channels;
[0017] A parameter λ is introduced to control the degree of defogging:
[0018]
[0019] Since the above operation is differentiable, λ can be optimized through backpropagation to make the defogging filter more conducive to pose estimation;
[0020] In the sharpening filter, the sharpening of the image can highlight the details of the image, and the sharpening process can be expressed as:
[0021] F(x, η) = I(x) + η(I(x) - Gau(I(x))) (4)
[0022] Wherein, I(x) is the input image, Gau(I(x)) represents the Gaussian filter, η is a positive scaling factor, and this sharpening operation is differentiable with respect to both x and η, and x and η can be optimized through backpropagation;
[0023] The small convolutional neural network for predicting the filter hyperparameters consists of 4 convolutional blocks and 2 fully connected layers. Each convolutional block includes a 3×3 convolutional layer with a stride of 2 and a leaky Relu activation function. The output channels of these four convolutional layers are 16, 32, 32, and 32 respectively; the input of the convolutional neural network is the image in the harsh environment, and the output of the last fully connected layer is the predicted hyperparameters of various filters.
[0024] Furthermore, the image enhancement based on the deep learning method in step 1 is implemented based on the generative adversarial network method. The generative adversarial network includes two parts: a generative network and a discriminative network; wherein:
[0025] The generation network model consists of 16 layers. The first half consists of 6 convolutional layers and 2 pooling layers. After each convolutional layer, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 32, 32, 64, 64, 128, and 128. Pooling layers are added after the 3rd and 6th convolutional layers respectively. The second half consists of 8 deconvolutional layers. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 256, 256, 128, 128, 64, 64, 32, and 3. Through convolutional and deconvolutional operations, the weight parameters are adjusted to achieve the effect of image enhancement.
[0026] The discriminant network model consists of a fully convolutional network, including a total of 5 convolutional layers. After the first 4 convolutional layers, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 1, and the number of channels is 42, 96, 192, 384, and 3. A sigmoid activation function is added at the end of the network for feature mapping to normalize the results.
[0027] Furthermore, the specific process of implementing image enhancement based on the generative adversarial network is as follows: The image under harsh environmental conditions is input into the generation network. After the convolutional and deconvolutional operations of the generation network, an enhanced image is obtained. Then, the enhanced image and the image under normal conditions are input into the discriminant network for discrimination to distinguish between true and false and output a probability. When the output probability value is close to 1, it indicates that the input is an image under normal lighting conditions. When the discriminator cannot determine true or false, the image generated by the generation network at this time is the optimal image.
[0028] Let {m i , i = 1, 2,..., N} and {n i , i = 1, 2,..., N} represent the images under harsh environments and the images under normal conditions respectively. The adversarial loss can be defined as:
[0029]
[0030] where G represents the generation network and D represents the discriminant network.
[0031] The mean square error loss of the network model can be defined as:
[0032]
[0033] Finally, the adversarial loss and the mean square error loss are combined and configured with certain weights α and β to obtain the loss of the final generation network:
[0034] L t = αL a + βL m (7)
[0035] The loss of the discrimination network can be defined as:
[0036]
[0037] Further, the specific process of using the auto - encoding fusion grid for image fusion in step 2 is as follows: The pictures to be fused are input into the encoding layer. Through two convolutions, the convolution kernel size is 2×2, and the stride is 1. The output of the encoding layer is the input of the fusion layer. Then, in the fusion layer, the features of the hidden layer are fused using the Addition strategy. The output of the fusion layer is the input of the decoding layer. The decoding layer consists of three convolution operations, the convolution kernel size is 2×2, and the stride is 1. To ensure that the extraction of image detail features is not lost, there is no pooling operation in the auto - encoding fusion network.
[0038] Further, in step 3, the Darknet53 network model is used for feature extraction. The input of the network is the picture that has been enhanced by the filter, and the output is the features of the picture, which are used for subsequent semantic segmentation and key - point prediction.
[0039] Further, in step 3, semantic segmentation assigns a label to each pixel point superimposed on the image to distinguish different objects. More precisely, given N object classes, this is transformed into outputting a vector of dimension N + 1 at each spatial position, plus one dimension to represent the background;
[0040] The loss function is:
[0041]
[0042] where M represents the number of categories; y c is an indicator variable, 0 or 1. If the category is the same as the category of the sample, it is 1, otherwise it is 0; p c represents the predicted probability that the observed sample belongs to class c.
[0043] Further, in step 3, the SIFT algorithm is used for key - point prediction to detect the distinctive two - dimensional key points in the texture image and lift them to three - dimensional. Then, the FPS algorithm is applied to select the top N key points. In this way, the selected key points are not only evenly distributed on the object surface but also have distinct texture features and are easy to detect;
[0044] During the process of key - point prediction, for each pixel point, the offset d i (x) relative to the two - dimensional key point of the object to which it belongs is predicted. Let the two - dimensional position of the pixel point be d, and the true position of the two - dimensional key point be d i , and P be the segmentation mask. Then the loss during the training process is:
[0045]
[0046] Meanwhile, the confidence of each prediction point will also be output, which is obtained through the sigmoid function output by the network. For each 3D key point, 20 two-dimensional positions with the highest confidence are selected as candidate points for subsequent pose calculation.
[0047] In step 3, the regression pose is calculated based on the PnP algorithm of RANSAC to obtain the accurate six-dimensional pose of the object.
[0048] The present invention also provides a six-dimensional pose estimation system for objects suitable for harsh environments, which is used to implement the above-mentioned six-dimensional pose estimation method for objects suitable for harsh environments. It includes a computer memory and a processor, an image enhancement module, an image fusion module, and a six-dimensional pose estimation module. The image enhancement module enhances the picture based on both traditional and deep learning methods. The image fusion module uses an autoencoder fusion network to fuse the enhanced pictures. The six-dimensional pose estimation module estimates the six-dimensional pose of the object in a harsh environment through feature extraction, semantic segmentation, key point prediction, and regression pose. All specific data processing and calculation work in all modules are completed by the computer processor, and all units interact with the data in the computer memory.
[0049] Compared with the prior art, the present invention has the following advantages:
[0050] 1. By adding an image enhancement module before pose estimation, the present invention can complete six-dimensional pose estimation in harsh environments (such as foggy days and low light conditions).
[0051] 2. By fusing the results of traditional image enhancement methods and deep learning image enhancement methods, the spatio-temporal information contained in the image is expanded, uncertainty is reduced, and reliability is increased.
[0052] 3. The method of the present invention is easy to implement, and its application value is mainly reflected in the following aspects:
[0053] (1) It can ensure the safety and reliability of autonomous driving technology in harsh environments such as at night and on foggy days.
[0054] (2) It can ensure that the robot can normally complete the object grasping task in harsh environments such as at night and on foggy days. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the system framework diagram of the six-dimensional pose estimation method for objects suitable for harsh environments of the present invention;
[0056] Figure 2 is the image fusion flowchart;
[0057] Figure 3 is the picture in a harsh environment;
[0058] Figure 4 For the enhanced image;
[0059] Figure 5 For the enhanced pose estimation result;
[0060] Figure 6 For the pose estimation result of the existing method. Detailed implementation mode
[0061] The technical solution of the present invention will be specifically and detailedly described below in conjunction with the embodiments and drawings of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several modifications and improvements can also be made, which should also be regarded as belonging to the protection scope of the present invention.
[0062] A six-dimensional pose estimation method for objects suitable for harsh environments mainly consists of three parts: image enhancement, image fusion, and six-dimensional pose estimation. This method uses both traditional and deep learning methods to enhance the image, and then fuses the enhanced image using an autoencoder fusion grid. After fusion, it is input into the pose estimation part for pose estimation. The specific process is as Figure 1 shown.
[0063] 1. Enhance the image using both traditional and deep learning methods;
[0064] 1.1 Image enhancement based on traditional methods: Use an image enhancement sub-module composed of several differentiable filters and a small convolutional neural network for predicting filter hyperparameters. The image enhancement sub-module includes a sharpening filter and a defogging filter;
[0065] (1) In the defogging filter, use the fog image formation model described by the following equation:
[0066] I(x) = J(x)t(x) + A(1 - t(x)) (1)
[0067] In the formula, I(x) is the input image, J(x) is the output fog-free image, A is the global atmospheric light component, and t(x) is the transmittance;
[0068] According to the formula, an approximate value of t(x) can be obtained:
[0069]
[0070] In the formula, C represents the RGB three channels;
[0071] Introduce a parameter λ to control the degree of defogging:
[0072]
[0073] Since the above operation is differentiable, λ can be optimized through backpropagation to make the defogging filter more conducive to pose estimation;
[0074] (2) In the sharpening filter, the sharpening of the image can highlight the details of the image, and the sharpening process can be expressed as:
[0075] F(x,η) = I(x) + η(I(x) - Gau(I(x))) (4)
[0076] In the formula, I(x) is the input image, Gau(I(x)) represents the Gaussian filter, and η is a positive scaling factor. This sharpening operation is differentiable with respect to both x and η, and x and η can be optimized through backpropagation;
[0077] (3) The small convolutional neural network used to predict the filter hyperparameters consists of 4 convolutional blocks and 2 fully connected layers. Each convolutional block includes a 3×3 convolutional layer with a stride of 2 and a leaky Relu activation function. The output channels of these four convolutional layers are 16, 32, 32, and 32 respectively. The input of the convolutional neural network is the image in the harsh environment, and the output of the last fully connected layer is the predicted hyperparameters of various filters.
[0078] 1.2 Image enhancement based on the deep learning method is implemented based on the generative adversarial network method. The generative adversarial network includes two parts: the generative network and the discriminative network. Among them: The generative network model consists of 16 layers. The first half consists of 6 convolutional layers and 2 pooling layers. After each convolutional layer, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 32, 32, 64, 64, 128, and 128. Pooling layers are added after the 3rd convolutional layer and the 6th convolutional layer respectively. The second half consists of 8 deconvolutional layers. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 256, 256, 128, 128, 64, 64, 32, and 3. Through convolutional and deconvolutional operations, the weight parameters are adjusted to achieve the effect of image enhancement. The discriminative network model consists of a fully convolutional network, which includes a total of 5 convolutional layers. After the first 4 convolutional layers, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 1, and the number of channels is 42, 96, 192, 384, and 3. A sigmoid activation function is added at the end of the network for feature mapping to normalize the results.
[0079] The specific process of image enhancement is: The image under harsh environmental conditions ( Figure 3)In the input generation network, an enhanced image is obtained through the convolution and transposed convolution operations of the generation network. Then, the enhanced image and the image under normal conditions are input into the discriminative network for discrimination to distinguish between true and false, and a probability is output. When the output probability value is close to 1, it indicates that the input is an image under normal lighting conditions. When the discriminator cannot determine true or false, the image generated by the generation network at this time is the optimal image. Figure 4 )
[0080] Let {m i , i = 1, 2,..., N} and {n i , i = 1, 2,..., N} represent the images in harsh environments and the images under normal conditions respectively. The adversarial loss can be defined as:
[0081]
[0082] where G represents the generation network and D represents the discriminative network;
[0083] The mean squared error loss of the network model can be defined as:
[0084]
[0085] Finally, the adversarial loss and the mean squared error loss are combined and configured with certain weights α and β to obtain the loss of the final generation network:
[0086] L t = αL a + βL m (7)
[0087] The loss of the discriminative network can be defined as:
[0088]
[0089] 2. Use an auto - encoding fusion grid for image fusion;
[0090] The images to be fused are input into the encoding layer. Through two convolutions, the kernel size is 2×2 and the stride is 1. The output of the encoding layer is the input of the fusion layer. Then, in the fusion layer, the features of the hidden layer are fused using the Addition strategy. The output of the fusion layer is the input of the decoding layer. The decoding layer consists of three convolution operations, the kernel size is 2×2 and the stride is 1. To ensure that the extraction of image detail features is not lost, there is no pooling operation in the auto - encoding fusion network. The fusion process is as Figure 2 shown.
[0091] 3. Perform six - dimensional pose estimation through feature extraction, semantic segmentation, key - point prediction, and regression pose.
[0092] 3.1 Feature extraction: The Darknet53 network model is used. The input of the network is the image that has been enhanced by filters, and the output is the features of the image, which are used for subsequent semantic segmentation and key point prediction.
[0093] 3.2 Semantic segmentation: A label is assigned to each pixel point superimposed on the image to distinguish different objects. More precisely, given N object classes, this is transformed into outputting a vector of dimension N + 1 at each spatial position, plus one dimension to represent the background;
[0094] The loss function is:
[0095]
[0096] where M represents the number of classes; y c is an indicator variable, 0 or 1, which is 1 if the class is the same as the class of the sample, otherwise it is 0; p c represents the predicted probability that the observed sample belongs to class c.
[0097] 3.3 Key point prediction: The SIFT algorithm is used to detect distinctive two-dimensional key points in the texture image and lift them to three dimensions; then the FPS algorithm is applied to select the top N key points among them. In this way, the selected key points are not only evenly distributed on the object surface, but also have distinct texture features and are easy to detect;
[0098] During the process of key point prediction, for each pixel point, the offset d i (x) relative to the two-dimensional key point of the object it belongs to is predicted. Let the two-dimensional position of the pixel point be d, and the true position of the two-dimensional key point be d i , and P be the segmentation mask, then the loss during the training process is:
[0099]
[0100] At the same time, the confidence of each predicted point is also output, which is obtained through the sigmoid function output by the network. For each three-dimensional key point, 20 two-dimensional positions with the highest confidence are selected as candidate points for subsequent pose calculation.
[0101] 3.4 Regression pose is calculated based on the PnP algorithm of RANSAC to obtain the accurate six-dimensional pose of the object. Figure 6 is the pose estimation result of the existing method in harsh environments (such as foggy days and low light conditions), Figure 5 is the enhanced pose estimation result of the method of the present invention. Compared with Figure 3 the pictures in harsh environments, it shows that the method of the present invention can well complete six-dimensional pose estimation in harsh environments (such as foggy days and low light conditions).
[0102] A method for implementing the above-mentioned six-dimensional pose estimation of an object, including a computer memory, a processor, an image enhancement module, an image fusion module, and a six-dimensional pose estimation module; the image enhancement module enhances pictures based on both traditional and deep learning methods, the image fusion module uses an auto-encoder fusion network to fuse the enhanced pictures, and the six-dimensional pose estimation module estimates the six-dimensional pose of an object in a harsh environment through feature extraction, semantic segmentation, key point prediction, and pose regression. The specific data processing and calculation work in all modules are completed by the computer processor, and all units interact with the data in the computer memory.
Claims
1. A six-dimensional pose estimation method for objects applicable to harsh environments, characterized in that, It includes the following steps: Step 1, enhance the image by using both traditional and deep learning methods; Step 2, perform image fusion by using an auto-encoding fusion grid; Step 3, perform six-dimensional pose estimation through feature extraction, semantic segmentation, key point prediction, and regression pose; The specific process of performing image fusion by using an auto-encoding fusion grid in Step 2 is as follows: Input the pictures to be fused into the encoding layer, and perform two convolutional operations with a convolutional kernel size of 2×2 and a stride of 1; the output of the encoding layer is the input of the fusion layer, and then the features of the hidden layer are fused by using the Addition strategy in the fusion layer; the output of the fusion layer is the input of the decoding layer, and the decoding layer consists of three convolutional operations with a convolutional kernel size of 2×2 and a stride of 1; in order to ensure that the extraction of image detail features is not lost, there is no pooling operation in the auto-encoding fusion network.
2. The six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, wherein In Step 1, the enhancement of the image based on the traditional method is achieved by using an image enhancement sub-module composed of several differentiable filters and a small convolutional neural network for predicting the hyperparameters of the filters. The image enhancement sub-module includes a sharpening filter and a defogging filter; In the defogging filter, the fog image formation model described by the following equation is adopted: I(x) = J(x)t(x) + A(1 - t(x)) (1) In the formula, I(x) is the input image, J(x) is the fog-free image output, A is the global atmospheric light component, and t(x) is the transmittance; According to the formula, an approximate value of t(x) can be obtained: In the formula, C represents the three RGB channels; A parameter λ is introduced to control the degree of defogging: Since the above operation is differentiable, λ can be optimized through backpropagation to make the defogging filter more conducive to pose estimation; In the sharpening filter, the sharpening of the image can highlight the details of the image, and the sharpening process can be expressed as: F(x, η) = I(x) + η(I(x) - Gau(I(x))) (4) In the formula, I(x) is the input image, Gau(I(x)) represents the Gaussian filter, and η is a positive scaling factor. This sharpening operation is differentiable with respect to both x and η, and x and η can be optimized through backpropagation; The small convolutional neural network for predicting the hyperparameters of the filters consists of 4 convolutional blocks and 2 fully connected layers. Each convolutional block includes a 3×3 convolutional layer with a stride of 2 and a leaky Relu activation function. The output channels of these four convolutional layers are 16, 32, 32, and 32 respectively; the input of the convolutional neural network is the image in a harsh environment, and the output of the last fully connected layer is the predicted hyperparameters of various filters.
3. A six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, characterized in that, In Step 1, the enhancement of the image based on the deep learning method is realized based on the method of generative adversarial network. The generative adversarial network includes two parts: a generative network and a discriminative network; among them: The generation network model consists of 16 layers. The first half is composed of 6 convolutional layers and 2 pooling layers. After each convolutional layer, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 32, 32, 64, 64, 128, and 128. Pooling layers are added after the 3rd and 6th convolutional layers respectively. The second half is composed of 8 deconvolutional layers. The convolutional kernel size is 3×3, the stride is 2, and the number of channels is 256, 256, 128, 128, 64, 64, 32, and 3. Through convolutional and deconvolutional operations, the weight parameters are adjusted to achieve the effect of image enhancement. The discriminant network model consists of a fully convolutional network, including a total of 5 convolutional layers. After the first 4 convolutional layers, batch normalization and leaky Relu activation functions are added. The convolutional kernel size is 3×3, the stride is 1, and the number of channels is 42, 96, 192, 384, and 3. A sigmoid activation function is added at the end of the network for feature mapping, and the result is normalized.
4. The six - dimensional pose estimation method for an object applicable to harsh environments according to claim 3, wherein, The specific process of image enhancement based on the generative adversarial network is as follows: The image under harsh environmental conditions is input into the generation network. After the convolutional and deconvolutional operations of the generation network, an enhanced image is obtained. Then, the enhanced image and the image under normal conditions are input into the discriminant network for discrimination to distinguish between true and false, and a probability is output. When the output probability value is close to 1, it indicates that the input is an image under normal lighting conditions. When the discriminator cannot determine true or false, the image generated by the generation network at this time is the optimal image. Let {m i , i = 1, 2, ..., N} and {n i , i = 1, 2, ..., N} represent the images in harsh environments and the images under normal conditions respectively. The adversarial loss can be defined as: Among them, G represents the generation network, and D represents the discriminant network. The mean square error loss of the network model can be defined as: Finally, the adversarial loss and the mean square error loss are combined and configured with certain weights α and β to obtain the loss of the final generation network: L t = αL a + βL m (7) The loss of the discriminant network can be defined as:
5. The six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, characterized in that, In step 3, the Darknet53 network model is used for feature extraction. The input of the network is the image that has been enhanced by the filter, and the output is the feature of the image, which is used for subsequent semantic segmentation and key point prediction.
6. The six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, characterized in that, In step 3, semantic segmentation assigns a label to each pixel point superimposed on the image to distinguish different objects. More precisely, given N object classes, this is transformed into outputting a vector with a dimension of N + 1 at each spatial position, plus a dimension to represent the background. The loss function is: where M represents the number of categories; y c is an indicator variable, either 0 or 1, which is 1 if the category is the same as the sample's category, otherwise 0; p c represents the predicted probability that the observed sample belongs to category c.
7. A six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, characterized in that, In step 3, key point prediction uses the SIFT algorithm to detect distinctive two-dimensional key points in the texture image and lift them to three dimensions; then the FPS algorithm is applied to select the top N key points among them. In this way, the selected key points are not only evenly distributed on the object surface but also have distinct texture features and are easy to detect. During the key point prediction process, for each pixel point, the offset d i (x) of the pixel point relative to the two-dimensional key point of the object to which it belongs is predicted. Let the two-dimensional position of the pixel point be d and the true position of the two-dimensional key point be d i . Given P as the segmentation mask, the loss during the training process is as follows: At the same time, the confidence of each predicted point is also output. This confidence is obtained through the sigmoid function output by the network. For each three-dimensional key point, 20 two-dimensional positions with the highest confidence are selected as candidate points for subsequent pose calculation.
8. A six - dimensional pose estimation method for an object applicable to harsh environments according to claim 1, characterized in that In step 3, the regression pose is calculated based on the PnP algorithm of RANSAC to obtain the accurate six-dimensional pose of the object.
9. An object six-dimensional pose estimation system applicable to harsh environments, characterized in that: A method for six - dimensional pose estimation of an object applicable to harsh environments described in any one of claims 1 - 8, comprising a computer memory and a processor, an image enhancement module, an image fusion module, and a six - dimensional pose estimation module; the image enhancement module enhances pictures based on both traditional and deep - learning methods, the image fusion module fuses the enhanced pictures using an auto - encoder fusion network, the six - dimensional pose estimation module performs six - dimensional pose estimation of an object in a harsh environment through feature extraction, semantic segmentation, key - point prediction, and pose regression. Specific data processing and calculation work in all modules are completed by the computer processor, and all units have data interaction with the computer memory.
Citation Information
Patent Citations
Low-illumination image enhancement method based on Retinex and deep learning
CN111968044A
Pose estimation system and method based on improved YOLO6D algorithm
CN113436251A