Object Pose Estimation Method Based on Weighted Feature Fusion and Optical Flow Estimation Correction
Through the object pose estimation method modified by weighted feature fusion and optical flow estimation, the performance degradation of object pose estimation in real environment is solved, the robustness of the model and the accuracy of pose estimation are improved, and the memory usage is reduced.
Patent Information
- Application Number
- CN202211343439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-29
AI Technical Summary
The existing object posture estimation technology faces the influence of factors such as noise, occlusion, weak texture, light changes and messy backgrounds in real environments, resulting in a degradation of estimation performance.
An object pose estimation method based on weighted feature fusion and optical flow estimation correction is adopted, including data enhancement, weighted feature fusion, rotation and translation network module, optical flow estimation correction and loss function optimization, to improve the robustness of the model and pose estimation quality.
It improves the robustness of the model, reduces the sensitivity to images, enhances feature extraction capabilities, reduces the model memory usage, and improves the accuracy and efficiency of pose estimation.
Smart Images

Figure CN115620203B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pattern recognition and computer vision, and particularly to an object pose estimation method based on weighted feature fusion and optical flow estimation correction. Background Art
[0002] Deep learning plays an important research value in the academic field and also has many applications in industry. With the in-depth application of convolutional neural networks, deep learning has greatly promoted the development of object pose estimation technology.
[0003] 6D object pose estimation is a forward-looking technology in the field of computer vision and has great application potential in fields such as the metaverse, virtual reality, robot operation, and autonomous driving. For example: The autonomous driving system must identify the direction and distance of the vehicle ahead and timely adjust the direction and speed of its own vehicle to avoid traffic accidents; during the operation of a robot, if the position or pose of an object is inconsistent with the preset one, the manipulator is very likely to malfunction and there are safety hazards. These fields all require accurately and quickly obtaining the translation and rotation information of the target object in the three-dimensional rectangular coordinate system of the x-axis, y-axis, and z-axis from images. In recent years, both the academic community and the industrial community have been continuously exploring new object pose estimation technologies, publishing relevant papers in various computer vision conferences and setting new records on public datasets.
[0004] Although great progress has been made in pose estimation technology, there are still many challenges in real environments. External factors such as noise, occlusion, weak texture, illumination changes, and interference from cluttered background environments will all affect the performance of object pose estimation. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an object pose estimation method based on weighted feature fusion and optical flow estimation correction, which can effectively estimate the pose of an object in an RGB image.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: An object pose estimation method based on weighted feature fusion and optical flow estimation correction, comprising the following steps:
[0007] Step S1: Obtain two-dimensional RGB image information for data augmentation and load the pre-trained weights of the model;
[0008] Step S2: Use the yolov5 model for weighted feature fusion to extract the position features of the target object;
[0009] Step S3: Add a rotation and translation network module to regress and generate the information of six degrees of freedom of the three-dimensional translation and three-dimensional rotation of the object;
[0010] Step S4: Use the optical flow estimation method to correct the pose information generated in step S3 to further improve the quality of pose estimation and obtain the final result.
[0011] In a preferred embodiment, the step S1 specifically includes the following steps:
[0012] Step S11: Obtain a public object pose estimation training data set from the Internet and obtain relevant annotations of the training data;
[0013] Step S12: Random rotation and scaling are applied to the pose estimation training set images, and the pose truth is modified accordingly to match the enhanced image; when the two-dimensional image is rotated along the center point of the image by a rotation angle θ, the three-dimensional rotation R and three-dimensional translation t of the object must also be rotated around the z-axis by a rotation angle θ; this rotation around the z-axis is represented by an offset vector Δr, as shown below:
[0014]
[0015] Where T represents the vector transpose symbol, the rotation angle θ is a random value, which is subject to uniform sampling in the range [0°, 360°]; the rotation offset vector ΔR of the object is obtained from Δr, and the rotation matrix R of the object after random rotation enhancement is aug and the translation matrix t aug The calculation formula is as follows:
[0016] R aug =ΔR·R
[0017] t aug =ΔR·t
[0018] Need to adjust the object's three-dimensional translation t=(t x , t y , t z ) T t in z components; using scaling parameter f scale Scale the image, and the translation matrix t′ after scaling enhancement aug The calculation formula is as follows:
[0019]
[0020] Among them, the scaling parameter f scale It follows uniform sampling in the range [0.6, 1.4], t x ,t y and t z Respectively represent the components of the object's three-dimensional translation on the x-axis, y-axis, and z-axis;
[0021] Step S13: Enhance the color space of the dataset images and add Gaussian noise to the images;
[0022] Use an automatic data augmentation method to improve the model performance and enhance the generalization between the dataset and the model; The automatic data augmentation method is specifically as follows: Adjust the contrast and brightness of the input images, and use two parameters g and h to adjust the contrast and brightness, and enhance the color space of the images; where g represents the number of times the image enhancement transformation is applied to an image, and h represents the degree of each image enhancement transformation;
[0023] Add Gaussian noise enhancement to the automatic data augmentation method. The channel additive Gaussian noise is sampled from a normal distribution with a range of ; where the number of times g of the enhancement transformation for each image follows a uniform distribution of integers in the range [0, 4], and the degree h of each image enhancement transformation follows a uniform distribution of integers in the range [1, 15];
[0024] Step S14: Load the pre-trained weights of the yolov5 model on the COCO public dataset to improve the model training speed.
[0025] In a preferred embodiment, the step S2 specifically includes the following steps:
[0026] Step S21: Replace the feature fusion module in the yolov5 model with a BiFPN structure, perform bidirectional cross-scale connection and weighted feature fusion on the backbone network features of multiple scales to obtain cross-scale fused features, where the calculation formula for weighted feature fusion is as follows:
[0027]
[0028] O represents the result after weighted feature fusion, i represents the index of the i-th feature I i , j represents the index of the j-th weight w j , w i represents the weight corresponding to the feature I i , w j represents the j-th weight;
[0029] The weights can ensure that their values are greater than or equal to 0 through the ReLU activation function, and then are normalized to between 0 and 1. ε is a constant, and its value is set to 0.0001. This weighted fusion operation is faster than the softmax fusion method, and the accuracy is comparable to that of the softmax fusion method;
[0030] Step S22: The calculation formula for combining bidirectional cross-scale connection and weighted features is as follows:
[0031]
[0032]
[0033] Among them, represents the intermediate feature from the top-down of the P5 feature layer to the P4 feature layer output by the yolov5 backbone network, represents the output feature from the bottom-up of the P3 feature layer to the P4 feature layer output by the yolov5 backbone network. Resize represents the upsampling or downsampling operation, and Conv represents the convolution operation. represents the information of the P4 feature layer output by the yolov5 backbone network, and w1 represents the weight from to . represents the information of the P5 feature layer output by the yolov5 backbone network, and w2 represents the weight from to . represents the output feature after the P4 feature layer undergoes the cross-scale fusion operation, represents the output feature after the P3 feature layer undergoes the cross-scale fusion operation, and w′1 represents the weight from to . w′2 represents the weight from to . w′3 represents the weight from to .
[0034] In a preferred embodiment, the steps specifically include the following steps:
[0035] Step S31: Select the axis angle as the rotation representation, specifically manifested as a rotation subnet, which predicts a rotation vector for each target object instead of the conventional three-dimensional rotation R∈SO(3); where represents the set of real numbers, and SO(3) represents the three-dimensional special orthogonal group, which is a 3×3 matrix; the network structure of the rotation subnet is similar to the regression network in the object detection algorithm EfficientDet, and an iterative refinement module is added on the basis of the regression network to concatenate the rotation vector r init output by the regression network along the channel dimension and the output Δconvr of the last convolutional layer of the regression network; the iterative refinement module takes the output r init of the regression network as the input and regresses Δconvr, and the calculation formula of the obtained rotation vector is as follows:
[0036] r = r init +Δconvr
[0037] Step S32: The added iterative refinement module in the rotation module consists of a depthwise separable convolution, group normalization, and SiLU activation to form a separable convolutional layer; D iter represents the number of separable convolutional layers, which depends on the number of current feature layers D iter The calculation formula of is as follows:
[0038]
[0039] Among them, represents the floor operation;
[0040] The output r of the regression network in step S31 init , after passing through an iterative refinement module, the rotation vector r is obtained, which is set as the input of the next iterative refinement module, and the iterative refinement operation is repeated N iter times, and the calculation formula of N iter is as follows:
[0041]
[0042] Except for the output layer, the number of channels in all layers is the same as that of the classification network in the yolov5 network, and the output layer is determined by the number of anchor boxes and rotation parameters;
[0043] Step S33: Use group normalization instead of batch normalization, so that the model can train the rotation subnet from scratch with a batch size of 1, rather than the minimum batch size of 32 required for batch normalization; the number of times N group of performing group normalization is calculated as follows:
[0044]
[0045] Among them, W bifpn represents the number of channels of BiFPN and the yolov5 prediction network;
[0046] Step S34: The network structure of the translation subnet in the translation module is different from that of the rotation subnet described in step S32 in that the output translation vector is applicable to each prior box obtained by yolov5; instead of directly regressing all components of the translation vector t = (t x , t y , t z ), the task is instead to predict the pixel coordinates of the target object with the center point c = (c T , c x , c y ) T in the two-dimensional image and the component t z of the translation vector; using the coordinates of the target center point c and the component t of the translation vectorz and the intrinsic parameters of the camera, the t of the translation vector t x and t y The calculation formulas for the components are as follows:
[0047]
[0048]
[0049] where p x and p y respectively represent the center point p=(p x , p y ) T of the two-dimensional image in the coordinates of the x-axis and y-axis, f x and f y respectively represent the focal lengths of the camera in the x-axis and y-axis;
[0050] For each prior box, predict the offset in pixels from the center of the prior box to the center point c of the corresponding target object, that is, predict the offset from the current point in the given feature map to the center point; normalize the offset from each layer of the feature pyramid using the stride of the input feature map.
[0051] In a preferred embodiment, the specific method of step S4 is as follows:
[0052] Step S41: Use the object pose information regressed in step S3 to generate a rendered image F d Input the optical flow network and approach the ground truth pose image F im under the guidance of optical flow estimation. Calculate the warped feature map F through the warping function w , and this mapping process is denoted as flow d→im , as follows:
[0053] F w =w(Fd, flow d→im )
[0054] where the warping function is a bilinear function, and the warping in channel 1 is calculated as follows:
[0055]
[0056] where I is the bilinear interpolation kernel, q d =(q xd , q yd ) T is the two-dimensional coordinate after the generation result F r of step S3 generates the rendered image, q w =(q xw, q yw ) T is F r The structure F generated by optical flow estimation w of the two-dimensional coordinates, and δ is a learnable parameter;
[0057] The estimated optical flow flow r→im is cascaded with the feature map extracted from the ground truth pose F im to obtain the feature after the cascading operation by F w and to obtain the final pose estimation result through calculation;
[0058] Step S42: Use the loss function based on PoseLoss and ShapeMatch. For asymmetric objects, L asym The definition of the loss is as follows:
[0059]
[0060] where represents the true rotation matrix of the object obj calculated by the Rodriguez rotation formula from the true rotation vector , Rot(r, obj) represents the predicted rotation matrix of the object obj calculated by applying the Rodriguez rotation formula through the rotation vector r predicted by the network, represents the pose ground truth translation vector, t represents the translation vector predicted by the network, represents the three-dimensional model point set of the object, and m represents the number of three-dimensional points;
[0061] The loss function uses the pose ground truth and the estimated 6D pose to transform the target area, and then calculates the average point distance between the transformed model points; this method allows the model to be directly optimized according to the metrics of the measurement performance; when the rotation and translation losses are calculated independently of each other, no additional hyperparameters are required to balance the partial losses;
[0062] The loss function L sym The calculation formula is as follows:
[0063]
[0064] where represents the true rotation matrix of the target obj1 calculated by the Rodriguez rotation formula from the true rotation vector , Rot(r, obj2) represents the predicted rotation matrix of the target obj2 calculated by applying the Rodriguez rotation formula through the rotation vector r predicted by the network; L sym and L asymSimilar, instead of strictly calculating the distance between the matching points of two transformed point sets, it considers the minimum distance from each point to any point in the other transformed point set;
[0065] The complete loss function L trans is defined as follows:
[0066]
[0067] Compared with the prior art, the present invention has the following beneficial effects:
[0068] 1. It provides a directly applicable data augmentation method for object pose estimation methods, which can be applied to other pose estimation methods to improve the robustness of the model and reduce the sensitivity of the model to images.
[0069] 2. Through the weighted feature fusion operation, it can effectively improve the feature extraction ability of the yolov5 model and obtain a larger receptive field and clear target position.
[0070] 3. New rotation and translation modules are added to the yolov5 object detection algorithm structure, expanding it into an end-to-end direct regression network for pose estimation tasks, without the complex calculations of Perspective-n-Points in general pose estimation tasks. The batch quantization structure can effectively reduce the memory occupancy of the model.
[0071] 4. Using optical flow estimation to correct the model regression results, making full use of the rich information of the object's three-dimensional structure, and effectively improving the quality of pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a flowchart of the method implementation of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0074] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0075] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0076] As Figure 1 shown, the present invention provides an object pose estimation method based on weighted feature fusion and optical flow estimation correction, comprising the following steps:
[0077] Step S1: Obtain two-dimensional RGB image information for data augmentation and load the pre-trained weights of the model. Specifically, it includes the following steps:
[0078] Step S11: Obtain the publicly available object pose estimation training dataset from the network and obtain the relevant annotations of the training data:
[0079] Step S12: Randomly rotate and scale the images in the pose estimation training set, and make corresponding modifications to the pose ground truth to match the augmented images. When rotating a two-dimensional image along the center point of the image by a rotation angle θ, the three-dimensional rotation R and three-dimensional translation t of the object must also rotate around the z-axis by the rotation angle θ. This rotation around the z-axis can be represented by an offset vector Δr, as follows:
[0080]
[0081] where T represents the vector transpose symbol, and the rotation angle θ is a random value uniformly sampled from the range [0°, 360°). The rotation offset vector ΔR of the object is obtained from Δr, and the rotation matrix R aug and translation matrix t aug of the object after random rotation augmentation are calculated as follows:
[0082] R aug = ΔR·R
[0083] t aug = ΔR·t
[0084] To make image scaling a data augmentation method for pose estimation, it is also necessary to adjust the t x , t y , t z ) T component in t = (t z ). The image is scaled using the scaling parameter f scale , and the translation matrix t' aug of the object after scaling augmentation is calculated as follows:
[0085]
[0086] where the scaling parameter f scale is uniformly sampled from the range [0.6, 1.4], and t x , t y and t zrespectively represent the components of the three-dimensional translation of an object along the x-axis, y-axis, and z-axis.
[0087] Step S13: Enhance the color space of the dataset images and add Gaussian noise to the images;
[0088] Use an automatic data augmentation method to improve the model performance and enhance the generalization between the dataset and the model. The automatic data augmentation method consists of multiple augmentation methods, such as adjusting the contrast and brightness of the input images. Use two parameters g and h to adjust the contrast and brightness, and enhance the color space of the images. Among them, g represents the number of times the image augmentation transformation is applied to an image, and h represents the degree of each image augmentation transformation.
[0089] Regarding the problem that applying image rotation and cropping augmentation in the two-dimensional object detection task to the object pose estimation field will lead to a mismatch problem between the input image and the ground truth 6D pose, we remove these augmentations from the automatic data augmentation method and add Gaussian noise. The channel additive Gaussian noise is sampled from a normal distribution with a range of The number of augmentation transformations g for each image follows a uniform distribution of integers in the range [0, 4], and the degree h of each image augmentation transformation follows a uniform distribution of integers in the range [1, 15].
[0090] Step S14: Load the pre-trained weights of the yolov5 model on the COCO public dataset to improve the model training speed;
[0091] Step S2: Use the yolov5 model with weighted feature fusion to extract the position features of the target object. Specifically, it includes the following steps:
[0092] Step S21: Replace the feature fusion module in the yolov5 model with the BiFPN structure, perform bidirectional cross-scale connection and weighted feature fusion on the backbone network features of multiple scales to obtain cross-scale fused features. The calculation formula for weighted feature fusion is as follows:
[0093]
[0094] O represents the result after weighted feature fusion, i represents the index of the i-th feature I i , j represents the index of the j-th weight w j , w i represents the weight corresponding to the feature I i , w j represents the j-th weight.
[0095] The weights can ensure that their values are greater than or equal to 0 through the ReLU activation function, and then through Normalize to between 0 and 1. ε is a constant, and for the purpose of avoiding numerical instability, it is set to 0.0001. This weighted fusion operation is faster than the softmax fusion method and has comparable accuracy to the softmax fusion method.
[0096] Step S22: The calculation formula for combining bidirectional cross-scale connections and weighted features is as follows:
[0097]
[0098]
[0099] Among them, represents the intermediate feature from the P5 feature layer output by the yolov5 backbone network to the P4 feature layer from top to bottom, represents the output feature from the P3 feature layer output by the yolov5 backbone network to the P4 feature layer from bottom to top. Resize represents the upsampling or downsampling operation, and Conv represents the convolution operation. represents the information of the P4 feature layer output by the yolov5 backbone network, and w1 represents the weight from to represents the information of the P5 feature layer output by the yolov5 backbone network, and w2 represents the weight from to represents the output feature of the P4 feature layer after the cross-scale fusion operation, represents the output feature of the P3 feature layer after the cross-scale fusion operation. ′1 represents the weight from to w′2 represents the weight from to to
[0100] The BiFPN structure further considers the weight information and context information on the basis of the PANet structure to balance different scales, can obtain a larger receptive field, clear target position and rich semantic information, and is very important for the feature extraction of object postures.
[0101] Step S3: Add a rotation and translation network module to regress and generate the information of six degrees of freedom of the three-dimensional translation and three-dimensional rotation of the object. Specifically, it includes the following steps:
[0102] Step S31: Select the axis angle as the rotation representation because it requires fewer parameters than quaternions and is more convenient to calculate. The rotation module is specifically manifested as a rotation subnet that predicts a rotation vector for each target object Rather than the conventional three-dimensional rotation \(R\in SO(3)\). Among them, represents the set of real numbers, and \(SO(3)\) represents the three-dimensional special orthogonal group, which is a \(3\times3\) matrix. The structure of the rotation subnet is similar to the regression network in the object detection algorithm EfficientDet. An iterative refinement module is added on the basis of the regression network to concatenate the rotation vector \(r\) output by the regression network init and the output \(\Delta conv_r\) of the last convolutional layer of the regression network along the channel dimension. The iterative refinement module takes the output \(r\) of the regression network init as the input and regresses \(\Delta conv_r\). The calculation formula of the obtained rotation vector is as follows:
[0103] \(r = r\) init +\(\Delta conv_r\)
[0104] Step S32: The iterative refinement module added in the rotation module consists of depthwise separable convolution, group normalization, and SiLU activation to form a separable convolutional layer. \(D\) iter represents the number of separable convolutional layers, which depends on the number of layers of the current feature layer \(D\) iter The calculation formula is as follows:
[0105]
[0106] Among them, represents the floor operation.
[0107] The output \(r\) of the regression network in step S31 init , after passing through an iterative refinement module, obtains the rotation vector \(r\), which is set as the input of the next iterative refinement module, and the iterative refinement operation is repeated \(N\) iter times. The calculation formula of \(N\) iter is as follows:
[0108]
[0109] Except for the output layer, the number of channels in all layers is the same as that of the classification network in the yolov5 network. The output layer is determined by the number of anchor boxes and rotation parameters.
[0110] Step S33: Using group normalization instead of batch normalization can reduce the minimum batch size required during training, enabling the model to train the rotation subnet from scratch with a batch size of 1, while the minimum batch size required for batch normalization is 32. Using group normalization greatly reduces the amount of memory required by the model during training. The number of times \(N\) group of performing group normalization is calculated as follows:
[0111]
[0112] Among them, W bifpn Indicates the number of channels of BiFPN and yolov5 prediction networks.
[0113] Step S34: The translation subnet network structure of the translation module is basically the same as the rotation subnet described in S32, except that the output translation vector Applicable to each prior box obtained by yolov5. We do not directly regress the translation vector t = (t x , t y , t z ) T Instead, the task is to predict the center point c = (c x , c y ) T The pixel coordinates of the target object in the two-dimensional image and t z Component. Using the coordinates of the target center point c and the component t of the translation vector z And the intrinsic parameters of the camera, the translation vector t x and t y The calculation formula of the component is as follows:
[0114]
[0115]
[0116] Among them, p x and p y They represent the center point of the two-dimensional image p=(p x , p y ) T The coordinates on the x-axis and y-axis, f x and f y Represents the focal length of the camera on the x-axis and y-axis respectively.
[0117] For each prior box, we predict the offset in pixels from the center of the prior box to the center point c of the corresponding target object, that is, the offset from the current point in the given feature map to the center point. In order to maintain the relative spatial relationship, the offset is normalized from each layer of the feature pyramid using the stride of the input feature map.
[0118] Step S4: Use the optical flow estimation method to correct the pose information generated in step S3 to further improve the quality of pose estimation and obtain the final result. Specifically, the following steps are included:
[0119] Step S41: Generate a rendered image F using the object posture information obtained by regression in step S3 d Input the optical flow network, and under the guidance of the optical flow estimation, the pose truth image Fim Close, through the warping function Calculate the warping feature map f w , and this mapping process is denoted as flow d→im , as follows:
[0120]
[0121] Among them, the warping function is a bilinear function, and the warping in channel 1 is calculated as follows:
[0122]
[0123] Among them, I is the bilinear interpolation kernel, q d =(q xd , q yd ) T is the generated result F of step S3 r The two-dimensional coordinates after generating the rendered image, q w =(q xw , q yw ) T is the two-dimensional coordinates of the structure f generated by F r through optical flow estimation, and δ is a learnable parameter. w
[0124] Concatenate the estimated optical flow flow r→im with the feature map extracted from the ground truth pose F im to obtain the feature after the concatenation operation by F w and through calculation to obtain the final pose estimation result.
[0125] Step S42: Use a loss function based on PoseLoss and ShapeMatch, which considers not only three-dimensional rotation but also three-dimensional translation. For asymmetric objects, the definition of the L asym loss is as follows:
[0126]
[0127] Among them, represents the true rotation matrix of the object obj calculated by the true rotation vector through the Rodriguez rotation formula, Rot(r, obj) represents the predicted rotation matrix of the object obj calculated by the predicted rotation vector r of the network through applying the Rodriguez rotation formula, represents the pose ground truth translation vector, t represents the translation vector predicted by the network, Represents a 3D model point set of an object, and m represents the number of 3D points.
[0128] The loss function uses the ground truth pose and the estimated 6D pose to transform the target region, and then calculates the average point distance between the transformed model points. This method allows the model to be directly optimized according to the metrics of the measurement performance. When the rotation and translation losses are calculated independently of each other, no additional hyperparameters are required to balance the partial losses.
[0129] To handle symmetric objects, the corresponding loss function L sym The calculation formula is as follows:
[0130]
[0131] Where Indicates that the target obj1 is rotated by the ground truth rotation vector The ground truth rotation matrix calculated by the Rodriguez rotation formula, and Rot(r, obj2) represents the predicted rotation matrix calculated by applying the Rodriguez rotation formula to the rotation vector r predicted by the network for the target obj2. L sym Is similar to L asym Similar, but instead of strictly calculating the distance between the matching points of the two transformed point sets, it considers the minimum distance from each point to any point in the other transformed point set. This helps to avoid unnecessary computational effort when dealing with symmetric objects during training.
[0132] The complete loss function L trans Is defined as follows:
[0133]
[0134] The above are the preferred embodiments of the present invention. All changes made according to the technical solutions of the present invention, when the functional effects produced do not exceed the scope of the technical solutions of the present invention, fall within the protection scope of the present invention.
Claims
1. An object pose estimation method based on weighted feature fusion and optical flow estimation correction, characterized in that It includes the following steps: Step S1: Obtain two-dimensional RGB image information for data augmentation and load the pre-trained weights of the model; Step S2: Use the yolov5 model for weighted feature fusion to extract the position features of the target object; Step S3: Add a rotation and translation network module to regress and generate the information of six degrees of freedom of the three-dimensional translation and three-dimensional rotation of the object; Step S4: Use the optical flow estimation method to correct the pose information generated in Step S3 to further improve the quality of pose estimation and obtain the final result; The specific steps of Step S2 include the following steps: Step S21: Replace the feature fusion module in the yolov5 model with the BiFPN structure, perform bidirectional cross-scale connection and weighted feature fusion on the backbone network features of multiple scales to obtain cross-scale fusion features, and the calculation formula for weighted feature fusion is as follows: O represents the result after weighted feature fusion, and i represents the index of the i-th feature I i ; j represents the index of the j-th weight w j ; w i represents the feature I i corresponding weight, w j represents the j-th weight; The weights are ensured to be greater than or equal to 0 through the ReLU activation function, and then normalized to between 0 and 1 through with ε being a constant set to 0.0001; The calculation formula for combining bidirectional cross-scale connection and weighted features is as follows: Among them, represents the intermediate feature from the top-down of the P5 feature layer to the P4 feature layer output by the yolov5 backbone network, represents the output feature from the bottom-up of the P3 feature layer to the P4 feature layer output by the yolov5 backbone network. Resize represents the upsampling or downsampling operation, and Conv represents the convolution operation. represents the P4 feature layer information output by the yolov5 backbone network. w1 represents from to the weight, represents the P5 feature layer information output by the yolov5 backbone network. w2 represents from to the weight, represents the output feature after the P4 feature layer undergoes the cross-scale fusion operation, represents the output feature after the P3 feature layer undergoes the cross-scale fusion operation. w'1 represents from to the weight, w'2 represents from to the weight, w'3 represents from to the weight; The specific method of Step S4 is: Step S41: Generate a rendered image F using the object pose information obtained by regression in step S3 d Input the optical flow network and approach the ground truth pose image F under the guidance of optical flow estimation im by means of a warping function to calculate the warped feature map F w , and this mapping process is denoted as flow d→im , as follows: where the warping function is a bilinear function, and the warping in channel l is calculated as follows: where I is the bilinear interpolation kernel, and q d =(q xd , q yd ) T is the two-dimensional coordinate of the generated result F r after generating the rendering image, q w =(q xw , q yw ) T is the two-dimensional coordinate of the structure F r generated by optical flow estimation from F w , and δ is a learnable parameter; The estimated optical flow flow r→im is cascaded with the feature map extracted from the pose ground truth F im to obtain the feature after the cascading operation from F w and to obtain the final pose estimation result through calculation; Step S42: Use a loss function based on PoseLoss and ShapeMatch. For asymmetric objects, the L asym loss is defined as follows: Among them, indicates that the object obj is the true rotation matrix calculated by the true rotation vector through the Rodriguez rotation formula, and Rot(r, obj) represents the predicted rotation matrix calculated by applying the Rodriguez rotation formula to the rotation vector r predicted by the network for the object obj. represents the true pose translation vector, t represents the translation vector predicted by the network, M represents the set of three-dimensional model points of the object, and m represents the number of three-dimensional points. The loss function uses the pose ground truth and the estimated 6D pose to transform the target area, and then calculates the average point distance between the transformed model points; when the rotation and translation losses are calculated independently of each other, no additional hyperparameters are required to balance the partial losses; Loss function L sym The calculation formula is as follows: Among them, indicates that the target obj1 is the true rotation matrix calculated by the true rotation vector through the Rodriguez rotation formula, and Rot(r, obj2) represents the predicted rotation matrix obtained by applying the Rodriguez rotation formula to the rotation vector r predicted by the network for the target obj2; L sym is similar to L asym However, instead of strictly calculating the distance between the matching points of the two transformed point sets, it considers the minimum distance from each point to any point in the other transformed point set. The complete loss function L trans is defined as follows:
2. The object pose estimation method based on weighted feature fusion and optical flow estimation correction according to claim 1, wherein The specific steps of Step S1 include the following steps: Step S11: Obtain the publicly available object pose estimation training dataset from the network and obtain the relevant annotations of the training data; Step S12: Randomly rotate and scale the images in the pose estimation training set and modify the pose ground truth accordingly to match the augmented images; when rotating a two-dimensional image by an angle θ along the center point of the image, the three-dimensional rotation R and three-dimensional translation t of the object must also rotate around the z-axis by the angle θ; this rotation around the z-axis is represented by the offset vector Δr, as follows: Where, T represents the vector transpose symbol, the rotation angle θ is a random value that follows uniform sampling within the range of [0°, 360°]; the rotation offset vector ΔR of the object is obtained from Δr, and the rotation matrix R of the object after random rotation augmentation aug and the translation matrix t aug The calculation formula is as follows: R aug = ΔR·R t aug = ΔR·t It is necessary to adjust the three-dimensional translation t = (t x , t y , t z ) T in t z component; use the scaling parameter f scale to scale the image, and the calculation formula of the translation matrix t' aug after scaling enhancement is as follows: Among them, the scaling parameter f scale obeys uniform sampling within the range of [0.6, 1.4], and t x , t y and t z respectively represent the components of the three-dimensional translation of the object along the x-axis, y-axis, and z-axis; Step S13: Enhance the color space of the dataset images and add Gaussian noise to the images; Use the automatic data augmentation method to improve the model performance and enhance the generalization between the dataset and the model; the automatic data augmentation method is specifically: adjust the contrast and brightness of the input image, use the two parameters g and h to adjust the contrast and brightness, and enhance the color space of the image; where g represents the number of times the image augmentation transformation is applied to an image, and h represents the degree of each image augmentation transformation; Add Gaussian noise enhancement to the automatic data augmentation method. The channel additive Gaussian noise is sampled from a normal distribution with a range of ; where, the number of enhancement transformations g for each image follows a uniform distribution of integers in the range [0, 4], and the degree h of each image enhancement transformation follows a uniform distribution of integers in the range [1, 15]; Step S14: Load the pre-trained weights of the yolov5 model on the COCO public dataset to improve the model training speed.
3. The object pose estimation method based on weighted feature fusion and optical flow estimation correction according to claim 1, wherein The specific steps of Step S3 include the following steps: Step S31: Select the axis angle as the rotation representation, which is specifically manifested as a rotation subnet that predicts a rotation vector for each target object instead of the conventional three-dimensional rotation R ∈ SO(3); where represents the set of real numbers, and SO(3) represents the three-dimensional special orthogonal group, which is a 3×3 matrix; the network structure of the rotation subnet is similar to the regression network in the object detection algorithm EfficientDet, and an iterative refinement module is added on the basis of the regression network to cascade the rotation vector r init output by the regression network along the channel dimension and the output Δconvr of the last convolutional layer of the regression network; the iterative refinement module takes the output r init of the regression network as the input and regresses Δconvr, and the calculation formula of the obtained rotation vector is as follows: r = r init + Δconvr Step S32: The added iterative refinement module in the rotation module consists of depthwise separable convolution, group normalization, and SiLU activation to form a separable convolutional layer; D iter represents the number of separable convolutional layers, which depends on the number of layers of the current feature layer D iter The calculation formula of is as follows: Among them, represents the floor operation; The output r of the regression network in step S31 init , after passing through an iterative refinement module, a rotation vector r is obtained, which is set as the input of the next iterative refinement module, and the iterative refinement operation is repeated N iter times. The calculation formula of N iter is as follows: Except for the output layer, the number of channels of all layers is the same as that of the classification network in the yolov5 network, and the output layer is determined by the number of anchor boxes and rotation parameters; Step S33: Use group normalization instead of batch normalization so that the model can train the rotation subnet from scratch with a batch size of 1, rather than the minimum batch size of 32 required for batch normalization; the number of times N group for performing group normalization is calculated as follows: Among them, W bifpn represents the number of channels of the BiFPN and yolov5 prediction networks; Step S34: The difference between the translation subnet network structure of the translation module and the rotation subnet described in step S32 lies in the output translation vector which is applicable to each prior box obtained by yolov5; instead of directly regressing all components of the translation vector t=(t x ,t y ,t z ), the task is divided into predicting the pixel coordinates of the target object with the center point c=(c T x ,c y ) T in the two-dimensional image and the t z component; using the coordinates of the target center point c, the components t z of the translation vector, and the intrinsic parameters of the camera, the calculation formulas for the t x and t y components of the translation vector t are as follows: Among them, p x and p y respectively represent the center point p = (p x , p y ) T of the two-dimensional image on the x-axis and y-axis coordinates, f x and f y respectively represent the focal lengths of the camera on the x-axis and y-axis; For each prior box, predict the offset in pixels from the center of the prior box to the center point c of the corresponding target object, that is, predict the offset from the current point in the given feature map to the center point; normalize the offset from each layer of the feature pyramid using the stride of the input feature map.
Citation Information
Patent Citations
Indoor attitude estimation method based on thermodynamic map for single image
CN109063301A
Pose estimation method and system of target object and robot
CN113409384A