Intelligent pose estimation method for maintainability test part based on pixel information
Through the improved intelligent pose estimation method of pixel information, combined with the network extraction features of activation function and attention mechanism, the problem of inaccurate pose information caused by occlusion is solved, and high-precision component pose estimation is achieved, which improves the effect of virtual and real fusion test.
Patent Information
- Application Number
- CN202510943679.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the maintenance test of the fusion of virtual reality and augmented reality, component occlusion leads to inaccurate acquisition of position information, which affects the test accuracy.
Using an intelligent pose estimation method based on pixel information, the feature extraction network is improved to enhance positioning accuracy in occlusion by combining the backbone network with activation function and channel attention mechanism, and using pixel key point voting and PnP algorithm to estimate the 6D pose of the component.
It improves the accuracy and robustness of component position estimation, especially in the case of occlusion, which can effectively predict the component position pose, and improves the accuracy and immersion of virtual and real fusion tests.
Smart Images

Figure CN120451274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual-reality fusion maintainability testing, and in particular to an intelligent pose estimation method for maintainability test components based on pixel information. Background Art
[0002] Maintainability refers to the ability to maintain or restore an equipment's original condition when repaired according to prescribed procedures and methods under specified conditions and within a specified timeframe. Maintainability assessments can verify the maintainability of weaponry. However, traditional maintainability assessments rely on multiple factors, including physical prototypes and test sites. While these assessments offer high reliability, they are also costly and require significant human and material resources. Consequently, new approaches to maintainability assessment are being sought. With the development of virtual reality technology, it has also been gradually applied to maintainability assessments. Using computer simulation and virtual reality technologies, virtual maintenance environments can be constructed using digital models of equipment prototypes, maintenance personnel, and maintenance tools. These virtual environments can then simulate maintenance processes for evaluation. While this approach is less expensive than traditional methods, it makes it difficult for maintenance personnel to perceive the forces and touch during assembly and disassembly, resulting in a poor user experience. This results in a lack of realism and a significant disconnect between the maintenance process and reality. In recent years, with the rapid development of technologies such as augmented reality and computer vision, maintainability testing and assessment based on the fusion of virtual and real worlds has become a new research hotspot. This testing model combines the advantages of traditional maintainability testing with virtual reality-based maintainability testing. By superimposing a virtual environment on a small number of physical objects, the system seamlessly integrates the physical and virtual objects, creating a cost-effective and immersive ship equipment maintenance scenario for maintenance personnel. Frequently operated equipment components are simulated primarily using physical devices, producing a completely realistic sense of touch and force. Furthermore, less frequently operated environmental equipment is simulated primarily using virtual simulation, with stereo registration technology superimposed on the surrounding equipment, effectively addressing space, resource, and cost constraints.
[0003] Maintainability testing is an important means of ensuring the integrity and continued performance of equipment. However, physical maintainability testing is subject to site constraints and is costly, while virtual maintainability testing often lacks force and tactile feedback, resulting in insufficient test accuracy. Therefore, virtual-reality integrated maintainability testing has gradually developed into a new, cost-effective testing strategy. This strategy involves superimposing a small number of key physical prototypes with an external virtual environment to create an immersive test scenario. This involves virtual environment registration, collision detection between virtual and real components, and maintenance path planning. All of these processes require accurate component pose information. However, due to the complexity of the maintainability test environment, component occlusion can lead to inaccurate component pose information. Therefore, a pixel-based intelligent pose estimation method for maintainability test components is proposed. Summary of the Invention
[0004] The present invention aims to provide an intelligent pose estimation method for maintainability test components based on pixel information, so as to solve the problem of inaccurate component pose information acquisition due to component occlusion caused by the complexity of the maintainability test environment.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] A method for intelligent pose estimation of a maintainability test component based on pixel information comprises the following steps:
[0007] S1, extract the features of the input image through the backbone network combining activation function and channel attention mechanism;
[0008] S2. The backbone network outputs the semantic segmentation result of the image, selects key points, and represents the vector field of all pixels in the image pointing to the key points;
[0009] S3, using the pixel key point voting method to determine the confidence representation of the two-dimensional key point;
[0010] S4. Estimate the 6D pose of the component by using the pose algorithm.
[0011] According to a further technical solution, in step S1, the backbone network is one of ResNet18, ResNet34, ResNet50 and ResNet101.
[0012] In a further technical solution, in step S1, the activation function is
[0013]
[0014] in Represents a simple and efficient two-dimensional space.
[0015] According to a further technical solution, in step S2, selecting key points includes the following steps:
[0016] S201, Collection Represents n points on the model surface, let set A represent the selected key points, set B represents the remaining key points, initially, set A is an empty set, set B is , select the initial key point as the point P0 with the largest distance between the point in set B and the center of mass of the point cloud, then A= ,B= ;
[0017] S202. Calculate the distances between other key points in set B and the initial key point, select the point with the largest distance to P0 as P1, add P1 to set A, and remove P1 from set B.
[0018] S203: Calculate the distance between set A and set B, select the point with the largest distance as the next key point, and repeat until k key points are selected.
[0019] A further technical solution is that in step S203, when k key points are selected, the minimum distance between set A and set B is calculated.
[0020]
[0021]
[0022] Among them, through represents the distance between a and b;
[0023] When k-1 key points are selected, since point b i In calculation When , the above formula can be simplified to
[0024]
[0025]
[0026]
[0027] Further technical solution, in step S3, each pixel P points to the key point X k Direction distance vector V k (p) is
[0028]
[0029] According to the results of semantic segmentation, find the relevant pixels of the component objects in the image, randomly select a pair of pixel points in the target object image, and divide their intersection points into two groups. As the key points of the object to be predicted , randomly select multiple pairs of pixels and repeat N times to obtain a set of key points with predicted objects , indicating the possible locations of key points, and finally all pixels of the component object vote for these hypothetical key points. Voting score Defined as
[0030]
[0031] in, Represents pixels Belong to the target , represents the exponential function, The condition is met when the threshold is 0.99.
[0032] A further technical solution is to use the PnP pose algorithm to estimate the 6D pose of the component in step S4, and to calculate the key points. Average value and covariance matrix Make an estimate
[0033]
[0034]
[0035] According to the mean of each key and covariance matrix , solve the rotation matrix of the target component through the minimum Mahalanobis distance and translation vectors
[0036]
[0037]
[0038] in, represents the coordinates of the three-dimensional key points, express The two-dimensional projection of the object is first rotated by the PnP algorithm based on the four key points. and translation vectors Initialize it to minimize the covariance matrix trajectory, then minimize the reprojection error through a nonlinear optimization algorithm, and finally solve the pose matrix.
[0039] The principles and beneficial effects of this technical solution are at least:
[0040] This paper improves the original PVNet algorithm, primarily by improving the feature extraction network. By introducing an attention mechanism and a FReLU activation function into the feature extraction network, the improved algorithm effectively enhances the network's ability to extract component features. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is the SE-PVNet backbone network structure diagram;
[0042] Figure 2 Schematic diagram of the insertion position of the FReLU activation function, where Figure 2 (a) Schematic diagram of the location of the FReLU activation function inserted in the jump connection with the same input and output dimensions. Figure 2 (b) Schematic diagram of the location of the FReLU activation function inserted into the skip connection with different input and output dimensions;
[0043] Figure 3 Schematic diagram of the difference between ReLU activation function and FReLU activation function, where Figure 3 (a) is a schematic diagram of the ReLU activation function structure. Figure 3 (b) is a schematic diagram of the FReLU activation function structure;
[0044] Figure 4 This is a comparison diagram of the ResNet and SE-ResNet structures, where Figure 4 (a) is a schematic diagram of the ResNet structure. Figure 4 (b) is a schematic diagram of the SE-ResNet structure;
[0045] Figure 5 A comparison chart of the experimental results of SE-PVNet and PVNet pose estimation loss functions;
[0046] Figure 6 Comparison chart of SE-PVNet and PVNet key point voting loss function experimental results;
[0047] Figure 7 A comparison chart of the experimental results of SE-PVNet and PVNet scene segmentation loss functions;
[0048] Figure 8 Schematic diagram of the comparison of distance and angle errors between SE-PVNet and PVNet;
[0049] Figure 9 Schematic diagram of the comparison of two-dimensional mapping indicators between SE-PVNet and PVNet;
[0050] Figure 10 This is a visualization diagram of the training process results, where Figure 10 (a) is a schematic diagram of the results of 5 rounds of iterative training. Figure 10 (b) is a schematic diagram of the results of 10 rounds of iterative training. Figure 10 (c) is a schematic diagram of the training results after 20 rounds of iteration. Figure 10 (d) is a schematic diagram of the results of 30 rounds of training. Figure 10 (e) is a schematic diagram of the results of 60 rounds of training. Figure 10 (f) is a schematic diagram of the training results after 120 rounds of iteration.
[0051] Figure 11 is the pose estimation result of the filter element under occlusion and truncation conditions, where Figure 11 (a) is a schematic diagram of unobstructed recognition. Figure 11 (b) is a schematic diagram of 10% occlusion recognition. Figure 11 (c) is a schematic diagram of 20% occlusion recognition. Figure 11 (d) is a schematic diagram of 40% occlusion recognition. Figure 11(e) is a schematic diagram of 50% occlusion recognition. Figure 11 (f) is a schematic diagram of 60% occlusion recognition;
[0052] Figure 12 is the pose estimation result of the wrench and socket, Figure 12 (a) is a schematic diagram of the wrench recognition results. Figure 12 (b) is a schematic diagram of the sleeve recognition results. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0054] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.
[0055] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0056] It should be emphasized here that the step marks mentioned below do not limit the order of the steps, but it should be understood that the steps can be executed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be executed simultaneously.
[0057] The core of the PVNet pose estimation method lies in predicting the direction vector from each pixel to a keypoint, and obtaining the coordinates of the keypoints by constructing a vector field voting method. This method is robust to occlusion and truncation. Even if some keypoints are invisible, the method leverages the rigid body properties to predict the pose of the occluded keypoints based on the vector field of the visible pixels, thereby estimating the object's pose.
[0058] Since maintainability test scenarios are generally complex and involve numerous surrounding components, occlusion is likely to occur during maintenance personnel's operation and between placed components. This paper introduces an attention mechanism based on PVNet to enhance the network's global modeling capabilities and improve the positioning accuracy of key points under occlusion. Furthermore, the FReLU activation function is introduced into the feature extraction network to enhance the network's feature extraction capabilities.
[0059] The details are as follows:
[0060] like Figure 1As shown in FIG, a component intelligent pose estimation method based on pixel information and PVNet includes the following steps:
[0061] S1, extract the features of the input image through the backbone network combining activation function and channel attention mechanism;
[0062] Residual networks can effectively solve the problems of gradient vanishing and gradient exploding caused by excessively deep network layers. This paper uses a residual network as the backbone network. Residual networks include one of ResNet18, ResNet34, ResNet50, and ResNet101. Because the more network layers there are, the longer the training time required. Considering time and hardware considerations, this paper uses ResNet18 as the backbone network for image feature extraction in this embodiment.
[0063] The ResNet18 network mainly consists of input, 4 convolutional layers, 8 residual blocks, average pooling layer, etc. Each convolutional layer uses a 3×3 convolution kernel and the commonly used ReLU activation function to extract the features of the target object in the image. Each residual block mainly consists of two convolutional layers and a skip connection, which is mainly used to solve the gradient disappearance and gradient explosion problems. The skip connection allows the input information to directly skip one or more layers and be used to add the output of the subsequent layers. There are two main ways of skip connection, such as Figure 2 The jump connection shown by the solid line deepens the network depth by concatenating the layers, while the input and output dimensions of the network remain unchanged. The jump connection shown by the dotted line is used to change the dimension of the output features. When the input and output dimensions are different, the dimensions are matched by using 1×1 convolution.
[0064] In addition, if Figure 2 As shown in the figure, the present invention also introduces the FReLU activation function for visual tasks in the ResNet18 network. It is a function specifically for computer vision tasks. By adding an additional spatial condition to extend the ReLU function, the increase in computational overhead is very small. It can activate insensitive information in the network space without changing the convolution and capture complex visual information.
[0065] The ReLU activation function is commonly used in existing technologies and has achieved good results in tasks such as image recognition, image segmentation, and target detection. However, it is not sensitive to certain information in space. The FReLU activation function is a function specifically designed for computer vision tasks and is an extension of the ReLU function. Its expression is:
[0066]
[0067] in Represents a simple and effective two-dimensional space, which can achieve simple and efficient spatial context feature extraction. The FReLU activation function can combine the pixel features in the residual block, which helps to obtain spatial information when applied in the convolutional network, enhance the sensitivity of the activation space, and thus improve the performance and feature expression ability of the model. Figure 3 As shown, Figure 3 (a) represents the ReLU activation function, Figure 3 (b) represents the FReLU activation function. From the figure, we can see the pixel space modeling ability of the FReLU activation function.
[0068] Compared to traditional convolutional neural networks, the residual structure in ResNet can build deeper networks and avoid the vanishing gradient phenomenon caused by the deepening of the network layers. Replacing the ReLU function in the residual network with the FReLU function makes the residual network more robust when facing visual tasks. However, the residual network extracts information at a single scale and cannot fully utilize the intermediate level features. Therefore, the present invention will introduce an attention mechanism module to improve the residual network, increase the network's information extraction capabilities, and make it more suitable for the task of estimating the pose of components under occlusion. By adding the channel attention mechanism SENet to the ResNet18 feature extraction network, the performance of the network can be improved by increasing the weight of unimportant information channels and reducing the weight of unimportant and erroneous information.
[0069] After adding the SENet attention mechanism module to the residual structure, the changes between the data are as follows Figure 4 Wherein, r is the dimensionality reduction ratio. When r is small, the global information transmitted by the previous layer can be preserved. In this embodiment, r is set to 16.
[0070] S2, the network outputs the semantic segmentation result of the input target component image and the vector field representation of all pixels pointing to key points in the image;
[0071] This paper uses the FPS algorithm to obtain the K key points of the object. The specific principles are as follows:
[0072] Assume that there are n points on the model surface, use Indicates that K points need to be selected, and the distance between each point and other points needs to be the farthest. Set A represents the selected points, and set B represents the remaining points.
[0073] S201. Initially, set A is empty, and set B = {b1b2...bn} is all the points on the surface of the object CAD model. To avoid inconsistent results caused by randomly selecting the initial point P0 from set B, the distance between the points in set B and the centroid of the point cloud is calculated, and the maximum value is selected as the initial point. At this time, A = ,B= .
[0074] S202. After the initial point is selected, the distance between the initial point and other points in set B is calculated to select the point with the largest distance from P0 as point P1, and P1 is added to set A and removed from set B.
[0075] S203. Calculate the distance between the two sets A and B, select the point with the largest distance as the next key point, and repeat the above process until K key points are selected. In this embodiment, K is set to 8. Although the above process can achieve the selection of key points, as the number of key points selected increases, the number of points in set A increases. When selecting the kth key point, it is necessary to calculate the distance between {a1a2...ak−1} and {b1b2...bn−k+1}. Each time a key point is selected, (n−k+1)(k−1) distance values need to be calculated, which will occupy a large amount of computing resources. Each time a point is selected, the distance between the points in set A and the points in set B must be calculated, which involves a large amount of repeated calculations. Therefore, the calculation of the repeated part needs to be optimized.
[0076] When selecting the kth key point, calculate the minimum distance between set A and set B
[0077]
[0078]
[0079] Among them, through Represents the distance between a and b.
[0080] When selecting k−1 key points, since point b i In calculation b has been calculated i and The minimum distance is obtained, and direct calculation will result in a lot of repeated calculations. Therefore, the above formula can be simplified to:
[0081]
[0082]
[0083]
[0084] After the above optimization, when selecting k key points, only the calculation The distance value of each point can be calculated, thereby reducing a lot of computing resources.
[0085] S3, using the RANSAC voting method to determine the confidence representation of the two-dimensional key points;
[0086] The present invention adopts the RANSAC voting algorithm to vote on the vector field and predict the direction of each pixel in the image, so that the network pays more attention to the local information of the image to reduce the impact of occlusion or truncation.
[0087] Depend on Figure 1 It can be seen that the network mainly outputs the semantic segmentation of the image and the vector field prediction information of the image pixels, with the pixel point P corresponding to the semantic segmentation label and each pixel pointing to the key point X k The distance vector V is represented by the direction vector Vk(p). k (p) is defined as follows
[0088]
[0089] Semantic segmentation divides the input image into two categories: foreground (target object) and background. Based on the voting algorithm, the key points of the target object in the image are selected. Based on the results of semantic segmentation, the relevant pixels of the component object in the image are found. A pair of pixels are randomly selected in the target object image and their intersection is selected. As the key points of the object to be predicted Randomly select multiple pairs of pixels and repeat N times to obtain a set of key points with predicted objects , indicating the possible locations of key points. Finally, all pixels of the component object vote for these hypothetical key points. Voting score The definition is as follows
[0090]
[0091] in, Represents pixels Belong to the target , represents the exponential function, The condition is met when the threshold is 0.99. The higher the voting score, the greater the possibility that the position is a key point. The confidence score of each hypothetical key point is obtained through the voting method, which provides a basis for solving the pose through the subsequent PnP algorithm.
[0092] S4. Estimate the 6D pose of the component by using the PnP algorithm.
[0093] In order to use the PnP algorithm, the key points need to be Average value and covariance matrix Make an estimate:
[0094]
[0095]
[0096] After obtaining the 2D key point positions of the parts, the PnP algorithm is used to solve the 6D pose of the parts. and covariance matrix , solve the rotation matrix of the target component through the minimum Mahalanobis distance and translation vectors
[0097]
[0098]
[0099] in, represents the coordinates of the three-dimensional key points, express The two-dimensional projection of the object is first rotated by the PnP algorithm based on the four key points. and translation vectors Initialize it to minimize the covariance matrix trajectory, then minimize the reprojection error through a nonlinear optimization algorithm, and finally solve the pose matrix.
[0100] The present invention takes the filter element as the object, performs training under the Linux system, and conducts experimental comparison and analysis on SE-PVNet and PVNet on the constructed filter element dataset.
[0101] The Adam optimizer is used and the model is trained for a total of 150 epochs. The initial learning rate is 0.001 and the learning rate decays by 50% every 20 epochs. The batch_size is set to 4 and the input image size is 640×480.
[0102] The pose estimation results are evaluated using the 2D Projection Metric and the Average 3D Distance of Model Vertices (ADD).
[0103] The present invention has constructed a total of 5850 filter cartridge images, of which 1350 are real images and 4500 are rendered images. Figures 5 to 9 As shown in the figure, the change curve of the loss function during training is as follows Figure 5 、 Figure 6 、 Figure 7 As shown, the ordinate represents the loss function value and the abscissa represents the number of iterations. Figure 8 、 9 This is the accuracy change curve. During the training process, the evaluation indicators are used to evaluate the training results every 5 rounds. The vertical axis represents the accuracy of the algorithm, and the horizontal axis represents the number of iterations.
[0104] Depend on Figures 5 to 9 The pose estimation loss function graph shows that the loss function value decreases rapidly at the beginning of training, gradually converges around 3×104 iterations, and then gradually decreases with increasing iterations, eventually converging to around 0.025. The convergence trends of the keypoint voting function and scene segmentation function are similar to those of the pose estimation loss function, with both functions converging around 3×104 iterations.
[0105] Select different iterations during the training process to visualize the pose estimation results, such as Figure 10 As shown in , the dark bounding box represents the prediction result, and the light bounding box is the true value. Figure 10 (d) It can be seen that after 30 iterations, the predicted bounding box has returned to the surrounding of the parts. After 120 training iterations, Figure 10 As shown in (f), the predicted bounding box is basically well overlapped with the real bounding box.
[0106] After training, the algorithm's performance was tested on the filter cartridge dataset using the aforementioned evaluation metrics. The test results are shown in Table 1. As can be seen from Table 1, under the same experimental conditions, the improved algorithm achieved a 1.54% improvement in the two-dimensional mapping metric accuracy of 91.84%, and a 1.91% improvement in the average distance metric accuracy of 83.69%.
[0107] Table 1 Evaluation indicators for filter cover pose estimation
[0108] method Two-dimensional mapping indicators Average distance indicator PVNet 90.30% 81.78% SE-PVNet 91.84% 83.69%
[0109] This paper also verifies the pose recognition effect of SE-PVNet algorithm in actual situations. Figure 11 As shown in the figure, it can be seen that the SE-PVNet algorithm can also estimate the pose of the filter element when the occlusion and cutoff area are less than 60%.
[0110] In addition, this paper also selects two parts, a wrench and a socket, to verify the generalization performance of the SE-PVNet algorithm. The pose estimation effect is as follows: Figure 12 shown.
[0111] The experimental results show that the improved algorithm is more accurate than the original PVNet algorithm on the dataset, and can also achieve good results in the case of occlusion.
[0112] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0113] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0114] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for intelligent pose estimation of maintainability test components based on pixel information, characterized in that: The steps include: S1, extract the features of the input image through the backbone network combining activation function and channel attention mechanism; S2. The backbone network outputs the semantic segmentation result of the image, selects key points, and represents the vector field of all pixels in the image pointing to the key points; S3, using the pixel key point voting method to determine the confidence representation of the two-dimensional key point; S4. Estimate the 6D pose of the component by using the pose algorithm.
2. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 1, characterized in that: In step S1, the backbone network is one of ResNet18, ResNet34, ResNet50 and ResNet101.
3. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 1, characterized in that: In step S1, the activation function is: ; in Represents a simple and efficient two-dimensional space.
4. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 1, characterized in that: In step S2, selecting key points includes the following steps: S201, Collection Represents n points on the model surface, let set A represent the selected key points, set B represents the remaining key points, initially, set A is an empty set, set B is , select the initial key point as the point P0 with the largest distance between the point in set B and the center of mass of the point cloud, then A= ,B= ; S202, calculate the distance between other key points in set B and the initial key point, select the point with the largest distance to P0 as P1, add P1 to set A, and remove P from set B. 1; S203: Calculate the distance between set A and set B, select the point with the largest distance as the next key point, and repeat until k key points are selected.
5. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 4, characterized in that: In step S203, when k key points are selected, the minimum distance between set A and set B is calculated. ; ; ; Among them, through represents the distance between a and b; When k-1 key points are selected, since point b i In calculation When , the above formula can be simplified to: ; ; 。 6. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 1, characterized in that: In step S3, each pixel P points to a key point X k Direction distance vector V k (p) is: ; According to the results of semantic segmentation, find the relevant pixels of the component objects in the image, randomly select a pair of pixel points in the target object image, and divide their intersection points into two groups. As the key points of the object to be predicted , randomly select multiple pairs of pixels and repeat N times to obtain a set of key points with predicted objects , indicating the possible locations of key points, and finally all pixels of the component object vote for these hypothetical key points. Voting score Defined as: ; in, Represents pixels Belong to the target , represents the exponential function, The condition is met when the threshold is 0.
99.
7. The method for intelligent pose estimation of a maintainability test component based on pixel information according to claim 1, characterized in that: In step S4, the 6D pose of the component is estimated using the PnP pose algorithm, and the key points Average value and the covariance matrix Make estimates; ; ; According to the mean of each key and the covariance matrix , solve the rotation matrix of the target component through the minimum Mahalanobis distance and translation vectors ; ; ; in, represents the coordinates of the three-dimensional key points, express The two-dimensional projection of the object is first rotated by the PnP algorithm based on the four key points. and translation vectors Initialize it to minimize the covariance matrix trajectory, then minimize the reprojection error through a nonlinear optimization algorithm, and finally solve the pose matrix.
Citation Information
Patent Citations
Workpiece pose estimation method based on dense prediction and grabbing system
CN115147488A
Single-view pose estimation method and system based on multi-modal input and attention mechanism
CN115861418A
Cited By
PVNet and double-branch feature fusion object pose detection method and device
CN120747226A