Method, device and product for 6d pose estimation of a target workpiece

By generating polarization anti-reflection and point cloud anti-occlusion feature vectors, and performing weighted fusion based on reflection intensity and occlusion rate, the problems of strong reflection and occlusion are solved, achieving high-precision and robust 6D pose estimation.

CN122636632APending Publication Date: 2026-08-25CHONGQING UNIV OF ARTS & SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611140333.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy and insufficient robustness in 6D pose estimation when faced with strong reflections and occlusions in industrial environments, and multimodal fusion methods have poor adaptability.

Method used

By generating polarization anti-reflective feature vectors and point cloud anti-occlusion feature vectors, and performing weighted fusion based on reflectivity coefficient and occlusion rate, the contribution of different modal features is dynamically adjusted, and pose estimation is performed by combining deep learning and geometric optimization algorithms.

Benefits of technology

It improves the accuracy and robustness of pose estimation in scenarios with both strong reflection and occlusion, and achieves adaptive fusion of multi-source information to adapt to complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636632A_ABST
    Figure CN122636632A_ABST
Patent Text Reader

Abstract

The application discloses a kind of methods, equipment and products for 6D pose estimation to target workpiece, belong to machine vision technical field.Method includes: based on the multi-angle polarized image of target workpiece, generates polarized anti-glare feature vector and glare intensity coefficient;Based on the three-dimensional point cloud of target workpiece, generates point cloud anti-shielding feature vector and shielding rate;According to the glare intensity coefficient and shielding rate of real-time calculation, dynamically weighted fusion is carried out to polarization feature and point cloud feature, and fusion feature vector is generated;Finally, based on the fusion feature vector, generate 6D pose initial solution, and obtain the final 6D pose estimation by iterative optimization.The application can resist the interference of strong glare and complex shielding simultaneously, realize high-precision, high-robustness 6D pose estimation, and be suitable for industrial automation assembly, robot precision grabbing and other fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine vision technology, and in particular to a method, apparatus and product for 6D pose estimation of a target workpiece. Background Technology

[0002] In the fields of intelligent manufacturing and industrial automation, 6D pose estimation, which is the accurate estimation of the position and orientation of a target workpiece in three-dimensional space, is a core technological prerequisite for robots to achieve tasks such as autonomous grasping, precision assembly, and intelligent inspection. However, the complexity of actual industrial environments poses significant challenges to 6D pose estimation.

[0003] On the one hand, many industrial workpieces (such as metal parts and high-gloss plastic parts) have smooth surfaces that are prone to strong specular reflection under certain lighting conditions, i.e., highlights or reflections. This causes traditional RGB image-based vision methods to fail because the image texture information of the reflective areas is severely damaged, making feature extraction and matching extremely difficult, thus making it impossible to accurately calculate the workpiece pose.

[0004] On the other hand, occlusion is a common problem in industrial settings, such as the stacking of workpieces, mutual occlusion of objects on conveyor belts, or occlusion of workpieces by the robot arm itself. Occlusion leads to incomplete 3D point cloud data, resulting in missing geometric information and structural voids. Point cloud registration algorithms that rely on complete geometric shapes will experience a sharp decline in performance under such circumstances, leading to large errors or even complete failure in pose estimation.

[0005] Currently, while existing multimodal fusion methods attempt to combine image and point cloud information, most employ fixed weights for feature stitching or simple linear fusion. This static fusion strategy cannot adaptively adjust the contribution of different modal features based on dynamic changes in the intensity of reflections and the degree of occlusion in the scene. Therefore, its robustness and accuracy remain limited when facing complex conditions where reflections and occlusions coexist. Thus, there is an urgent need for a high-precision, highly robust 6D pose estimation technique that can effectively handle both strong reflections and severe occlusion interference and adaptively fuse multi-source information. Summary of the Invention

[0006] The purpose of this application is to provide a method, device and product for 6D pose estimation of a target workpiece, aiming to solve the technical problems of weak anti-interference ability of single-modal methods and poor adaptability of multi-modal fusion methods in the prior art, resulting in low accuracy and insufficient robustness of 6D pose estimation in complex scenes with both reflection and occlusion.

[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for 6D pose estimation of a target workpiece, comprising: Based on the multi-angle polarization image of the target workpiece, a polarization anti-reflection feature vector and a reflectance intensity coefficient are generated; Based on the three-dimensional point cloud of the target workpiece, generate point cloud anti-occlusion feature vector and occlusion rate; The polarization anti-reflective feature vector and the point cloud anti-occlusion feature vector are weighted and fused according to the reflectivity coefficient and the occlusion rate to generate a fused feature vector; Based on the fused feature vector, a 6D pose estimate of the target workpiece is generated.

[0008] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for 6D pose estimation of a target workpiece.

[0009] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method for 6D pose estimation of a target workpiece.

[0010] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for 6D pose estimation of a target workpiece.

[0011] Based on the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, device, and product for 6D pose estimation of a target workpiece, achieving adaptive and intelligent fusion of different modal information. It not only fuses polarization images and 3D point clouds but also innovatively proposes dynamically adjusting the fusion weights of the two modal features based on real-time perceived reflectivity coefficients and occlusion rates. When reflectivity is severe, the system automatically increases the weight of polarization features; when occlusion is severe, it increases the weight of the completed point cloud features. This adaptive mechanism enables the invention to flexibly respond to various changing industrial scenarios, exhibiting stronger environmental adaptability and robustness compared to fixed-weight fusion methods. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1This is a flowchart illustrating a method for 6D pose estimation of a target workpiece in one embodiment of this application.

[0014] Figure 2 This is a schematic diagram illustrating the generation process of polarization anti-reflective feature vectors in one embodiment of this application.

[0015] Figure 3 This is a schematic diagram illustrating the process of generating point cloud anti-occlusion feature vectors in one embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0018] The method for 6D pose estimation of a target workpiece in this embodiment is executed by a computer device, which can be a terminal computing device or a server. The terminal computing device can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, etc. The server can be a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0019] The method for 6D pose estimation of a target workpiece provided in this application specifically includes: Step S110: Based on the multi-angle polarization image of the target workpiece, generate polarization anti-reflection feature vector and reflectance intensity coefficient.

[0020] Step S120: Based on the 3D point cloud of the target workpiece, generate the point cloud anti-occlusion feature vector and occlusion rate.

[0021] Step S130: The polarization anti-reflection feature vector and the point cloud anti-occlusion feature vector are weighted and fused according to the reflection intensity coefficient and the occlusion rate to generate a fused feature vector.

[0022] Step S140: Based on the fused feature vector, generate a 6D pose estimate of the target workpiece.

[0023] Specifically, this method can be divided into four interrelated stages. First, based on the multi-angle polarization image of the target workpiece, the method generates a polarization anti-reflective feature vector and a reflectivity coefficient; this step aims to effectively suppress high-gloss reflectivity interference on the workpiece surface using polarization information. Second, the method generates a point cloud anti-occlusion feature vector and an occlusion rate in parallel based on the 3D point cloud of the target workpiece; this step aims to overcome the problem of incomplete workpiece information caused by occlusion through the processing and completion of 3D geometric information. Next, using the outputs of the first two steps as input, the polarization anti-reflective feature vector and the point cloud anti-occlusion feature vector are weighted and fused according to the reflectivity coefficient and occlusion rate to generate a more comprehensive and interference-resistant fused feature vector. Finally, based on this fused feature vector, a solution model is used to generate a 6D pose estimate of the target workpiece. These four steps are interconnected and work synergistically to ensure the accuracy and robustness of the final pose estimate.

[0024] The core technical idea of ​​the above method lies in recognizing the advantages of polarized images in handling reflection problems and the advantages of 3D point clouds in describing geometric structures, and creatively combining the two. Specifically, this invention does not employ simple information superposition, but instead designs an adaptive fusion strategy. This strategy can perceive the severity of reflection and occlusion in the scene in real time and dynamically adjust the contribution weights of the two information sources in the final decision accordingly. By extracting polarization anti-reflection feature vectors that can resist strong reflection interference from multi-angle polarized images and point cloud anti-occlusion feature vectors that can resist the influence of occlusion from 3D point clouds, and then intelligently weighting them according to the quantified reflection intensity coefficient and occlusion rate, an accurate 6D pose estimate is finally calculated based on the high-quality fused feature vectors. This scheme fundamentally solves the problem of insufficient robustness of traditional single-modal methods or fixed-weight fusion methods in complex dynamic scenes.

[0025] In some embodiments, based on a multi-angle polarized image of the target workpiece, a polarization anti-reflection feature vector and a reflectance intensity coefficient are generated. Specifically, this includes: calculating a Stokes vector based on the multi-angle polarized image, and obtaining a polarization degree parameter map and a polarization angle parameter map from the Stokes vector; performing region separation of the multi-angle polarized image based on the polarization degree parameter map and the polarization angle parameter map, wherein the region separation includes separating reflective regions from diffuse reflection regions; generating a polarization anti-reflection feature vector based on the multi-angle polarized image after region separation; and calculating the reflectance intensity coefficient based on the proportion of reflective regions and the average polarization degree in the multi-angle polarized image.

[0026] Preferably, generating polarization anti-reflection feature vectors based on multi-angle polarized images after region separation specifically includes: using an improved ResNet18 network with a channel-spatial dual attention module added after the residual block to process the multi-angle polarized images after region separation to generate polarization anti-reflection feature vectors.

[0027] An improved ResNet18 network incorporates a channel-spatial dual attention module after the residual block in the standard ResNet18 architecture. The channel attention module analyzes the importance of different feature channels, automatically enhancing the weights of polarization feature channels that contribute significantly to distinguishing the identity and pose of target workpieces, while suppressing channels associated with reflective noise. The spatial attention module learns the importance of different spatial regions in the image, allowing the network to focus more on diffuse reflection regions rich in texture and structural information, while ignoring highlight reflection regions lacking information. By employing this improved ResNet18 network with a channel-spatial dual attention module after the residual block to process multi-angle polarized images after region separation, the anti-reflective and representational capabilities of the generated feature vectors are significantly improved.

[0028] The following is an example of generating polarization anti-reflective eigenvectors and reflectivity coefficients.

[0029] 1. Use a split-focus plane polarization camera (such as FLIR BFS-U3-51S5P-C) to acquire four-way polarization images of the target workpiece at 0°, 45°, 90°, and 135°, with the image resolution set to 1280×720.

[0030] 2. Preprocessing of the acquired polarization images: An adaptive Gaussian filtering algorithm is used for noise reduction. The filter kernel diameter is dynamically adjusted according to the noise intensity in different regions of the image. While removing salt-and-pepper noise and Gaussian noise, detailed information on polarization degree and polarization angle is preserved. Two-dimensional polarization correction is performed. A standard polarizer and an integrating sphere light source are used to complete polarization direction calibration and radiation calibration to eliminate system errors.

[0031] 3. Based on the preprocessed polarization image, calculate the Stokes vector (S0, S1, S2), where S0 is the total light intensity, S1 is the difference between the horizontal and vertical polarization intensities, and S2 is the difference between the 45° and 135° polarization intensities. Then, calculate the polarization degree parameter map and the polarization angle parameter map. The polarization degree ranges from [0,1], and the polarization angle ranges from [0,π].

[0032] 4. The Otsu adaptive thresholding algorithm is used to perform initial segmentation on the polarization degree parameter map and the polarization angle parameter map respectively. Combined with the region growing algorithm, the reflective seed points obtained from the initial segmentation are used as the core to gradually grow a complete reflective region. Then, through morphological dilation and erosion operations, small noise and holes in the reflective region are removed to obtain an accurate reflective region mask. The reflective region mask is then superimposed with the original polarization image to obtain a polarization image that separates the reflective region from the diffuse reflection region, laying the foundation for subsequent feature extraction.

[0033] 5. Input the polarized image, which separates the reflective and diffuse reflection regions, into the improved ResNet18 network described above. The improved ResNet18 network then outputs a polarization anti-reflective feature vector.

[0034] And, in one example, the reflectivity coefficient = (Area of ​​reflective region / Total area of ​​polarized image) × Mean polarization degree. The mean polarization degree is calculated based on the polarization degree parameter map, i.e., the parameter values ​​at each location point in the map are averaged.

[0035] In some embodiments, based on the 3D point cloud of the target workpiece, a point cloud anti-occlusion feature vector and an occlusion rate are generated. Specifically, this includes: registering the 3D point cloud of the target workpiece with a CAD model; jointly determining and locating occlusion regions in the 3D point cloud by comparing the geometric differences between the 3D point cloud and the CAD model, and combining the point cloud local density, normal vector deviation, and curvature abrupt change index; determining the surface type of the occlusion region based on the geometric prior information of the CAD model, and guiding the sampling process of the random sampling consensus algorithm based on the surface type to complete the 3D point cloud; generating a point cloud anti-occlusion feature vector based on the completed 3D point cloud; and determining the occlusion rate based on the proportion of points in the occlusion region in the completed 3D point cloud.

[0036] Preferably, the method of generating point cloud anti-occlusion feature vectors based on the completed 3D point cloud specifically includes: using an improved PointNet++ network with added local feature enhancement module and attention mechanism to process the completed 3D point cloud to generate point cloud anti-occlusion feature vectors.

[0037] The improved PointNet++ network adds a local feature enhancement module and an attention mechanism to the standard PointNet++ architecture. The local feature enhancement module typically employs a multi-scale grouping (MSG) strategy, which can simultaneously capture geometric details of the point cloud at different neighborhood scales. This is particularly important for accurately describing the completed region, as it better integrates the features of the original and completed point clouds. The attention mechanism calculates the association weights between points, allowing the network to focus on geometric structures (such as edges and corners) that play a crucial role in pose estimation, while suppressing the contributions of flat regions or noisy points. By using this improved PointNet++ network with added local feature enhancement and attention mechanisms to process the completed 3D point cloud, it is possible to extract feature vectors that are insensitive to occlusion and contain rich geometric details.

[0038] The following is an example of generating point cloud anti-occlusion feature vectors and occlusion rates.

[0039] 1. Three-dimensional point cloud data of the target workpiece is acquired using LiDAR, with a point cloud density of 10,000 points / frame. Outliers are removed using a statistical filtering algorithm, with the number of neighboring points set to 50 and the standard deviation threshold set to 1.5. Abnormal points that are too far from their neighbors are eliminated. Voxel mesh downsampling is used to normalize the number of point clouds to 2048 points, reducing computation while preserving the core geometric features of the point cloud. The normal vector and curvature are estimated using the PCA method, and the neighboring point set of each point is calculated to construct the covariance matrix. The normal vector and curvature are obtained through eigenvalue decomposition for subsequent occlusion area detection.

[0040] 2. The preprocessed 3D point cloud is precisely registered with the CAD model of the target workpiece. The ICP algorithm (Iterative Closest Point Algorithm) is used to align the 3D point cloud with the CAD model, and the distance between each point in the point cloud and the surface of the CAD model is calculated. Combining the local density, normal vector deviation, and curvature change characteristics of the point cloud, density threshold, normal vector deviation threshold, and curvature change threshold are set. When the local density of the point cloud is lower than the density threshold, the normal vector deviation is greater than the normal vector deviation threshold, or the curvature change is greater than the curvature change threshold, the area is determined to be an occluded area, and the location and range of the occluded area are accurately located.

[0041] 3. Based on the geometric prior of the CAD model, determine the surface type (plane, cylinder, sphere, etc.) of the occluded area and adaptively select the corresponding fitting model; use geometric prior to guide the sampling of the fitting model, restricting the sampling points within the geometric constraints of the CAD model, reducing invalid sampling and improving sampling efficiency; iteratively fit using the RANSAC algorithm, calculate the number of consistency points for each fitting model, select the fitting model with the most consistency points as the optimal fitting model, generate a complete point cloud of the occluded area based on the optimal fitting model, ensuring completion accuracy ≤2.5mm; fuse the completed point cloud with the original 3D point cloud to obtain a complete 3D point cloud.

[0042] 4. Input the completed 3D point cloud into the improved PointNet++ network described above, and the improved PointNet++ network will output the point cloud anti-occlusion feature vector.

[0043] And, in one example, occlusion rate =Number of point cloud points in the occluded area / Total number of point cloud points in the 3D point cloud.

[0044] In some embodiments, the polarization anti-reflection feature vector and the point cloud anti-occlusion feature vector are weighted and fused based on the reflectance intensity coefficient and the occlusion rate, specifically including: Step S131: Calculate the first weight and the second weight based on the reflectivity coefficient and the occlusion rate. ; ; in, and These are the first weight and the second weight, respectively. and These are the reflectivity and occlusion rate, respectively. A constant that is less than a threshold (e.g., 0.002); It is usually a very small value (e.g., 0.001), which is used to prevent the denominator from being zero in the ideal situation of no reflection and no obstruction.

[0045] Step S132: The polarization anti-reflection feature vector and the point cloud anti-occlusion feature vector are weighted and fused using the first weight and the second weight.

[0046] The formula for weighted fusion can be expressed as: ×Polarization Anti-reflective Feature Vector+ × Point cloud anti-occlusion feature vector.

[0047] The above fusion formula cleverly achieves dynamic weight allocation: when the reflective intensity When it increases, The proportion of polarization anti-reflective feature vectors increases accordingly, making them dominant in the fusion process; when the occlusion rate When it increases, The proportion of occlusion resistance increases accordingly, making the contribution of the point cloud anti-occlusion feature vector greater. This adaptive mechanism ensures that the fused features maximize effective information and suppress interference information in any scenario.

[0048] In some embodiments, generating a 6D pose estimate of a target workpiece based on a fused feature vector specifically includes: inputting the fused feature vector into a pose calculation model to calculate an initial 6D pose solution for the target workpiece; reprojecting the feature points of the CAD model of the target workpiece onto a multi-angle polarization image based on the initial 6D pose solution and calculating the reprojection error; and iteratively optimizing the initial 6D pose solution with minimizing the reprojection error as the optimization objective to output a target 6D pose solution.

[0049] Specifically, the process of generating 6D pose estimation for a target workpiece based on fused feature vectors includes two stages: initial solution generation and iterative optimization. First, the fused feature vectors generated after weighted fusion are input into a pre-trained pose calculation model (e.g., a lightweight fully connected network). This model can map high-dimensional features to pose parameters, thereby quickly calculating the initial 6D pose solution for the target workpiece. However, this initial solution may contain some errors. Therefore, to improve accuracy, a second stage of optimization is required. In this stage, several 3D feature points of the target workpiece's CAD model are used to calculate their 2D projection coordinates in a multi-angle polarization image based on the obtained initial 6D pose solution, using a camera projection model. The error between these projection coordinates and the coordinates of the corresponding feature points actually detected in the image is then calculated; this error is the reprojection error.

[0050] The subsequent iterative optimization process aims to minimize the reprojection error, iteratively optimizing the initial 6D pose solution to output the final target 6D pose solution. Specifically, this optimization process employs the Levenberg-Marquardt algorithm. The Levenberg-Marquardt algorithm is an efficient nonlinear least squares estimation algorithm that combines the advantages of gradient descent and Gauss-Newton methods. When solving pose optimization problems, it ensures both convergence stability and fast convergence speed. In each iteration, the algorithm calculates an optimal pose update based on the gradient information of the current reprojection error, used to correct the current pose estimate. This process is repeated until the reprojection error converges to below a sufficiently small threshold, or until the preset maximum number of iterations is reached. The output pose at this point is the finely optimized target 6D pose solution, with a much higher accuracy than the initial solution.

[0051] As can be seen from the above description, the present invention has the following technical effects: First, this invention improves the accuracy and robustness of 6D pose estimation in complex scenarios with both strong reflections and severe occlusion. By introducing polarization vision, information such as polarization degree is used to effectively separate and suppress strong reflection interference from the workpiece surface. Simultaneously, by combining prior CAD model data to complete the point cloud of incomplete areas, the geometric structure of occluded regions is restored. These two methods address the core challenges of reflection and occlusion from both image and point cloud perspectives.

[0052] Secondly, this invention achieves adaptive and intelligent fusion of information from different modalities. It not only fuses polarization images and 3D point clouds, but also innovatively proposes dynamically adjusting the fusion weights of the two modal features based on real-time perceived reflectivity coefficients and occlusion rates. When reflectivity is severe, the system automatically increases the weight of polarization features; when occlusion is severe, it increases the weight of the completed point cloud features. This adaptive mechanism enables the invention to flexibly respond to various changing industrial scenarios, exhibiting stronger environmental adaptability and robustness compared to fixed-weight fusion methods.

[0053] Finally, this invention balances accuracy and efficiency in pose estimation. It employs a two-stage strategy of "deep learning to generate an initial solution + geometric iterative optimization." The initial pose solution is rapidly regressed from powerful fusion features using a neural network, and then fine-tuned using the Levenberg-Marquardt algorithm to minimize the reprojection error. This method combines the global optimization capabilities of deep learning with the local convergence accuracy of traditional optimization algorithms, ensuring high accuracy while meeting the real-time requirements of industrial scenarios.

[0054] In an example workflow, the method for 6D pose estimation of a target workpiece specifically includes: First, during the data acquisition phase, a split-focus plane polarization camera (e.g., FLIR BFS-U3-51S5P-C) is used to acquire four-way polarization images of the target workpiece at a resolution of 1280×720, including 0°, 45°, 90°, and 135°. At the same time, a lidar is used to acquire three-dimensional point cloud data of the workpiece, with the point cloud density set to 10,000 points per frame.

[0055] In the polarization processing branch, such as Figure 2 As shown, the acquired polarization image is first preprocessed. An adaptive Gaussian filtering algorithm is used to dynamically adjust the filter kernel size according to the local noise intensity of the image to effectively remove noise and preserve details. Simultaneously, a standard polarizer and integrating sphere light source are used to perform polarization orientation calibration and radiometric calibration on the camera to eliminate systematic errors. Then, the Stokes vector and DoLP (degree of polarization) and AoP (angle of polarization) parameter maps are calculated. Next, the Otsu adaptive thresholding algorithm is used to perform preliminary segmentation of the DoLP and AoP parameter maps, and combined with region growing and morphological operations (dilation, erosion) to obtain an accurate reflective region mask. Finally, the preprocessed image is input into an improved ResNet18 network (with a channel-spatial dual attention module added after the residual block). After lightweight optimization, the network outputs a 256-dimensional normalized polarization anti-reflective feature vector.

[0056] In the point cloud processing branch, such as Figure 3As shown, the collected point cloud data first undergoes preprocessing. Statistical filtering is used to remove outliers, with the number of neighboring points set to 50 and the standard deviation threshold set to 1.5. Then, voxel mesh downsampling is used to normalize the point cloud to 2048 points. The normal vector and curvature of each point are estimated using the PCA method. Next, the ICP algorithm is used to precisely register the point cloud with the workpiece CAD model, and occlusion regions are located by combining the set density, normal vector deviation, and curvature abrupt change thresholds. For occlusion regions, the surface type is determined based on the geometric prior of the CAD model, and a geometric prior-guided RANSAC algorithm is used for point cloud completion, ensuring a completion accuracy error of less than 2.5 mm. Finally, the completed point cloud is input into an improved PointNet++ network (with the addition of a local feature enhancement module MSG and an attention mechanism), and the network outputs a 256-dimensional normalized point cloud anti-occlusion feature vector.

[0057] During the adaptive fusion phase, the reflectivity coefficient is calculated in real time. and occlusion rate Reflectivity coefficient = (Area of ​​reflective region / Total area of ​​polarized image) × Mean degree of polarization. Occlusion rate =Number of point cloud points in the occluded area / Total number of point cloud points in the 3D point cloud. Then, according to the formula... and Calculate the dynamic weights, where Set to 0.001. Dynamically weightedly fuse the 256-dimensional polarization features and the 256-dimensional point cloud features, and then perform L2 normalization to finally output a 512-dimensional fused feature vector.

[0058] In the pose calculation stage, a 512-dimensional fused feature vector is input into a lightweight network containing two fully connected layers to quickly generate an initial 6D pose solution (4-dimensional quaternions + 2-dimensional translation vector). Subsequently, with the goal of minimizing the reprojection error, the Levenberg-Marquardt algorithm is used to iteratively optimize the initial solution. The iteration process continues until the reprojection error is less than a threshold of 0.5 pixels, at which point the final high-precision 6D pose is output.

[0059] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0060] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0061] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0062] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0063] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0064] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0065] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0066] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for 6D pose estimation of a target workpiece, characterized in that, include: Based on the multi-angle polarization image of the target workpiece, a polarization anti-reflection feature vector and a reflectance intensity coefficient are generated; Based on the three-dimensional point cloud of the target workpiece, generate point cloud anti-occlusion feature vector and occlusion rate; The polarization anti-reflective feature vector and the point cloud anti-occlusion feature vector are weighted and fused according to the reflectivity coefficient and the occlusion rate to generate a fused feature vector; Based on the fused feature vector, a 6D pose estimate of the target workpiece is generated.

2. The method for 6D pose estimation of a target workpiece according to claim 1, characterized in that, Based on the multi-angle polarization image of the target workpiece, a polarization anti-reflective feature vector and a reflective intensity coefficient are generated, specifically including: The Stokes vector is calculated based on the multi-angle polarization image, and the polarization degree parameter map and polarization angle parameter map are obtained from the Stokes vector; The region separation of the multi-angle polarized image is achieved based on the polarization degree parameter map and the polarization angle parameter map, and the region separation includes separating the reflective region from the diffuse reflection region. The polarization anti-reflective feature vector is generated based on the multi-angle polarization image after region separation; The reflectivity coefficient is calculated based on the proportion of reflective areas and the average degree of polarization in the multi-angle polarized image.

3. The method for 6D pose estimation of a target workpiece according to claim 2, characterized in that, Based on the multi-angle polarization image after region separation, a polarization anti-reflective feature vector is generated, specifically including: An improved ResNet18 network with a channel-spatial dual attention module added after the residual block is used to process the multi-angle polarization image after region separation to generate the polarization anti-reflective feature vector.

4. The method for 6D pose estimation of a target workpiece according to claim 1, characterized in that, Based on the 3D point cloud of the target workpiece, an anti-occlusion feature vector and occlusion rate of the point cloud are generated, specifically including: The three-dimensional point cloud of the target workpiece is registered with the CAD model; By comparing the geometric differences between the 3D point cloud and the CAD model, and combining the local density of the point cloud, normal vector deviation, and curvature abrupt change, the occlusion area in the 3D point cloud is jointly determined and located. The surface type of the occluded region is determined based on the geometric prior information of the CAD model, and the sampling process of the random sampling consensus algorithm is guided based on the surface type to complete the three-dimensional point cloud; Based on the completed 3D point cloud, the point cloud anti-occlusion feature vector is generated. The occlusion rate is determined based on the proportion of point clouds in the occluded region in the completed 3D point cloud.

5. The method for 6D pose estimation of a target workpiece according to claim 4, characterized in that, Based on the completed 3D point cloud, the point cloud anti-occlusion feature vector is generated, specifically including: An improved PointNet++ network with added local feature enhancement module and attention mechanism is used to process the completed 3D point cloud to generate the point cloud anti-occlusion feature vector.

6. The method for 6D pose estimation of a target workpiece according to claim 1, characterized in that, The polarization anti-reflective feature vector and the point cloud anti-occlusion feature vector are weighted and fused based on the reflectivity coefficient and the occlusion rate, specifically including: Calculate the first weight and the second weight based on the reflectivity coefficient and the occlusion rate: ; ; in, and These are the first weight and the second weight, respectively. and These are the reflectivity and occlusion rate, respectively. It is a local minimum; The polarization anti-reflection feature vector and the point cloud anti-occlusion feature vector are weighted and fused using the first weight and the second weight.

7. The method for 6D pose estimation of a target workpiece according to claim 1, characterized in that, Based on the fused feature vector, a 6D pose estimate of the target workpiece is generated, specifically including: The fused feature vector is input into the pose calculation model to calculate the initial 6D pose of the target workpiece. The feature points of the CAD model of the target workpiece are reprojected into the multi-angle polarization image according to the initial 6D pose and the reprojection error is calculated. With minimizing the reprojection error as the optimization objective, the initial 6D pose solution is iteratively optimized to output the target 6D pose solution.

8. The method for 6D pose estimation of a target workpiece according to claim 7, characterized in that, The iterative optimization employs the Levenberg-Marquardt algorithm.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for 6D pose estimation of a target workpiece according to any one of claims 1-8.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for 6D pose estimation of a target workpiece as described in any one of claims 1-8.