An industrial part pose estimation data acquisition platform and data processing method
By designing a data acquisition platform that includes an adjustable light source and an ArUco calibration board, and combining adaptive illumination adjustment and reprojection error filtering, the problems of high dataset construction complexity and error accumulation in existing technologies are solved, and efficient and accurate pose estimation dataset construction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU UNIV OF TECH
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-31
AI Technical Summary
Existing datasets suffer from high computational complexity, complex annotation processes, severe error accumulation, and low automation when constructing pose estimation datasets for weakly textured, highly reflective industrial parts, making it difficult to meet the needs of engineering applications.
A data acquisition platform consisting of a horizontal base, a rotating platform, an ArUco calibration board, a camera, and an adjustable light source is adopted. Combined with adaptive illumination adjustment and a data filtering method based on reprojection error, multi-view RGB-D data acquisition is achieved, and the pose is automatically generated by using reference coordinates provided by the ArUco calibration board.
It reduces the degree of human intervention, improves the efficiency of dataset construction and annotation accuracy, ensures the temporal consistency and spatial accuracy of pose estimation, and constructs a structured industrial part pose dataset.
Smart Images

Figure CN122492825A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an industrial part pose estimation data acquisition platform and data processing method. Background Technology
[0002] With the development of industrial automation and intelligent manufacturing technologies, object pose estimation has been widely used in robot grasping, automated assembly, industrial inspection, and flexible manufacturing. To improve the recognition accuracy and operational stability of pose estimation algorithms in complex industrial scenarios, it is usually necessary to construct a large-scale, accurately labeled, and representative industrial part pose dataset for model training, testing, and performance evaluation.
[0003] Existing pose estimation datasets include the LINEMOD dataset and the YCB-Video dataset. The LINEMOD dataset contains 15 low-texture household items, each corresponding to over 1100 frames of multi-view images, primarily used for 6D pose evaluation of targets with weak textures. The YCB-Video dataset provides accurate 6D pose annotations for 21 YCB objects across 92 videos and 133,827 frames, mainly targeting everyday items and general-purpose objects in desktop scenarios. These datasets have played a significant role in advancing research on general object pose estimation; however, their object types and acquisition scenarios are still primarily focused on home environments, which differ significantly from the material properties, scale distribution, and workspace environment of actual industrial parts.
[0004] To improve the representativeness of pose estimation evaluations in industrial scenarios, datasets targeting industrial objects have emerged. The T-LESS dataset contains 30 industrially relevant objects with no significant texture, emphasizing similarities between objects in color, reflectivity, and geometry. The MP6D dataset focuses on 6D pose estimation for metal parts, containing 20 metal parts and explicitly introducing occlusion, lighting variations, and multi-object scenes. Although these datasets have begun to expand into industrial scenarios, based on publicly available information, their primary goal remains serving as benchmark datasets for algorithm evaluation. They are still insufficient to directly meet the needs of engineering applications for rapid deployment of data acquisition, automatic pose generation, and structured annotation construction for specific small to medium-sized, weakly textured, and highly reflective industrial parts.
[0005] Compared to everyday consumer goods or ordinary targets with obvious textures, industrial parts, especially machined metal parts, typically exhibit a lack of texture information, smoother surfaces, a tendency to reflect light under certain lighting and viewing angles, numerous details in local edges and hole / groove structures, and a high degree of similarity in outline or size among different parts. These properties significantly increase the difficulty of pose estimation; highly glossy reflective surfaces introduce false edges in RGB images, leading to inaccurate depth measurements. For small to medium-sized industrial parts, these characteristics can easily cause local highlights, shadow occlusion, edge blurring, and holes, abrupt changes, and outliers in the depth map during actual data acquisition, thereby increasing the difficulty of feature extraction, point cloud construction, and pose determination.
[0006] Regarding data acquisition and pose annotation processes, existing public datasets and related research typically require manually assigning an initial pose to the first frame or keyframe during pose annotation, followed by frame-by-frame registration or subsequent optimization to obtain a relatively accurate pose label. Existing methods, to avoid manually annotating all video frames, only manually assign the target pose to the first frame of each video, further optimizing it with depth information; or they provide manually created CAD models and semi-automatic reconstruction models, and accurately annotate the 6D pose of test images from the system's sampling perspective. While these methods can achieve high-precision poses, they usually require significant manual intervention in the first frame and rely on subsequent frame-by-frame or viewpoint point cloud alignment and correction, making the system setup and annotation process quite complex.
[0007] Furthermore, existing local point cloud registration methods are highly sensitive to initial pose and easily get trapped in local optima. Therefore, related research introduces global search to reduce initialization dependence. Model-based registration methods show a significant decrease in accuracy and robustness under conditions of locally visible point clouds and large initial errors. For industrial parts with weak textures, smooth surfaces, and some reflectivity, the depth map itself is more susceptible to interference from local missing points, depth jumps, and outliers. Directly using frame-by-frame point cloud registration for pose annotation is not only computationally intensive but also more likely to introduce error accumulation in continuous frame processing, thus affecting the consistency and reliability of dataset labels. At the same time, existing dataset construction processes often separate data acquisition, pose solving, data filtering, annotation output, and dataset organization, resulting in low overall automation and hindering rapid reuse and batch construction in engineering scenarios.
[0008] Therefore, there is an urgent need to provide a pose estimation data acquisition platform and data processing method for weakly textured and highly reflective industrial parts. This would enable stable acquisition of multi-view RGB-D data of industrial parts, high-quality imaging under controlled lighting, automatic generation of the true pose of industrial parts based on reference coordinate constraints, and automated output of target region segmentation masks and pose labels, while simplifying the system structure and reducing deployment complexity. This would improve the efficiency of dataset construction, annotation accuracy, and temporal consistency, and provide a reliable data foundation for the training and application of pose estimation algorithms in robot grasping, automated assembly, and industrial inspection scenarios. Summary of the Invention
[0009] The present invention provides an industrial part pose estimation data acquisition platform and data processing method to solve the problems existing in the prior art.
[0010] The technical solutions adopted in this invention are as follows:
[0011] An industrial part pose estimation data acquisition platform, including
[0012] Horizontal base;
[0013] An ArUco calibration plate with feature points of known geometry, the ArUco calibration plate being used to support industrial parts;
[0014] A rotating platform, which is mounted on a horizontal base, is used to drive the ArUco calibration plate to rotate.
[0015] A camera, which is fixedly mounted on a horizontal base, is used to simultaneously acquire RGB images and depth images of industrial parts;
[0016] A first light source and a second light source are symmetrically arranged on both sides of the rotating platform. The position and height of the first light source and the second light source relative to the horizontal base are adjustable, and the incident angle of the first light source and the second light source is adjustable.
[0017] Furthermore, both the first light source and the second light source are LED fill lights, and both the first light source and the second light source are equipped with a diffuser or a diffuser structure at their light-emitting ends.
[0018] Furthermore, the first light source and the second light source are respectively mounted on a horizontal base via light source brackets. The light source brackets include a longitudinal lifting structure, a lateral telescopic structure, and an angle adjustment structure, which are used to adjust the height, horizontal distance, and incident angle of the corresponding light source, respectively.
[0019] Furthermore, the ArUco calibration plate is provided with a plurality of ArUco marks, which are arranged in a circular or arrayed manner, and the corner points of the ArUco marks constitute the feature points of the known geometric structure.
[0020] This invention also discloses a method for processing pose estimation data of industrial parts, implemented based on the aforementioned data acquisition platform, comprising the following steps:
[0021] S1: Fix the industrial parts onto the ArUco calibration plate, drive the ArUco calibration plate and the industrial parts to rotate synchronously through a rotating platform, and simultaneously acquire multi-view RGB images and depth images of the industrial parts through a camera.
[0022] S2: Detect feature points of the ArUco calibration board in each frame of the image, extract the two-dimensional pixel coordinates of the feature points, combine the known three-dimensional spatial coordinates of the feature points in the ArUco calibration board coordinate system, solve the pose transformation matrix of the ArUco calibration board in the camera coordinate system, and filter the data frames based on the reprojection error.
[0023] S3: In the initial frame, the depth image is enhanced, and the point cloud of the target region is extracted based on the target region constraint. The initial pose of the industrial part in the camera coordinate system is calculated by point cloud registration. A fixed spatial transformation relationship between the coordinate system of the industrial part and the coordinate system of the calibration board is established by combining the pose transformation matrix of the ArUco calibration board in the initial frame. In subsequent frames, the true pose of the industrial part in the camera coordinate system is calculated based on the pose transformation matrix of the ArUco calibration board in the current frame and the fixed spatial transformation relationship.
[0024] S4: Based on the real pose, project the 3D model of the industrial part onto the image plane to generate the target region segmentation mask and 2D bounding box. Associatively store the RGB image, depth image, segmentation mask, 2D bounding box and real pose to construct a structured industrial part pose dataset.
[0025] Further, in S1, the pixel grayscale values within the target area are statistically analyzed to construct a lighting evaluation function. Based on this lighting evaluation function, the light source brightness, incident angle, and light source distance are adjusted in a closed-loop iterative manner. The lighting evaluation function is:
[0026] ,
[0027] in, The grayscale mean is... The standard deviation of grayscale For the target grayscale value, These are the weighting coefficients;
[0028] A quadratic polynomial response surface model is constructed to relate the light source brightness, incident angle, and distance to the light source to the illumination evaluation function. Based on the response surface model, the optimal range of illumination parameters is determined.
[0029] Further, in S2, the three-dimensional feature points in the calibration plate coordinate system are projected onto the image plane through the pose transformation matrix to obtain the predicted pixel coordinates. The reprojection error between the predicted pixel coordinates and the actual detected pixel coordinates is calculated, and the overall pose solution result of the current frame is evaluated based on the root mean square error. When the root mean square error is greater than a preset threshold, the corresponding data frame is filtered out.
[0030] Furthermore, the reprojection error is:
[0031] ,
[0032] in, These are the actual corner pixel coordinates obtained from the detection. The predicted pixel coordinates of the feature points projected onto the image plane after pose transformation;
[0033] The root mean square error is:
[0034] ,
[0035] in, This represents the number of valid feature points used in the calculation.
[0036] Furthermore, in S3, the depth image quality enhancement processing includes:
[0037] Depth continuity constraint is used to remove pixels whose depth difference with neighboring pixels exceeds the depth continuity threshold; normal consistency constraint is used to remove pixels whose angle with the normal of neighboring points exceeds the normal angle threshold; and based on the target region mask, only the 3D points corresponding to the target region are back-projected to obtain the target region point cloud.
[0038] Furthermore, in S3, the fixed spatial transformation relationship is established by the following formula:
[0039]
[0040] in, This is a fixed spatial transformation matrix from the coordinate system of the industrial parts to the coordinate system of the calibration plate. Let be the pose transformation matrix of the calibration board in the camera coordinate system in the initial frame. This is the pose transformation matrix of the industrial parts in the camera coordinate system in the initial frame;
[0041] In subsequent frames, the true pose of the industrial part in the camera coordinate system is calculated using the following formula:
[0042] ,
[0043] in, This is the pose transformation matrix of the calibration board in the camera coordinate system in the current frame. This represents the true pose of the industrial parts in the current frame within the camera coordinate system.
[0044] Furthermore, in S4, the vertices of the 3D model of the industrial part are transformed to the camera coordinate system through the real pose, and projected onto the image plane in combination with the camera intrinsic parameters to obtain a 2D projection area. A target region segmentation mask is generated based on the 2D projection area, and a 2D bounding box is generated based on the minimum bounding rectangle of the projection point set.
[0045] The present invention has the following beneficial effects:
[0046] (1) By constructing an industrial part pose estimation data acquisition platform that includes a horizontal base, a rotating platform, an ArUco calibration plate, a camera, and a first and second light source that are symmetrically arranged on both sides and whose position, height, incident angle and illumination intensity are adjustable, stable acquisition of multi-view RGB-D data for weak texture and high reflectivity industrial parts is realized. The constructed dataset is consistent with the actual industrial scene in terms of object type, material properties and scale distribution, which improves the representativeness of the pose estimation algorithm in the industrial scene.
[0047] (2) By introducing an adaptive illumination adjustment strategy based on gray-scale statistical feedback, the brightness of the light source, the incident angle and the distance of the light source are adjusted in a closed loop. This effectively suppresses the interference of local highlights and shadows on the surface of highly reflective industrial parts, improves the consistency and uniformity of the gray-scale distribution of the image, and thus improves the stability of corner detection of ArUco calibration board and edge feature extraction of industrial parts.
[0048] (3) The ArUco calibration board provides feature points with known geometric structures as reference coordinates. Only the point cloud registration is needed in the initial frame to establish a fixed spatial transformation relationship between the industrial part coordinate system and the calibration board coordinate system. There is no need to manually annotate the initial pose of each frame of data, which reduces the degree of human intervention.
[0049] (4) In subsequent frames, the actual pose of the industrial parts is directly derived by matrix multiplication of the pose transformation matrix of the ArUco calibration board in the camera coordinate system in the current frame with the fixed spatial transformation relationship. This transforms frame-by-frame point cloud registration into temporal pose propagation based on the fixed spatial transformation relationship, which significantly reduces computational complexity and avoids error accumulation and registration drift caused by frame-by-frame registration, thereby improving the temporal consistency and spatial accuracy of pose annotation.
[0050] (5) In the initial frame pose solution process, depth continuity constraints and normal consistency constraints are introduced to enhance the quality of the depth image. The target area point cloud is extracted by combining the target area constraints. This effectively eliminates abnormal points caused by high reflectivity and measurement errors, and improves the geometric consistency and registration stability of the scene point cloud.
[0051] (6) An integrated processing flow was realized, from the acquisition of raw RGB-D data, adaptive lighting optimization, calibration board pose detection, reprojection error screening, initial frame point cloud quality enhancement, fixed space transformation propagation, automatic generation of real pose to automatic construction of target region segmentation mask and 2D bounding box. RGB image, depth image, target region segmentation mask, 2D bounding box and real pose are associated and stored according to a unified frame index, and a structured industrial parts pose dataset is constructed, which improves the data construction efficiency and the degree of system automation. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the overall data acquisition platform for industrial part pose estimation.
[0053] Figure 2 This is a structural diagram of the light source and its support.
[0054] Figure 3 This is a structural diagram of the camera and camera bracket.
[0055] Figure 4 This is a schematic diagram of the ArUco calibration plate.
[0056] Figure 5(a) shows the contour distribution of the evaluation function in the brightness and distance parameter space.
[0057] Figure 5(b) shows the corresponding three-dimensional response surface plot.
[0058] Figure 6 Examples of RGB and depth images of industrial parts.
[0059] Figure 7 This is a schematic diagram of the GT pose transformation.
[0060] Figure 8 Example of GT pose visualization results.
[0061] Figure 9 Example of a segmentation mask for industrial parts. Detailed Implementation
[0062] The invention will now be further described with reference to the accompanying drawings.
[0063] This invention provides a data acquisition platform and data processing method for industrial parts pose estimation, primarily targeting the dataset construction needs of weakly textured, highly reflective metal industrial parts. It establishes a complete processing workflow integrating multi-view image acquisition, adaptive lighting optimization, automatic target pose generation, and data annotation output. Through a unified hardware platform and a unified data processing link, the entire process from raw image acquisition and realistic pose generation to structured dataset output can be completed, thereby reducing manual intervention and improving data construction efficiency, annotation accuracy, and system stability.
[0064] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.
[0065] S1: Build a data acquisition platform.
[0066] Figure 1 This is a schematic diagram of the overall platform for industrial part pose estimation data acquisition. The platform of this invention generally includes a first light source support 1, a second light source support 7, a first light source 2, a second light source 6, a camera support 3, a camera 4, an ArUco calibration plate 5, a rotating platform 8, and a horizontal base 9.
[0067] Through the collaborative design of the above structures, multi-view image acquisition of industrial parts under controlled lighting conditions is realized, and a stable data foundation is provided for subsequent calibration plate pose detection, multi-coordinate system transformation calculation and automatic generation of target object pose.
[0068] To improve the imaging quality of reflective industrial parts during image acquisition, this invention incorporates an adjustable illumination module in the image acquisition platform to provide stable illumination for the target object (i.e., the industrial part) and the calibration board. It should be noted that, compared to natural objects with rich textures and predominantly diffuse reflection, industrial parts typically have smooth surfaces, limited texture information, and significant specular reflection characteristics. During image acquisition, this can easily lead to localized highlights and shadows, resulting in unstable grayscale responses at the calibration board corners and blurred edge information of the target object. Consequently, this affects the accuracy and stability of corner detection, feature extraction, and pose determination. Therefore, it is necessary to optimize the illumination conditions during the image acquisition stage to create a stable and uniform imaging environment.
[0069] The first light source support 1, the second light source support 7, and the first light source 2 and the second light source 6 together constitute an illumination adjustment system, set on both sides of the horizontal base 9, to create a stable and controllable lighting environment to meet the image acquisition needs of low-texture, high-reflectivity metal industrial parts. This invention primarily targets small to medium-sized industrial parts with a maximum feature size ranging from 10mm to 180mm. These parts typically have smooth surfaces, weak texture information, and curved or angular structures, making them prone to specular reflection and localized shadows during visual acquisition, thus affecting the stability of feature extraction and pose determination. Based on these problems, this invention achieves precise control of illumination conditions through the coordinated design of light source position, incident angle, and illumination parameters.
[0070] Figure 2The diagram shows the structure of the light source and its support. In this embodiment, each light source support includes a longitudinal lifting rod 101, a transverse telescopic rod 102, a light source fixing block 103, a connecting block 104, and a base plate 105. The longitudinal lifting rod 101 is used to adjust the height of the light source to ensure that the illumination area can cover the target object. When acquiring data for industrial parts with a maximum feature size of 10mm to 180mm, the overall size of the target part is relatively small, while the light source needs to be flexibly adjusted according to the camera installation position and changes in the target size. Therefore, the longitudinal adjustment stroke is preferably 300mm, allowing the light source to be adjusted to a suitable height within a reasonable range, thereby ensuring that the illumination area stably covers the target.
[0071] The horizontal telescopic rod 102 is used to adjust the relative distance between the light source and the industrial part, and its adjustment stroke is preferably 100mm. During the acquisition of images of industrial parts with weak texture and high reflectivity, changes in the distance between the light source and the target directly affect the illumination distribution and imaging quality. When the illumination conditions deviate from a reasonable range, it can easily lead to enhanced specular reflection or insufficient illumination, resulting in local saturation or loss of detail in the image, affecting the acquisition of the part's edge features.
[0072] The light source fixing block 103 is used to adjust the incident angle of the light source, and its pitch adjustment range is preferably 0°~90°. By changing the incident direction of the light, the reflected light can be deviated from the camera's line of sight, thereby reducing the interference of specular reflection on imaging. This adjustment method can be adapted to industrial parts with different geometries and surface characteristics.
[0073] Specifically, the first light source 2 and the second light source 6 are preferably LED supplementary lights, and a milky white diffuser is used at their light-emitting ends to diffuse the light. The diffuser is fixedly installed at the front end of the lamp body and covers the light-emitting area, so that the directional light emitted by the LED is diffused before shining on the target object, thereby reducing the directionality of the light and improving the uniformity of illumination. The principle is that when light directly shines on a relatively smooth industrial part, the reflected light energy tends to concentrate and distribute along a specific direction; when this direction is close to the camera's line of sight, local highlight areas are easily formed in the image, making the gray value close to saturation, thus causing the loss of detail information. By using the diffuser, the influence of reflection on imaging can be reduced, and the consistency of gray-scale response in the target area can be improved.
[0074] The camera mount 3 and the camera 4 together constitute an image acquisition system, which is fixedly mounted on the horizontal base 9. It is used to achieve multi-view, highly stable image acquisition of industrial parts under single-camera conditions. Through multi-degree-of-freedom adjustment of the camera's spatial position and attitude, clear and complete images of the target object can be obtained from different observation angles, thereby meeting the requirements of subsequent pose determination.
[0075] Figure 3The diagram shows the structure of the camera and camera mount. The camera mount includes a column 301, a column connecting block 302, a horizontal rod 306, a sliding block 303, and a camera mounting base 304. The column connecting block 302 is connected to the column 301 and uses fixing screws to adjust the camera height, allowing the camera's main optical axis to be aligned with the target area. The horizontal rod 306 is connected to the column connecting block 302, and its horizontal rotation range is preferably 120°. This angle range can cover the main lateral observation directions of industrial parts, meeting the needs of multi-view data acquisition, while reducing the complexity of the mechanism while ensuring structural stability.
[0076] The sliding block 303 is fitted onto the horizontal rod 306 and can move along the horizontal rod. Its preferred range of movement is 300mm, used to adjust the relative distance between the camera and the industrial part. This parameter setting comprehensively considers the camera's working distance and the target size range. When the maximum feature size of the target part is 10mm to 180mm, position adjustment within this range allows the target object to fully enter the camera's field of view and maintain high spatial resolution in the image, thereby improving the reliability of feature extraction and pose determination. The sliding block 303 can also be adjusted at its angle to the horizontal direction to further change the camera's attitude to adapt to the observation needs of parts with different structures.
[0077] The camera mount 304 is connected to the sliding block 303, and its pitch angle adjustment range is preferably 90°, which is used to change the camera's viewing angle of the industrial parts from above, at eye level, or tilted, thereby acquiring image information from different perspectives. The camera 305 is fixed to the camera mount 304 with screws.
[0078] Camera 4 is preferably a depth camera, used to simultaneously acquire RGB images and depth images. In this embodiment, the Orbbec Gemini215 depth camera is preferably used, which can acquire stable depth information within a range of 0.1m to 1m. In the application scenario of small and medium-sized industrial parts targeted by this invention, the preferred working distance is 0.2m to 0.6m. Within this distance range, the target object can fully enter the field of view while maintaining high depth resolution and image clarity, thereby providing a reliable data foundation for feature point detection, point cloud construction, and pose determination on the calibration board.
[0079] Through the coordinated design of the above structure and parameters, the camera has the ability to adjust its spatial position and attitude with multiple degrees of freedom, thereby enabling the selection of a suitable observation angle according to the size, shape and surface characteristics of industrial parts, and achieving stable and high-quality multi-view image acquisition.
[0080] Figure 4This is a schematic diagram of the ArUco calibration plate. The ArUco calibration plate 5 is mounted on the rotating platform 8, and preferably adopts a structure in which multiple ArUco markers are distributed around it. In this embodiment, the calibration plate measures 430mm × 430mm × 5mm, with a total of 20 ArUco markers arranged on it. The side length of each marker is 64.5mm, and the interval between adjacent markers is 6.38mm. These dimensions and layout parameters are designed based on the overall size of the calibration plate and the camera's field of view, ensuring that the markers can be stably detected under different observation angles and distances.
[0081] ArUco markers are used to provide feature points with known geometric structures. By detecting the marked corner points in the image and establishing their correspondence with three-dimensional spatial coordinates, the pose of the calibration board in the camera coordinate system can be solved. Based on this pose information, it can be further used as a reference coordinate to calculate and annotate the pose of the target object.
[0082] Compared to single-marker detection methods, this invention employs a multi-marker joint detection strategy. When some marks cannot be effectively identified due to occlusion or changes in viewing angle, the pose can still be solved using the remaining visible marks, thus avoiding the problem of overall calculation failure due to single-point failure. By introducing redundant mark information, not only is the stability of calibration board pose estimation improved, but the robustness of the system under complex lighting and multi-view conditions is also enhanced.
[0083] The rotating platform 8 is placed on a horizontal base 9 and its position can be adjusted on the base. It is used to support the calibration plate and industrial parts and drive them to rotate stably around the central axis. The rotational speed of the rotating platform 8 is preferably set to 3 r / min. This parameter is determined based on the camera's frame rate, target size, and image sharpness requirements.
[0084] In this embodiment, the camera frame rate is 30fps. If the rotation speed is too high, the pose changes between adjacent frames will be too large, easily leading to motion blur and reducing the accuracy of corner detection and pose calculation. If the rotation speed is too low, the viewpoint changes that can be acquired per unit time will be less, affecting the acquisition efficiency. 3r / min is suitable for small to medium-sized, lightweight industrial parts, ensuring the continuity of pose changes between adjacent frames while balancing image clarity and acquisition efficiency, thus meeting the requirements of subsequent pose calculation and data annotation.
[0085] S2: Adaptive illumination adjustment and image acquisition.
[0086] Based on the above platform structure, image data acquisition is performed, and the specific process is as follows:
[0087] In this embodiment, a low-texture, highly reflective metallic industrial part with dimensions of 98mm × 51mm × 21mm is selected as the example object. First, the target industrial part is placed in the central area of the ArUco calibration plate 5, keeping it relatively fixed with the calibration plate throughout the acquisition process. Then, the rotating platform 8 is activated, driving the calibration plate and the industrial part to rotate around the central axis at a constant angular velocity, thereby forming a continuous and uniformly changing posture sequence of the target object from the camera's perspective, thus constructing multi-view observation conditions.
[0088] Before acquisition, the spatial position, illumination distance, and incident angle of the light source are adjusted using a light source bracket, and an adaptive adjustment mechanism based on grayscale statistical feedback is used to ensure stable grayscale response in the calibration plate area and the target edge area. Simultaneously, the camera height, horizontal position, and attitude are adjusted using a camera bracket to ensure the target is fully within the field of view, and the camera position remains fixed during the acquisition process.
[0089] In this embodiment, the resolution of the RGB and depth images is set to 1080×720, the acquisition frame rate is set to 30fps, and it is matched with the rotation speed of the rotating platform at 3r / min. This parameter combination ensures smooth and continuous pose changes between adjacent frames and avoids motion blur. After the rotating platform completes approximately two rotations, acquisition stops, yielding approximately 1500 to 2000 frames of RGB images and corresponding depth images.
[0090] To further enhance the diversity of perspectives, the position of the camera bracket can be changed, the incident direction of the light source can be adjusted, or the position of the rotating platform can be moved to repeat the above acquisition process, thereby obtaining multiple sets of data samples under different observation conditions.
[0091] During data acquisition, the depth camera simultaneously outputs RGB and depth images, and assigns a unique timestamp to each frame. After acquisition, the RGB image sequence, depth image sequence, and timestamp information are stored in a structured manner, and a frame-by-frame correspondence is established so that subsequent pose calculation results can be matched with the original image data.
[0092] To further ensure that the acquired images have stable brightness levels and good illumination uniformity, this invention designs an illumination quality evaluation, parameter optimization, and adaptive adjustment method based on grayscale statistics.
[0093] Based on the coordinated adjustment of light source brightness, light source distance, and incident angle, a lighting quality evaluation method based on ROI grayscale statistics is introduced. Specifically:
[0094] By constructing grayscale mean With gray standard deviation The system quantitatively characterizes the overall brightness level and illumination uniformity of the image; furthermore, by setting brightness constraint ranges and uniformity judgment thresholds, it determines whether the current illumination conditions meet the acquisition requirements.
[0095] Based on this, an illumination evaluation function S, including a brightness deviation term and a grayscale discrete term, is constructed. Combined with a quadratic polynomial response surface model, the relationship between light source brightness, incident angle, light source distance, and the evaluation function S is fitted and analyzed to obtain an optimal range of illumination parameters. Furthermore, based on grayscale statistical indicators and the evaluation function S, this invention constructs a closed-loop feedback adaptive illumination adjustment mechanism. By iteratively updating the light source parameters, the illumination conditions gradually converge to a more optimal state, thereby improving image acquisition quality, corner detection stability, and subsequent pose calculation accuracy.
[0096] To quantitatively describe illumination quality, this invention statistically analyzes the pixel grayscale values within the target region (ROI). Let the set of pixels corresponding to the target ROI in the image plane be... Then, the total number of pixels N participating in the statistics is defined as the cardinality of the pixel set, that is: For sets The pixels in the image are numbered in any order. Let the grayscale value corresponding to the i-th pixel be... Then the image grayscale mean Defined as:
[0097] ,
[0098] in, Used to characterize the overall brightness level of the sampling area, when A low light level indicates insufficient light; when... If the brightness is too high, it indicates a risk of overall overexposure or excessively strong highlights in certain areas.
[0099] To further characterize illumination uniformity, grayscale standard deviation is introduced. As an evaluation indicator, its calculation formula is as follows:
[0100]
[0101] in, This indicates the degree of dispersion of grayscale values relative to the mean. When the value is small, it indicates that the grayscale distribution at each sampling point is relatively concentrated and the illumination is relatively uniform; when... A larger value indicates a significant difference in brightness within the area, typically corresponding to a situation where highlights and shadows coexist.
[0102] Based on the above grayscale statistics, the present invention sets the illumination uniformity determination criteria as follows:
[0103] , ,
[0104] in, and These represent the allowable range of grayscale mean, used to constrain the overall brightness of the image to be within a reasonable range. This is the grayscale standard deviation threshold, used to determine the uniformity of illumination. When the above conditions are met simultaneously, the current illumination environment is considered stable and uniform, sufficient for subsequent visual processing.
[0105] It should be noted that the parameters are not fixed values, but should be adaptively set according to the surface material, reflective properties, and acquisition environment of the industrial parts. In practice, ROI images are acquired under different combinations of lighting parameters during the pre-acquisition stage, and their grayscale distribution is statistically analyzed.
[0106] Table 1. Experimental Results of Gray Scale Standard Deviation and Illumination Quality
[0107]
[0108] As can be seen from Table 1, when the grayscale standard deviation When the value is in the range of approximately 20 to 22, the image grayscale distribution is relatively concentrated, and the illumination uniformity is good; when... When the threshold value exceeds approximately 25, noticeable highlight or shadow areas begin to appear in the image, and the uneven lighting becomes significantly more pronounced. Based on the above experimental statistical results, the optimal grayscale standard deviation threshold is determined. The value is set to be within the range of 22 to 25, which is used as the basis for judging the uniformity of illumination.
[0109] Meanwhile, the allowable range of grayscale mean [ , The optimal setting is 30% to 70% of the imaging dynamic range. Here, the imaging dynamic range refers to the range of grayscale values in the image. The 30% to 70% range can effectively avoid images that are too dark or too exposed, thereby ensuring the stable extraction of effective feature information.
[0110] Based on this, in order to comprehensively evaluate the lighting quality, this invention constructs a lighting evaluation function:
[0111] ,
[0112] in, The target grayscale value is preferably the median of the grayscale dynamic range, which is used to characterize the desired overall brightness level. , where is the weighting coefficient, used to adjust the relative influence of the brightness deviation term and the grayscale standard deviation term in the evaluation function. In practical applications, to ensure that the brightness deviation term and the grayscale standard deviation term are of similar magnitude, thereby achieving a comprehensive evaluation of illumination uniformity and overall brightness, it is preferable to . The value is set to be within the range of 0.5 to 1.0. In this embodiment, α≈0.8 is chosen to balance the reasonableness of brightness and the uniformity of illumination.
[0113] Based on the aforementioned illumination evaluation function, different combinations of illumination parameters can be uniformly and quantitatively evaluated, and the optimal selection of illumination parameters can be achieved by comparing the values of the evaluation function S. When the evaluation function S obtains a smaller value, it indicates that the overall brightness of the image under the current illumination conditions is close to the target gray level and the gray level distribution is uniform, thus exhibiting superior imaging quality.
[0114] To describe the relationship between illumination parameters and image quality, this invention constructs a quadratic polynomial response surface model to fit and analyze the relationship between light source brightness, incident angle, light source distance, and evaluation function S. Its general form can be expressed as a quadratic polynomial function:
[0115] ,
[0116] in, , and These correspond to the brightness, incident angle, and distance in Table 1, respectively. The brightness setting value output by the light source controller is expressed as a percentage relative to the maximum output brightness of the light source; The angle between the main light-emitting direction of the light source and the plane where the calibration plate is located is set and read through the angle adjustment mechanism of the light source fixing block; The spatial distance from the center of the light-emitting surface of the light source to the center of the target area is determined by the displacement of the lateral telescopic mechanism of the light source support. If necessary, it can be directly measured using a ruler.
[0117] Figure 5(a) shows the contour distribution of the evaluation function in the brightness and distance parameter space, and Figure 5(b) shows the corresponding three-dimensional response surface plot.
[0118] Under a fixed incident angle of approximately 20°, the evaluation function exhibits a single low-value region distribution characteristic within the parameter space of light source brightness and distance, indicating a stable optimal parameter range. Contour plots and 3D response surface visualizations show that the optimal region is concentrated around 40%–45% brightness and approximately 250–270 mm distance, consistent with experimental observations. This verifies the effectiveness and stability of the proposed evaluation function and illumination adjustment method within this parameter range.
[0119] This invention constructs an adaptive illumination adjustment mechanism based on image feedback, using the aforementioned grayscale statistical index and illumination evaluation function. This mechanism uses the mean grayscale value of the Region of Interest (ROI). With gray standard deviation As the state feedback variable, with the evaluation function S as the optimization objective, the evaluation function S is gradually reduced by iteratively adjusting the parameters of light source brightness, light source distance, and incident angle, thereby achieving adaptive optimization of lighting conditions. The above process constitutes a closed-loop feedback adjustment mechanism based on the evaluation function S.
[0120] Specifically, during image acquisition, the image under the current lighting conditions is first obtained, and the corresponding ROI region is calculated. , and the evaluation function S; subsequently, based on the grayscale statistics and preset adjustment rules, the illumination parameters are adjusted, wherein: when When the target grayscale range is deviated from, the brightness and distance of the light source should be adjusted first; when When the set threshold is exceeded, the distance to the light source and the incident angle are adjusted first to improve the uniformity of illumination; when a highlight area is detected, the brightness of the light source is reduced and the incident angle is adjusted so that the reflected light avoids the direction of the camera's line of sight.
[0121] After updating the parameters, the image is reacquired and a new evaluation function S value is calculated, and compared with the result of the previous round. When the S value decreases, the current parameters are retained; when the S value increases, the corresponding adjustment parameters are corrected in the opposite direction, and the adjustment step size is appropriately reduced. Preferably, the light source brightness adjustment step size is set to 5%~10%, the light source distance adjustment step size is set to 10mm~20mm, and the incident angle adjustment step size is set to 5°~10°.
[0122] When the evaluation function S changes by less than a preset threshold during multiple consecutive iterations The parameter update process stops when the maximum number of iterations is reached, thus obtaining a stable combination of illumination parameters. Through the aforementioned finite step size and termination condition constraints, the illumination adjustment process gradually converges to a better state while ensuring stability, achieving adaptive optimization of illumination parameters.
[0123] In summary, this invention, through optimized light source structure design, coordinated control of illumination parameters, and an adaptive adjustment mechanism based on grayscale statistical feedback, can construct a stable, uniform, and controllable imaging illumination environment. This effectively suppresses localized highlights and shadows on highly reflective industrial parts surfaces, improving the consistency and uniformity of image grayscale distribution. Furthermore, stable and clear grayscale responses can be obtained in the calibration board corner areas and target object edge areas, thereby reducing corner detection errors, improving the reliability of feature extraction, and further enhancing the stability and accuracy of the pose solving process.
[0124] Figure 6This is an example of RGB and depth images of an industrial part acquired. Through the above acquisition process, multi-view RGB-D data of industrial parts with continuous pose changes, uniform viewpoint coverage, and stable imaging quality can be obtained. This provides a unified and reliable raw data foundation for subsequent calibration board pose detection, target object true pose solving, and automatic generation of pose annotation data.
[0125] S3: Automatic generation of target pose based on calibration board.
[0126] After completing the multi-view RGB-D data acquisition of S2 industrial parts, this invention proposes an automatic target pose generation method based on calibration board constraints, error screening, and fixed spatial transformation propagation. Building upon calibration board pose detection, this method further introduces a data quality screening mechanism with reprojection error constraints, an initial frame point cloud quality enhancement strategy for high-reflectivity scenes, and a temporal pose propagation model based on fixed spatial transformation relationships. This enables automated and high-precision generation of the target industrial part's true pose in the camera coordinate system.
[0127] Specifically, this invention first detects the corner features of the ArUco calibration board in the acquired images and combines them with the camera imaging model to solve the pose information of the calibration board in the camera coordinate system at each time point, thereby constructing a unified, stable and continuously changing reference coordinate system for the entire acquisition sequence. Subsequently, the reliability of the pose calculation results of each frame is evaluated by reprojection error analysis, and abnormal frames and low-quality data are filtered out to improve the accuracy, consistency and robustness of subsequent pose ground truth data.
[0128] Based on this, in order to address the problem that the high reflectivity of industrial parts surfaces can easily cause depth jumps, voids and noise points, this invention introduces depth continuity constraints and normal consistency constraints in the initial frame pose solving process, performs quality enhancement processing on the scene point cloud, and combines the target region constraint registration strategy to complete the high-precision matching between the target scene point cloud and the CAD model, thereby establishing a fixed spatial mapping relationship between the target object coordinate system and the calibration board coordinate system.
[0129] At any subsequent moment, by simply combining the pose of the calibration board in the current frame, the true pose of the target object in the camera coordinate system can be directly derived through the coordinate chain transfer relationship, without the need to repeatedly perform point cloud matching and iterative optimization frame by frame. Therefore, this invention transforms the traditional pose solving mode that relies on frame-by-frame point cloud registration into a temporal pose propagation mode based on reference coordinate constraints. This significantly reduces computational complexity while effectively suppressing error accumulation and registration drift problems, thereby improving the efficiency of true pose generation and spatial annotation accuracy, and providing reliable support for the high-quality construction of industrial part pose estimation datasets.
[0130] Specifically, in the pose detection process, ArUco markers are first detected in the RGB image, and the corner pixel coordinates of each marker are extracted. Let the two-dimensional pixel coordinates of the detected i-th corner point in the image be:
[0131] ,
[0132] in, The number of valid feature points detected. , These represent the horizontal and vertical pixel coordinates of the corner point in the image plane, respectively.
[0133] Since images are discrete sampled data, directly detected corner positions typically only have pixel-level accuracy. To improve the accuracy of corner localization, this invention further employs a sub-pixel-level optimization method to iteratively correct corner positions. This method constructs a grayscale distribution model within the corner's neighborhood and refines the corner position with the goal of minimizing grayscale error, thereby obtaining more accurate corner coordinates. After optimization, the impact of pixel quantization error on the pose calculation results can be effectively reduced.
[0134] After obtaining the precise two-dimensional pixel coordinates, the three-dimensional spatial coordinates of each corner point in the calibration plate coordinate system are determined based on the known geometry of the calibration plate. Since the calibration plate is a planar structure, its three-dimensional coordinates can be expressed as:
[0135] ,
[0136] in, The total number of valid feature points detected. , This represents the position coordinates of the i-th feature point in the calibration plate plane coordinate system. This establishes the two-dimensional image pixel point. With three-dimensional space points The correspondence between them provides geometric constraints for subsequent pose solving.
[0137] Based on the above correspondence, a pinhole camera imaging model is introduced to describe the projection relationship from a three-dimensional point to a two-dimensional image plane:
[0138] ,
[0139] in, , These are the two-dimensional pixel coordinates of the feature points on the calibration board in the image. , , These are the three-dimensional spatial coordinates of the corresponding feature points in the calibration plate coordinate system. Let be the camera extrinsic matrix, where For rotation matrix, This is a translation vector used to describe the spatial transformation relationship between the calibration board coordinate system and the camera coordinate system. As a scale factor, Camera intrinsic parameter matrix:
[0140] ,
[0141] in, , Indicates focal length. , The principal point coordinates are represented; this model describes the imaging process of a point in three-dimensional space in an image and is the basis for pose solving.
[0142] Based on the aforementioned 2D-to-3D correspondence, the pose parameters are solved by minimizing the reprojection error. The optimization objective is to minimize the error between the 3D point projected onto the image plane after the current pose transformation and the actual detected pixel position. The obtained rotation vector r is converted into a rotation matrix using the Rodrigues formula. :
[0143] ,
[0144] Finally, a homogeneous transformation matrix from the calibration board coordinate system to the camera coordinate system can be constructed:
[0145] ,
[0146] Where Tc−b represents the complete pose information of the calibration board in the camera coordinate system, including its spatial position and spatial orientation.
[0147] Table 2 Example of pose detection results for the calibration board in frame i.
[0148]
[0149] By repeating the above process for each frame of the acquired sequence, the continuous pose changes of the calibration board during the entire acquisition process can be obtained, providing basic data for subsequent multi-coordinate system transformation calculations and automatic generation of the target object's true pose.
[0150] To evaluate the accuracy and stability of pose calculation, this invention further introduces a reprojection error analysis mechanism to assess the quality of each frame's pose result. Specifically, the 3D feature points in the calibration board coordinate system are reprojected onto the image plane using the obtained pose transformation matrix to obtain the corresponding predicted pixel coordinates. And compared with the actual detected corner pixel coordinates The comparison is performed. The reprojection error of the i-th feature point is defined as:
[0151] ,
[0152] in, This represents the Euclidean distance error of the i-th feature point in the image plane.
[0153] Based on this, to further quantify the overall error level of the current frame, this invention uses the root mean square error (RMSE) as an evaluation metric, and its calculation formula is as follows:
[0154] ,
[0155] in, This represents the number of valid feature points involved in the calculation. RMSE comprehensively reflects the overall deviation of all feature points, and compared to single-point errors, it can more stably characterize the accuracy of the pose calculation results for the current frame.
[0156] Based on the above error evaluation results, this invention performs filtering processing on each frame of data in the acquisition sequence. Preferably, the RMSE threshold is set comprehensively based on factors such as camera resolution, calibration board size, and the reflective properties of industrial parts surfaces. In this embodiment, the RMSE threshold is preferably set to 1 pixel. When the RMSE is not greater than this threshold, the pose calculation result of that frame is considered reliable and can be used for subsequent pose calculation and data annotation. By introducing a quality assessment and filtering mechanism based on reprojection error, low-quality data can be effectively filtered out, and pose results with large errors can be avoided from participating in subsequent calculations, thereby improving the accuracy, consistency, and stability of the dataset in the pose annotation process.
[0157] After obtaining reliable calibration board pose information, this invention further calculates the ground truth (GT) pose of the target object in the camera coordinate system. Traditional methods typically perform point cloud registration to solve for the target pose for each frame of data. However, this method involves a large computational load and is easily affected by noise, initial values, and local optima, leading to cumulative errors in the pose results over time. Especially for highly reflective industrial parts, local missing values, depth jumps, and outliers are prone to appear in the depth map, further reducing the stability of point cloud registration.
[0158] Figure 7This invention presents a schematic diagram of ground truth (GT) pose transformation (i.e., GT pose is the true pose). Addressing the aforementioned issues, this invention proposes a pose solving method based on a fixed spatial transformation relationship. Furthermore, considering the high reflectivity and lack of texture information on industrial parts surfaces, it collaboratively introduces depth continuity constraints, normal consistency constraints, and region constraint registration mechanisms during the initial pose solving process. This jointly constrains the point cloud data quality, the registration area range, and the rationality of the pose solution, thereby improving the stability and accuracy of the initial pose solution. The core idea is to establish the relative spatial relationship between the target object and the calibration board through a high-precision point cloud registration in the initial frame. In subsequent frames, the pose of the target object is directly derived from the changes in the calibration board's pose. This transforms the traditional frame-by-frame point cloud registration calculation process into a pose propagation process based on a fixed spatial transformation, effectively avoiding the error accumulation and computational overhead caused by repeated frame-by-frame registration.
[0159] Specifically, in the initial frame, a scene point cloud is first generated based on the depth map acquired by the depth camera. Let the pixel coordinates in the depth map be... The corresponding depth value is Its neighborhood pixel set is denoted as To address the depth anomalies caused by high reflectivity on industrial parts surfaces, a continuity constraint is applied to the depth data before constructing the point cloud. The basic principle is that on a continuous surface of a real object, the depth change between adjacent pixels is usually gradual, while abnormal regions caused by reflection or measurement errors often exhibit significant abrupt depth changes. Therefore, the validity of a point as a valid surface point is determined by detecting the depth difference between a pixel and its neighborhood. When the following conditions are met:
[0160] ,
[0161] The 3D point corresponding to the pixel is retained when the pixel is active, otherwise it is discarded. Represents pixel depth value, For neighboring pixels, This is the depth continuity threshold.
[0162] Furthermore, each point is estimated in the point cloud space. unit normal vector The normal vector describes the spatial orientation of the local surface where the point is located. On a continuous, smooth surface, the normal vector directions of adjacent points should be basically consistent, while in noisy points or anomalous regions, the normal vector directions usually vary significantly. Therefore, by comparing the angles between the normal vectors of neighboring points, anomalous points are further filtered out. For neighboring points... The normal vector is When the following conditions are met: If a point is considered to satisfy the normal consistency constraint, it is considered to be discarded; otherwise, it is discarded. This is the normal angle threshold.
[0163] Based on this, to reduce the interference of the calibration board and background structure on the registration process, this invention first obtains the region information of the target object in the image plane. In the initial frame, since the pose of the target object is unknown, manual initialization is used to coarsely locate the region where the target object is located, generating an initial target region mask. .in, Represents pixels It belongs to the target object region. Represents pixels This area is not the target region. Only for those that meet the following conditions... The pixel corresponds to the 3D point and is back-projected to obtain the point cloud of the target area. It is used for the registration calculation of target objects and model point clouds in the initial frame, thereby transforming the point cloud registration process from global scene constraints to local constraints of the target region, so as to reduce the interference of background structure on the matching results and improve the registration stability.
[0164] parameter and The depth continuity threshold is determined based on the camera's accuracy, acquisition distance, and the size range of the industrial parts. Preferably, it is obtained through statistical analysis of local depth fluctuations and normal variation distributions. In this embodiment, acquisition is performed on an industrial part with dimensions of approximately 98mm × 51mm × 21mm at an acquisition distance of approximately 500mm. Experiments show that the 95th percentile of the local depth difference is approximately 2.0mm, and the 99th percentile is approximately 3.0mm. Therefore, the depth continuity threshold... Within the range of 1.0mm to 3.0mm, therefore More preferably, it is about 2.0 mm; the 90th percentile of the angle between the normals of adjacent points is about 29.2°, and the 95th percentile is about 39.4°. Therefore, the normal angle threshold τn is between 25° and 45°. More preferably, it is about 30°.
[0165] Based on the above parameter settings, while preserving the continuous surface structure of the main body of industrial parts, abnormal points caused by high reflectivity and measurement errors can be effectively eliminated, avoiding the introduction of depth abrupt change points and normal abnormal points into the point cloud registration process, thereby improving the geometric consistency and stability of the scene point cloud, and further enhancing the accuracy and robustness of subsequent pose solving.
[0166] Processed target area point cloud Point cloud of the CAD model of the target object Registration is performed. The ICP algorithm is used to register the two sets of point clouds. Essentially, it seeks an optimal rigid transformation so that the model point cloud, after rotation and translation, is as aligned as possible with the scene point cloud. The optimization objective can be expressed as:
[0167] ,
[0168] in, This represents the i-th point in the scene point cloud. This represents the corresponding point in the model's point cloud. For rotation matrix, Let N be the translation vector, and N be the number of points involved in the registration. By solving this optimization problem, the pose transformation matrix of the target object in the initial frame can be obtained. :
[0169] ,
[0170] This matrix aligns the CAD model with the real object in three-dimensional space and serves as the benchmark for the entire sequence pose calculation.
[0171] Meanwhile, based on the calibration board pose detection results in this step, the pose transformation matrix of the calibration board in the camera coordinate system in the initial frame can be obtained. Since the target object is fixedly placed on the calibration plate during the acquisition process, and the two remain relatively stationary throughout the entire acquisition process, a fixed spatial transformation relationship between the target object coordinate system and the calibration plate coordinate system can be established through the initial frame. The calculation method is as follows:
[0172] ,
[0173] in, This indicates the transformation relationship from the object's coordinate system to the calibration plate's coordinate system. This is the transformation matrix from the camera coordinate system to the calibration board coordinate system in the initial frame. This is the pose transformation matrix of the target object in the camera coordinate system in the initial frame, which is calculated by the point cloud registration method. This transformation relationship reflects the fixed installation position of the target object in the calibration board reference coordinate system, and therefore remains unchanged throughout the entire acquisition sequence.
[0174] At any subsequent time t, simply continue using the method described in this step to obtain the pose of the calibration board in the camera coordinate system. The true pose of the target object in the camera coordinate system can be directly calculated using the following coordinate chain relationship:
[0175]
[0176] in, Indicates time The true pose of the target object. The pose of the calibration board relative to the camera in the current frame. This establishes a fixed spatial transformation relationship in the initial frame. Essentially, the calculation process uses a calibration board as an intermediate reference coordinate system, mapping its pose changes over time to the target object, thus enabling continuous derivation of the target object's pose. Compared to traditional frame-by-frame point cloud registration methods, this method eliminates the need to repeatedly perform point cloud matching in each frame; pose calculation can be completed solely through the multiplication of homogeneous transformation matrices.
[0177] Figure 8 The above formula is used to obtain the GT pose visualization result of industrial parts from the camera's perspective. This method transforms the originally computationally complex, initial condition-sensitive, and error-accumulating 3D point cloud registration problem into a stable and efficient algebraic operation process. This not only significantly reduces the computational load and improves data annotation efficiency, but also effectively avoids the cumulative propagation of pose errors in the time series, thereby ensuring the temporal consistency and spatial accuracy of the entire dataset.
[0178] S4: Data labeling and dataset construction.
[0179] In step S3, the true pose of the target object in each frame is obtained. Then, based on this pose information, the acquired RGB and depth images can be automatically labeled to construct a structured dataset.
[0180] Specifically, for any time t, the three-dimensional coordinates of each vertex in the 3D CAD model of the target object in the object coordinate system are known. And the pose transformation matrix of the target object in the camera coordinate system at that moment. First, the 3D model points can be transformed to the camera coordinate system:
[0181] ,
[0182] in, and They are respectively The rotation matrix and translation vector in the diagram. This represents the three-dimensional coordinates of the model points in the camera coordinate system.
[0183] Subsequently, based on the pinhole camera imaging model, the three-dimensional points can be projected onto the image plane, and its expression is:
[0184] ,
[0185] in, For the camera intrinsic parameter matrix, , These are the projected pixel coordinates. The scale factor is used. By projecting the model surface point by point or by rasterizing it based on a triangular mesh, the two-dimensional projection area of the target object in the image can be obtained.
[0186] Figure 9 The pixel-level segmentation mask obtained by the above method for industrial parts realizes the automatic mapping process from three-dimensional pose to two-dimensional annotation, so that the generation of segmentation mask does not depend on manual annotation and can ensure the spatial consistency and accuracy of annotation results.
[0187] Furthermore, the collected data is structured and stored to form a standardized dataset format.
[0188] Specifically, the acquired RGB images, depth images, and corresponding segmentation masks are numbered according to frame order and stored in separate data directories. All types of data are named with the same frame index number to ensure the correspondence between different modal data.
[0189] At the same time, the target object pose information and related annotation data corresponding to each frame are uniformly stored in a structured annotation file, preferably organized in YAML format.
[0190] Each sample's annotation information includes the target object's rotation matrix cam_R_m2c, translation vector cam_t_m2c, 2D bounding box obj_bb in the image, and target category identifier obj_id in the camera coordinate system. The rotation matrix cam_R_m2c and translation vector cam_t_m2c are directly derived from the target object's ground truth (GT) pose calculated in step S3, representing the target object's spatial pose and position relative to the camera. The 2D bounding box obj_bb is automatically generated during the data annotation and dataset construction process in this step. Specifically, the points of the target object's 3D CAD model are transformed to the camera coordinate system using the GT pose matrix, projected onto the image plane using camera intrinsic parameters, and then the minimum bounding rectangle is calculated based on the resulting projection point set, thus obtaining the corresponding 2D bounding box. The target category identifier obj_id is a unique category number pre-assigned to each CAD model during the dataset construction phase, and is associated with the corresponding pose and bounding box information. The above method enables the unified organization and structured storage of target object pose parameters, 2D detection boxes, and category labels, providing standardized labeled data for subsequent tasks such as pose estimation, target detection, and instance segmentation.
[0191] In summary, the platform and processing flow constructed by this invention can achieve stable acquisition of multi-view RGB images and depth image data of weak texture and high reflectivity industrial parts in a controlled environment, and can automatically obtain the true pose information of the target object and pixel-level segmentation mask that strictly correspond to each frame of the image.
[0192] Based on the above processing results, a structured dataset can be formed, including RGB images, depth images, target object true pose labels, and segmentation masks. All types of data maintain a correspondence in the time series and have a unified reference coordinate in the spatial scale, thereby ensuring the consistency and accuracy of the data.
[0193] The constructed dataset features continuous pose changes, uniform view coverage, high annotation accuracy, and strong time series consistency. It can be directly used for pose estimation model training and testing, and can also be extended to tasks such as target detection, instance segmentation, and robot grasping and localization, providing high-quality data support for industrial vision algorithm development and practical engineering applications.
[0194] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be considered within the scope of protection of the present invention.
Claims
1. An industrial part pose estimation data collection platform, characterized by: include Horizontal base (9); An ArUco calibration plate (5) with characteristic points of known geometry, the ArUco calibration plate (5) being used to support industrial parts; A rotating platform (8) is mounted on a horizontal base (9) for driving the ArUco calibration plate (5) to rotate; Camera (4), which is fixedly mounted on a horizontal base (9) for synchronously acquiring RGB images and depth images of industrial parts; The first light source (2) and the second light source (6) are symmetrically arranged on both sides of the rotating platform (8), and the position and height of the first light source (2) and the second light source (6) relative to the horizontal base (9) are adjustable, and the incident angle of the first light source (2) and the second light source (6) is adjustable.
2. The industrial part pose estimation data collection platform of claim 1, wherein: Both the first light source (2) and the second light source (6) are LED fill lights, and both the first light source (2) and the second light source (6) are equipped with a diffuser or a diffuser structure at their light-emitting ends.
3. The industrial part pose estimation data collection platform of claim 1, wherein: The first light source (2) and the second light source (6) are respectively set on the horizontal base (9) by light source brackets. The light source brackets include a longitudinal lifting structure, a transverse telescopic structure and an angle adjustment structure, which are used to adjust the height, horizontal distance and incident angle of the corresponding light source, respectively.
4. The industrial part pose estimation data collection platform of claim 1, wherein: The ArUco calibration plate (5) is provided with multiple ArUco markers, which are arranged in a circular or arrayed manner. The corner points of the ArUco markers constitute the feature points of the known geometric structure.
5. An industrial part pose estimation data processing method, based on the data acquisition platform of any one of claims 1-4, characterized in that: Includes the following steps: S1: Fix the industrial parts on the ArUco calibration plate (5), drive the ArUco calibration plate (5) and the industrial parts to rotate synchronously through the rotating platform (8), and synchronously acquire multi-view RGB images and depth images of the industrial parts through the camera (4); S2: In each frame of the image, detect the feature points of the ArUco calibration board (5), extract the two-dimensional pixel coordinates of the feature points, combine the known three-dimensional spatial coordinates of the feature points in the coordinate system of the ArUco calibration board (5), solve the pose transformation matrix of the ArUco calibration board (5) in the camera coordinate system, and filter the data frames based on the reprojection error. S3: In the initial frame, the depth image is enhanced, and the target area point cloud is extracted based on the target area constraint. The initial pose of the industrial part in the camera coordinate system is calculated by point cloud registration. The fixed spatial transformation relationship between the industrial part coordinate system and the calibration board coordinate system is established by combining the pose transformation matrix of the ArUco calibration board (5) in the initial frame. In subsequent frames, the true pose of the industrial part in the camera coordinate system is calculated based on the pose transformation matrix of the ArUco calibration board (5) in the current frame and the fixed spatial transformation relationship. S4: Based on the real pose, project the 3D model of the industrial part onto the image plane to generate the target region segmentation mask and 2D bounding box. Associatively store the RGB image, depth image, segmentation mask, 2D bounding box and real pose to construct a structured industrial part pose dataset.
6. The industrial part pose estimation data processing method of claim 5, wherein: In S1, the pixel grayscale values within the target area are statistically analyzed to construct a lighting evaluation function. Based on this function, the light source brightness, incident angle, and light source distance are iteratively adjusted in a closed loop. The lighting evaluation function is as follows: , in, The grayscale mean is... The standard deviation of grayscale For the target grayscale value, These are the weighting coefficients; A quadratic polynomial response surface model is constructed to relate the light source brightness, incident angle, and distance to the light source to the illumination evaluation function. Based on the response surface model, the optimal range of illumination parameters is determined.
7. The industrial part pose estimation data processing method as described in claim 5, characterized in that: In S2, the three-dimensional feature points in the calibration board coordinate system are projected onto the image plane through the pose transformation matrix to obtain the predicted pixel coordinates. The reprojection error between the predicted pixel coordinates and the actual detected pixel coordinates is calculated, and the overall pose solution result of the current frame is evaluated based on the root mean square error. When the root mean square error is greater than a preset threshold, the corresponding data frame is filtered out.
8. The industrial part pose estimation data processing method as described in claim 5, characterized in that: The reprojection error is: , in, These are the actual corner pixel coordinates obtained from the detection. The predicted pixel coordinates of the feature points projected onto the image plane after pose transformation; The root mean square error is: , in, This represents the number of valid feature points used in the calculation.
9. The industrial part pose estimation data processing method as described in claim 5, characterized in that: In S3, the depth image quality enhancement processing includes: Depth continuity constraint is used to remove pixels whose depth difference with neighboring pixels exceeds the depth continuity threshold; normal consistency constraint is used to remove pixels whose angle with the normal of neighboring points exceeds the normal angle threshold; and based on the target region mask, only the 3D points corresponding to the target region are back-projected to obtain the target region point cloud.
10. The industrial part pose estimation data processing method as described in claim 5, characterized in that: In S3, the fixed spatial transformation relationship is established by the following formula: in, This is a fixed spatial transformation matrix from the coordinate system of the industrial parts to the coordinate system of the calibration plate. Let be the pose transformation matrix of the calibration board in the camera coordinate system in the initial frame. This is the pose transformation matrix of the industrial parts in the camera coordinate system in the initial frame; In subsequent frames, the true pose of the industrial part in the camera coordinate system is calculated using the following formula: , in, This is the pose transformation matrix of the calibration board in the camera coordinate system in the current frame. This represents the true pose of the industrial parts in the current frame within the camera coordinate system.
11. The industrial part pose estimation data processing method as described in claim 5, characterized in that: In S4, the vertices of the 3D model of the industrial part are transformed to the camera coordinate system through the real pose, and then projected onto the image plane in combination with the camera intrinsic parameters to obtain a 2D projection area. A target region segmentation mask is generated based on the 2D projection area, and a 2D bounding box is generated based on the minimum bounding rectangle of the projection point set.