Openable container three-dimensional reconstruction method based on structural characteristics and operation interaction
By extracting plane profile features in a home environment and using the Welsch loss function and gradient information, the three-dimensional reconstruction accuracy and static dynamic component distinction problems of openable containers are solved, and a three-dimensional reconstruction with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510316329.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing three-dimensional reconstruction technology is difficult to effectively deal with openable containers that lack texture in the home environment, especially planar structures, which leads to low reconstruction accuracy and difficulty in distinguishing between static substrates and dynamic components, resulting in inaccurate reconstruction results.
By extracting plane contour features as geometric constraints, combining Welsch loss function and truncated symbol distance function of gradient information, point cloud registration is improved, and static and dynamic components are distinguished through operational interaction, and the three-dimensional reconstruction process is optimized using the aggregate hierarchical clustering algorithm and iterative closest point algorithm.
It improves the reconstruction accuracy in planar structure scenarios, enhances the robustness of point cloud registration, effectively separates static and dynamic components, and improves the quality and speed of three-dimensional reconstruction.
Smart Images

Figure CN120388144A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional reconstruction, and relates to a three-dimensional reconstruction method for an openable container based on structural features and operation interaction. Background Art
[0002] In the field of domestic service robots, especially when operating on openable containers such as cabinets, drawers, refrigerators, etc. in the home environment, the robot needs to have the ability to perform three-dimensional reconstruction and understanding of the container and search for and classify the items inside the container. To this end, robots usually rely on three-dimensional reconstruction technology to understand the structure of the environment and perform operations. The existing three-dimensional reconstruction technologies mainly use the following several types:
[0003] Texture-based reconstruction methods: These methods rely on the texture information on the surface of objects in the scene and calculate the geometric structure of the object by identifying the texture. However, for openable containers in the home environment, texture information is usually lacking. Especially for containers with a planar structure, the surface is usually relatively simple and lacks texture, resulting in low accuracy of texture-based methods when dealing with these structures. It is unable to effectively capture the relative position relationship between planes, thus affecting the overall reconstruction effect and making it difficult to be effectively applied.
[0004] Iterative closest point algorithm (ICP): The ICP algorithm is usually used to estimate the pose of an object or scene by matching multiple three-dimensional point clouds. This algorithm is usually applicable to relatively complex scenes with rich texture features, but it will be affected by noise and mismatched points on simple geometric structures such as planes, resulting in inaccurate reconstruction results. Especially in a complex home environment, the accuracy of point cloud registration is greatly reduced due to changes in lighting and viewing angles. Mismatched points have a serious impact on the accuracy of the reconstruction result, especially when the boundary between target objects or between an object and the background is relatively blurred. The robustness and accuracy of ICP often cannot meet the actual application requirements.
[0005] Deep learning methods: In recent years, deep learning technology has been widely used in three-dimensional reconstruction, especially convolutional neural networks (CNNs) and image segmentation technology. However, these methods usually rely on a large amount of labeled data and often have difficulty dealing with scenes containing simple geometric features, such as common openable containers in the home. For these containers, existing deep learning models may not be able to effectively infer the complete three-dimensional structure from a single image or video stream.
[0006] In addition, since an openable container generally includes a static base body and movable components such as doors, drawers, etc., existing 3D reconstruction methods usually have difficulty effectively distinguishing these dynamic components and the static base body. Especially when the dynamic components move or change, it is difficult to fuse the data of the static structure and the dynamic components. This results in the mixing of the moving parts and the static parts of the container during the reconstruction process, making it difficult to accurately restore the complete structure of the container. Summary of the Invention
[0007] In order to solve the above technical problems existing in the prior art, the present invention proposes a 3D reconstruction method for an openable container based on structural features and operation interaction. By combining the geometric constraints of planar features, the robust processing of point cloud registration using the Welsch loss function, and the method of separating dynamic and static structures based on operation interaction, these defects in the prior art are overcome. The specific technical solution is as follows:
[0008] A 3D reconstruction method for an openable container based on structural features and operation interaction includes the following steps:
[0009] Step 1: Obtain consecutive frame depth images of the scene where the openable container is located, and extract the planar contour features of the depth images as geometric constraints for pose estimation of adjacent frames;
[0010] Step 2: Represent the 3D model using the truncated signed distance function GSDF with gradient information, and fuse the depth data to reconstruct the static 3D model of the openable container;
[0011] Step 3: Dynamically analyze the opening and closing states of the openable container, distinguish the dynamic components and the static base body of the openable container, and mark the dynamic components to optimize the 3D reconstruction process.
[0012] Further, in Step 1, the planar contour features are extracted from the depth images through an agglomerative hierarchical clustering algorithm.
[0013] Further, in Step 1, the Welsch function is used as the loss function for the distance between corresponding points in pose estimation.
[0014] Further, in Step 2, the truncated signed distance function GSDF with gradient information stores the gradient information of the point cloud data, and calculates the distance between each input point and the nearest point on the model surface using the first-order Taylor expansion method.
[0015] Further, in Step 2, the depth data is preprocessed, specifically including: taking the k-th frame depth image R k as the input, performing depth image denoising, normal vector calculation, and planar contour extraction operations to obtain the denoised 3D point cloud V k and the normal vector N of the point cloud kand the planar contour C in the depth image k .
[0017] Further, in step 2, the 3D model represented by GSDF is reconstructed by fusing depth data, specifically including: for the input k-th frame depth image R k , calculating its normal vector N k and the relative pose T g,k , the weight value W k (p) is always equal to 1, and the truncated signed distance value is calculated by the following equation:
[0018]
[0019] where q = π(p) represents the perspective projection process of the camera, represents finding the pixel point with the closest distance in the depth image; after obtaining the truncated signed distance value, the GSDF value is updated using the following formula:
[0020]
[0021] Further, the normal vector is calculated using the approximate least squares method, where it is assumed that the depth values of the target pixel point and adjacent pixel points in the depth image are approximately equal, and a box filter is used to accelerate the calculation.
[0022] Further, in step 3, images of the openable or closable container are collected and the 3D model of the corresponding state is reconstructed, and the static part point cloud and the movable part point cloud of the openable container are extracted based on the structural information difference.
[0023] Further, the extraction of the static part point cloud and the movable part point cloud of the openable container based on the structural information difference specifically includes:
[0024] Let the 3D point cloud of the container with the movable part closed be P C , the RGB image information collected in this coordinate system be RGB C , the 3D point cloud of the container with the movable part opened be P O , the RGB image information collected in this coordinate system be RGB O ;
[0025] Detect the movable part in the images RGB C and RGB O , and construct the corresponding relationship between the movable parts in RGB C and RGB O based on the nearest neighbor search; combining the corresponding relationship between the RGB image and the 3D point cloud information, construct the corresponding relationship between the movable part point clouds in P C and P O ;
[0026] Calculate the pose transformation matrix between P C and P O using the Iterative Closest Point (ICP) algorithm, and transform P C and P O to the same coordinate system. Then, the pose transformation matrix between the static components in the two point clouds can be considered as the identity matrix, while the pose transformation matrix between the movable components is obtained by solving using the ICP algorithm; Based on the existing pose transformation matrix between the movable components, further extract the structural information of the static and movable components in P and P
[0027] If the three-dimensional point p in the point cloud P C is a point in the movable component s, its transformation matrix can transform p to the corresponding point p' in the point cloud P O . To determine the best correspondence between p and the component, apply all the obtained rigid body transformation matrices to the point p and calculate the distance d C to the closest point in P O : O d s (p):
[0028] d s (p) = ||T s · p - p'||
[0029] When d s (p) < 0.02 m, the constructed correspondence is considered available, and the following confidence index is constructed:
[0030]
[0031] After calculating the confidence for point p corresponding to all other components, the component with the highest confidence is the component corresponding to point p.
[0032] Furthermore, the points in the three-dimensional information of the closed state of the movable component that cannot establish a correspondence are all points of the static component, and the points in the three-dimensional information of the open state of the movable component that cannot establish a correspondence are all points of the movable component.
[0033] The method of the present invention has the following remarkable advantages:
[0034] 1. Improve the reconstruction accuracy in the plane structure scene: By combining the plane contour features and geometric constraints, improve the pose estimation between adjacent frames, and solve the problem of insufficient accuracy of the existing methods when dealing with simple geometric structures.
[0035] 2. Improved the robustness of point cloud registration: The Welsch loss function was adopted to suppress the influence of mismatched points, reducing the interference of noise and mismatched points on the results.
[0036] 3. Effectively separated dynamic and static components: By operating on the interaction information, the static base and dynamic components in the openable container were accurately distinguished, providing more reliable perception data for the robot to perform item access and storage, and improving the reconstruction accuracy and stability of the three-dimensional structure of the container.
[0037] 4. Improved the quality and speed of three-dimensional reconstruction: The truncated signed distance function based on gradient information greatly improved the efficiency of point cloud data processing and improved the reconstruction quality. Brief Description of the Drawings
[0038] Figure 1 is the overall flowchart of a three-dimensional reconstruction method for an openable container based on structural features and operation interactions according to an embodiment of the present invention;
[0039] Figure 2 is the flowchart of plane extraction based on agglomerative hierarchical clustering according to an embodiment of the present invention;
[0040] Figure 3 is an example diagram of plane extraction based on agglomerative hierarchical clustering according to an embodiment of the present invention;
[0041] Figure 4 is an example diagram of contour extraction based on edge detection according to an embodiment of the present invention;
[0042] Figure 5 is an example diagram of contour boundary extraction based on plane features according to an embodiment of the present invention;
[0043] Figure 6 Comparison diagram of the contour extraction effects between the method of the embodiment of the present invention and the LSD method;
[0044] Figure 7 is the algorithm flowchart of reconstructing a static three-dimensional model using GSDF and fusing depth data according to an embodiment of the present invention;
[0045] Figure 8 is the flowchart of point cloud extraction of the movable and static components of the openable container according to an embodiment of the present invention. Detailed Embodiment
[0046] In order to make the objectives, technical solutions, and technical effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings of the specification and embodiments.
[0047] As Figure 1 shown, this embodiment discloses a three-dimensional reconstruction method for an openable container based on structural features and operation interactions, including the following steps:
[0048] Step 1: Obtain consecutive frame depth images of the scene where the openable container is located through a depth camera or an RGB-D camera, extract the planar contour features of the depth images, and constrain the pose estimation of adjacent frames to solve the problem of relatively large pose estimation errors that may occur in a scene mainly constructed with simple planes.
[0049] Specifically, the appearance of an openable container is significantly different from that of a conventional three-dimensional object. For example, its surface structure is mainly planar and may lack the surface texture relied on by general three-dimensional reconstruction methods. Using the geometric features in the scene as constraints can improve the estimation of adjacent frame poses, such as line features, and the perpendicularity of adjacent planes. However, line features are usually not stable enough, and the requirement of having multiple mutually perpendicular planes is relatively demanding. Therefore, the present invention proposes to extract stable planar contour features in the scene as geometric constraints to enhance the robustness of relative pose estimation.
[0050] As Figure 2 shown, in this embodiment, stable planar contour features are extracted from the depth image through the Agglomerative Hierarchical Clustering (AHC) algorithm. The planar contour can effectively describe the structural features of the container, especially in scenes lacking texture or details.
[0051] The main steps of the AHC algorithm are as follows:
[0052] Construct a node graph, divide the input image into image blocks of the same size to define nodes, and use the adjacent relationship between image blocks to construct the node graph, as shown in (a) of Figure 2 ;
[0053] Merge similar nodes to achieve rough segmentation. According to the plane fitting error, merge nodes with errors within the threshold range, and iterate this process until all nodes are traversed, as shown in (b), (c), and (d) of Figure 2 ;
[0054] Optimize the segmentation boundary, try to associate other pixels adjacent to the rough segmentation boundary with the rough segmentation result, and perform node graph merging again to achieve precise segmentation of the image boundary, as shown in (e) of Figure 2 ;
[0055] By setting the size of the image block for each node and the minimum number of pixels included in each plane, the extraction of planes of different sizes can be achieved, as shown in Figure 3 . It can be seen that the extracted plane boundaries are relatively clear and can remove the influence of irrelevant data.
[0056] After obtaining the plane information in the input depth image, use the contour detection method for binary images to detect the contours of the planes in the image, asFigure 4 As shown in (b) of []. However, when there are holes in the input depth image, the hole contours will also appear in the edge detection results. To reduce redundant contour information, dilation and erosion operations are performed on the grayscale image representing the plane, and the improved contour extraction effect is as shown in Figure 4 (c) of [].
[0057] Set the boundary contours with a length greater than 2000 pixels as unavailable contours to avoid misdetecting the contours of the bottom boundary of the scene. Merge all available plane contours to obtain the structural contour pixels in the depth image. The contour extraction process is as shown in Figure 5 shown in
[0058] as Figure 6 shown in the comparison between the method of the present invention and the line segment extraction algorithm LSD (Line Segment Detector) algorithm. As can be seen from the figure, the scene straight line contours extracted by the LSD algorithm are usually discontinuous, as shown by the circles in the figure, and perform poorly in the indoor object images with continuous repeating textures. While the method of the present invention extracts the contours of the objects in the scene based on the structural information, can well extract the continuous contours of the objects, and is not affected by the continuous repeating textures in the objects.
[0059] Take the extracted plane contour features as geometric constraints to constrain the pose estimation between adjacent frames, that is, extract the contour boundary points of the plane structure in the depth image as the constraint points for solving the inter-frame pose, and weaken the pose solution error of the point-to-plane measurement method in the plane structure scene.
[0060] The plane contour features of each frame of image are aligned with the previous frame through geometric constraints to ensure the accuracy of pose estimation and improve the reconstruction effect in the textureless or simple geometric scene. The introduction of plane contour features significantly improves the accuracy of pose estimation in the scene with simple geometric structures such as planes and vertical planes, and reduces the possibility of pose drift. In actual operation, when the robot faces a cabinet, the edge contour features of the cabinet door are extracted from the depth image of the first frame, and then the same features are continuously extracted in the second frame of image, and geometric constraints are used to ensure the accurate pose estimation of the opening and closing actions of the cabinet door. This process provides an accurate pose reference for subsequent 3D reconstruction.
[0061] Aiming at the measurement noise and partial overlap problems faced in the input point cloud, and in order to further improve the accuracy of pose estimation, this embodiment uses the Welsch function as the loss function of the corresponding point distance in the pose estimation to suppress the influence of outliers and mismatched points. This minimizes the influence of mismatched points on the estimation result in complex scenes and improves the robustness of point cloud registration.
[0062] The Welsch function: Among them, v>0 is a settable hyperparameter. Since the function is monotonically increasing, matching points with large errors will lead to larger matching errors. Since the upper bound of the function value is 1, the loss function is insensitive to large matching errors caused by abnormal data in the input data and incorrect matching caused by partial overlap.
[0063] Step 2: Use the truncated signed distance function (GSDF) with gradient information to represent the 3D model and fuse the depth data to reconstruct the static 3D model of the openable container, such as Figure 3 As shown, including:
[0064] Depth data preprocessing: Take the kth frame depth image R k As input, perform depth image denoising, normal vector calculation and plane contour extraction operations to obtain the denoised three-dimensional point cloud V k , the normal vector N of the point cloud k and the plane contour C in the depth image k ;
[0065] Relative pose estimation: combined with C k 、V k and N k , calculate the pose matrix between the new input depth image and the constructed 3D model;
[0066] 3D information fusion: Use the pose T of the current frame g,k R k and N k integrated into the reconstructed model.
[0067] The GSDF is a hybrid structure that combines the advantages of explicit and implicit representation methods. Unlike the truncated signed distance function, in addition to the truncated signed distance value, it also stores the gradient information of the voxel and uses the first-order Taylor expansion to calculate the model surface point closest to the voxel without any additional calculation.
[0068] Specifically, the first frame of the input depth image is used as the global frame, GSDF voxels are constructed based on it, and a hash structure is used to manage all GSDF voxels to avoid allocating memory for areas without object information in the three-dimensional space and reduce memory consumption.
[0069] Use S to represent the GSDF field obtained by fusion of the depth images of the 1st to kth frames, and use S k (p) represents the point p∈R in the GSDF field under the global posture 3 The corresponding attribute value is calculated from the input depth image. k (p) stores the truncated signed distance value F at point p k (p), weight value W k (p) and the gradient value Nk (p), that is
[0070]
[0071] More specifically, based on the GSDF representation method, after obtaining the normal vector of the voxel through query, the distance from the input point to the nearest point on the model surface can be approximately calculated using the first-order Taylor expansion. When both the gradient and the signed distance are stored in the voxel, the point p on the three-dimensional surface closest to the voxel v j The nearest point p on the three-dimensional surface s can be calculated by Equation (2-2):
[0072] p s (v j ) = v j - ψ j g j (2-2)
[0073] where ψ j is the signed truncated distance at voxel v j ; g j is the gradient value at voxel v j .
[0074] When using the iterative closest point algorithm based on point-to-plane error to solve the relative pose, it is necessary to calculate the distance |d s (p)| from the input three-dimensional point to the nearest point on the model surface and the gradient of the corresponding point on the model For the input three-dimensional point p, its distance d S (p) to the model S and the gradient of the corresponding point on the model are:
[0075]
[0076] Based on the GSDF voxel representation, d s (p) and can be approximately calculated through one voxel lookup and first-order Taylor expansion:
[0077] d S (p) = ψ0 + (p - v j ) T g j (2-5)
[0078]
[0079] Since v j is regularly sampled in three-dimensional space, j can be calculated by the formula p / v s without using other tools to assist in determining the corresponding relationship.
[0080] Reconstructing a 3D model by fusing depth data through GSDF: For the input k-th frame depth image R k , calculate its normal vector N k and relative pose T g,k , the weight value W k (p) is always equal to 1, and the truncated signed distance value is calculated by the following equation:
[0081]
[0082] where q = π(p) represents the perspective projection process of the camera, represents finding the pixel point with the closest distance in the depth image. After obtaining the truncated signed distance value, the GSDF value is updated using the following formula:
[0083]
[0084] where the normal vector N k The calculation specifically includes:
[0085] For the target point and its neighboring points P = {p i |i = 1, 2,... k}, fit the plane equation n x x + n y y + n z z - d = 0, where (x, y, z) T represents a point on the plane, and (n x , n y , n z , d) are the plane parameters. Use the least squares method to calculate the normal vector (n x , n y , n z ) and the offset d are expressed as:
[0086]
[0087] The solution of the normal vector n in Equation (2-13) is the eigenvector corresponding to the smallest eigenvalue of the covariance matrix , where d is the mean value of the vertex positions to the distance of the calculated plane, that is This means that the complexity of calculating each vertex normal vector is proportional to the number of adjacent points k used, and it takes a long time.
[0088] To calculate the normal vector quickly, use the approximate least squares method to calculate the normal vector. Assume that the depth values of the target pixel point and its adjacent pixel points in the depth image are approximately equal, and use a box filter to accelerate the calculation. The process is as follows:
[0089] The point p = (x, y, z) in the space rectangular coordinate system TThe coordinates (r, θ, φ) in the spherical coordinate system can also be used T to represent, where θ represents the azimuth angle, φ represents the elevation angle, and r represents the distance to the center point of the camera. That is
[0090] p = rv (2-14)
[0091] where
[0092]
[0093] the coordinates (x, y, z) T corresponding spherical coordinates (r, θ, φ) T are:
[0094]
[0095] Dividing both sides of Equation (2-13) by d, assuming that the plane to be found does not pass through the origin, i.e., d!= 0, which is obviously true in the depth image, it is expressed as:
[0096]
[0097] where is the product of the normal vector n and the scale factor. Normalizing will obtain the normal vector n, that is The three-dimensional points are represented using the spherical coordinate system. Assuming that the depth values are close in the adjacent regions on the depth image, both sides of Equation (2-17) are divided by r i , and the optimized objective function can be simplified to:
[0098]
[0099] It can be solved as follows:
[0100]
[0101] where is only related to the position of the pixel on the image and has nothing to do with the specific depth value of a certain pixel position, and can be calculated in advance at the initialization stage of the program. The box filter technology can be used to accelerate the calculation.
[0102] In the 3D reconstruction process of this step, in addition to storing the conventional point cloud data, the gradient information of the voxels is also introduced. Through the method of first-order Taylor expansion, the 3D model surface point closest to the input point can be quickly calculated. Combining with the imaging characteristics of the depth camera, by using the truncated signed distance function, the noise problem of the point cloud can be effectively processed. Among them, in order to improve the processing speed and accuracy of the point cloud data, the present invention uses the approximate least squares method to calculate the normal vector, thereby accelerating the 3D reconstruction process.
[0103] This embodiment also points out that the distance between the target object and the depth camera, and the angle between the camera view angle and the normal of the measured surface will all affect the error of the measured depth value. The analysis of the measurement errors of several common RGB-D cameras shows that within the measurement range, the camera measurement error approximately increases with the increase of the distance from the object to the camera; when the angle between the camera view angle and the normal of the surface of the measured object is greater than 70°, the measurement error will increase significantly. Based on the above data, further select reasonable data in the depth image to be incorporated into the reconstructed 3D model, and the screening conditions are as follows:
[0104] Imaging distance: Combining the imaging characteristics of the RGB-D camera, select the maximum distance containing the information of the target object as the maximum reasonable distance of the depth image value, and this value varies accordingly in different scenarios.
[0105] Incident angle of the object imaging surface: When the incident angle formed by the camera line of sight and the normal of the measured surface is too large, the measured value may have a large error due to the Fresnel effect. Therefore, in order to exclude these potentially incorrect measurement values, combining the existing error information, only select the points where this angle is less than 70°.
[0106] Step 3: Dynamically analyze the opening and closing states of the openable container, distinguish the movable parts and static parts of the openable container, and mark the movable parts to optimize its 3D reconstruction process, providing reliable environmental perception information for the robot and improving the operation efficiency of the robot.
[0107] Reconstruct the static 3D structure of the movable parts in different motion states, extract the point clouds of the static parts and movable parts of the openable container based on the structural information differences, that is, use the structural differences to separate the point cloud data, ensuring the independence of the static parts and movable parts. In particular, movable parts such as drawers and doors are dynamically marked and processed through the interaction of the robot, thereby improving the accuracy and detail performance of the container 3D reconstruction.
[0108] Specifically, operate the robot to open or close the openable container, collect images and reconstruct the 3D model of the corresponding state through Step 2. Let the 3D point cloud of the container with the movable parts closed be P C , and the RGB image information collected in this coordinate system is RGB C ; the 3D point cloud of the container with the movable parts opened is P O , and the RGB image information collected in this coordinate system is RGB O . Then according to P C , RGB C , P O , RGB O The process of extracting the static parts and movable parts includes:
[0109] (1) Detect the RGB image C and the RGB O to find the movable parts, and build the correspondence between the movable parts in the RGB C and the RGB O based on the nearest neighbor search. Combining the correspondence between the RGB image and the 3D point cloud information, build the correspondence between the point clouds of the movable parts in the P C and the P O .
[0110] (2) Use the iterative closest point algorithm to calculate the pose transformation matrix between the P C and the P O , and transform the P C and the P O to the same coordinate system. In this way, the pose transformation matrix T c,o(s0) between the static parts in the two point clouds can be considered as the identity matrix, while the pose transformation matrix between the movable parts is obtained by using the iterative closest point algorithm.
[0111] (3) Based on the existing pose transformation matrix between the movable parts, further extract the structural information of the static parts and the movable parts in the P C and the P O .
[0112] If the 3D point p in the point cloud P C is a point in the movable part s, its transformation matrix can transform p to the corresponding point p' in the point cloud P O . To determine the best correspondence between p and the part, apply all the obtained rigid body transformation matrices to the point p and calculate its distance d O to the nearest point in the P s :
[0113] d s (p) = T s ·p - p'
[0114] When d s (p) < 0.02m, it is considered that the constructed correspondence is available, and the following confidence index is constructed:
[0115]
[0116] After calculating the confidence levels of point p corresponding to all other components, the component with the highest confidence level is the component corresponding to point p. Since some information exists differently when the movable component is in the open and closed states, they cannot establish available corresponding relationships. It is stipulated that the points in the three-dimensional information of the closed state of the movable component for which corresponding relationships cannot be established are all points of the static component, and the points in the three-dimensional information of the open state of the movable component for which corresponding relationships cannot be established are all points of the movable component. As Figure 4 shown.
[0117] Based on the above method content, taking the application in a domestic service robot platform as an example, this platform is equipped with the following hardware devices:
[0118] Depth camera or RGB-D camera: Used to capture the depth image and RGB image of the environment and obtain the three-dimensional information of the container. It is recommended to use Microsoft's Kinect sensor or Intel's RealSense camera, which has relatively high depth perception accuracy.
[0119] Robot arm and actuator: The robot is equipped with a robotic arm for performing tasks such as interacting with the container, such as opening drawers and placing items. The arm can be precisely adjusted in position through the controller.
[0120] Computing unit: Built-in embedded computing platform, such as NVIDIA Jetson platform, responsible for image processing, data calculation and execution of three-dimensional reconstruction algorithms.
[0121] The robot obtains the RGB-D image data of openable containers such as cabinets and drawers through the depth camera and transmits it to the computing unit. The computing unit processes the data according to the technical solution proposed by the present invention, generates a three-dimensional model of the container and executes the item sorting task.
[0122] Specifically, in actual operation, when the robot faces a cabinet, it uses a depth camera to capture the RGB-D image data of the cabinet, extracts the edge features of the cabinet door from the depth image of the first frame, then continues to extract the same features in the second frame image, and uses geometric constraints to ensure the accurate pose estimation of the opening and closing actions of the cabinet door. This process provides an accurate pose reference for subsequent 3D reconstruction. When the robot attempts to reconstruct the inside of a cabinet, the depth camera may have noise in the image due to light or object occlusion. Using gradient information, the present invention can filter out unreliable depth data during the 3D reconstruction process, thereby obtaining a more accurate 3D structure of the inside of the cabinet. When the robot opens the cabinet, by recording the opening and closing actions of the door and the static structure of the base, 3D models of the cabinet door and the cabinet body can be reconstructed respectively. During the reconstruction process, the motion state of the door is marked as a dynamic component, and the structure of the cabinet body is marked as a static part, so as to ensure the independence of the two in the 3D model. The robot can detect the positions of objects such as bottles and tableware through YOLOv7, and combine with the 3D reconstruction model to determine the specific positions of the items in space, so as to accurately perform the tasks of storing or taking out items.
[0123] In summary, the present invention proposes to use planar contour features to assist camera pose estimation to solve the problem of large pose estimation errors that may occur in the scene of an openable container composed of planes; use the Welsch function as the loss function for the distance between corresponding points in pose estimation to improve the robustness of dealing with outliers. Use a truncated signed distance function with gradient information to represent the 3D model, and combine with the imaging characteristics of the depth camera to filter out significantly incorrect depth information, thereby improving the reconstruction quality. Based on the interaction between the robot and the openable container, extract the static and dynamic structure information of the openable container, provide reliable environmental perception information for the robot, and significantly improve the reconstruction accuracy and the efficiency of robot operation.
[0124] The above is only the preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the implementation process of the present invention has been described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A three-dimensional reconstruction method for an openable container based on structural features and operational interactions, characterized in that, Including the following steps: Step 1: Obtain continuous frame depth images of the scene where the openable container is located, and extract the planar contour features of the depth images as geometric constraints for pose estimation of adjacent frames; Step 2: Represent the 3D model using the truncated signed distance function GSDF with gradient information, and fuse depth data to reconstruct the static 3D model of the openable container; Step 3: Dynamically analyze the opening and closing states of the openable container, distinguish the dynamic components and static matrix of the openable container, and mark the dynamic components to optimize its 3D reconstruction process.
2. The three-dimensional reconstruction method of the openable container according to claim 1, wherein In Step 1, the planar contour features are extracted from the depth images through the agglomerative hierarchical clustering algorithm.
3. The three-dimensional reconstruction method of the openable container according to claim 1, wherein In Step 1, the Welsch function is adopted as the loss function for the distance between corresponding points in pose estimation.
4. The three-dimensional reconstruction method of the openable container according to claim 1, characterized in that In Step 2, the truncated signed distance function GSDF with gradient information stores the gradient information of the point cloud data, and calculates the distance between each input point and the nearest point on the model surface using the first-order Taylor expansion method.
5. The three-dimensional reconstruction method of the openable container according to claim 4, characterized in that, In step 2, the depth data is preprocessed, specifically including: taking the k-th frame depth image R k As input, perform depth image denoising, normal vector calculation and plane contour extraction operations to obtain the denoised three-dimensional point cloud V k , the normal vector N of the point cloud k and the plane contour C in the depth image k .
6. The three-dimensional reconstruction method of the openable container according to claim 5, wherein, In step 2, the 3D model represented by GSDF is reconstructed by fusing depth data, specifically including: for the input k-th frame depth image R k , calculating its normal vector N k and relative pose T g,k , weight value W k (p) is always equal to 1, and the truncated signed distance value is calculated by the following equation: where \(q = \pi(p)\) represents the perspective projection process of the camera, which means finding the pixel point with the closest distance in the depth image; after obtaining the truncated signed distance value, the GSDF value is updated using the following formula:
7. The three-dimensional reconstruction method of the openable container according to claim 6, wherein The normal vector is calculated using the approximate least squares method, where it is assumed that the depth values of the target pixel point and adjacent pixel points in the depth image are approximately equal, and a box filter is used to accelerate the calculation.
8. The three-dimensional reconstruction method of the openable container according to claim 1, wherein, In Step 3, images of the openable container in the open or closed state are collected and the 3D models in the corresponding states are reconstructed, and the point clouds of the static components and movable components of the openable container are extracted based on the structural information differences.
9. The three-dimensional reconstruction method of the openable container according to claim 8, wherein, The extraction of the point clouds of the static components and movable components of the openable container based on the structural information differences specifically includes: Let the three-dimensional point cloud of the container with the movable part closed be P C , and the RGB image information collected in this coordinate system be RGB C , let the three-dimensional point cloud of the container with the movable part open be P O , and the RGB image information collected in this coordinate system be RGB O ; Detect the RGB image C and RGB O Detect the movable parts in it, and build the correspondence between the movable parts in RGB C and RGB O based on the nearest neighbor search; Combine the correspondence between the RGB image and the three-dimensional point cloud information to build the correspondence between the point clouds of the movable parts in P C and P O ; Calculate P using the Iterative Closest Point algorithm C and P O The pose transformation matrix between them, transform P C and P O to the same coordinate system, then the pose transformation matrix between the static components in the two point clouds can be considered as the identity matrix, while the pose transformation matrix between the movable components is obtained by solving using the Iterative Closest Point algorithm; Based on the existing pose transformation matrix between movable components, for P C and P O , further extract the structural information of the static and movable components; if the three-dimensional point p in the point cloud P C is a point in the movable component s, its transformation matrix can transform p to the corresponding point p' in the point cloud P O . To determine the best correspondence between p and the component, apply all the obtained rigid body transformation matrices to the point p and calculate the distance d O to the nearest point in P s (p): d s (p) = T s ·p - p' When d s (p) < 0.02 m, it is considered that the constructed corresponding relationship is available, and the following confidence index is constructed: After calculating the confidence of point p corresponding to all other components, the component with the highest confidence is the component corresponding to point p.
10. The three-dimensional reconstruction method of the openable container according to claim 9, wherein, The points that cannot establish a correspondence in the 3D information of the closed state of the movable component are all points of the static component, while the points that cannot establish a correspondence in the 3D information of the open state of the movable component are all points of the movable component.