Method for estimating poses of multiple targets, robot control method and system, and product
Patent Information
- Application Number
- PCT/CN2025/073614
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art is difficult to effectively estimate the position of multiple targets in complex environments, and cannot meet the needs of high-precision control of robots.
By acquiring real-time video data, segmenting pixel data, updating the pose using differential rendering and minimizing errors, real-time estimation of multi-objective poses is achieved.
High-precision real-time estimation of multi-objective poses is achieved, and the robot can be supported in complex environments.
Smart Images

Figure CN2025073614_12062025_PF_FP_ABST
Abstract
Description
Method for estimating multi-target poses, robot control method, system and product
[0001] Cross-references
[0002] This application claims priority to PCT patent application No. PCT / CN2024 / 073943, entitled “ROBOT CONTROL BASED ON REAL-TIME ENVIRONMENT DIGITAL TWIN RECONSTRUCTION,” filed on January 25, 2024. This PCT application No. PCT / CN2024 / 073943 is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates to the technical field of computer image processing, and in particular to a method for estimating multi-target poses, a robot control method, system, and product. Background Art
[0004] For example, US Patents 10565747B2 and 10565747B2 describe a method for estimating the pose of objects in static images using a particle filter and an autoencoder. The method also describes establishing a differentiable rendering pipeline for a given static two-dimensional image input. However, existing technologies lack methods for estimating the pose of multiple objects in complex environments, making them incapable of meeting the requirements for high-precision robot control. Summary of the Invention
[0005] The technical problem to be solved by the present disclosure is to overcome the above-mentioned defects in the prior art and provide a method for estimating multi-target poses, a robot control method, system and product.
[0006] The present disclosure solves the above technical problems through the following technical solutions:
[0007] The present disclosure provides a method for estimating multi-target poses, the method comprising:
[0008] Acquiring real-time video data, wherein the real-time video data includes first image frames of the multiple targets;
[0009] Based on the first image frame, obtaining first high-dimensional segmented pixel data, wherein the first high-dimensional segmented pixel data includes first pixel data of each object, and the first pixel data is a three-dimensional or higher-dimensional array;
[0010] Based on a first preset pose, performing differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array;
[0011] Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, the first preset pose is updated to obtain a first target pose.
[0012] Preferably, the real-time video data further includes at least a second image frame of the multiple targets;
[0013] After obtaining the target posture, the method further includes:
[0014] Based on the second image frame, obtaining second high-dimensional segmented pixel data;
[0015] Based on the first target pose, performing differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array;
[0016] Based on minimizing the error between the second high-dimensional segmented pixel data and the second high-dimensional data, updating the first target pose to obtain a second target pose;
[0017] Based on other image frames of the real-time video data, the posture corresponding to the previous image frame is updated in sequence to obtain other target postures.
[0018] Preferably, the step of obtaining first high-dimensional segmented pixel data based on the first image frame includes:
[0019] Segmenting each object in the first image frame using a segmentation algorithm to obtain a mask array;
[0020] The first high-dimensional segmented pixel data is obtained based on the mask array and the channel data of the first image frame.
[0021] Preferably, the step of performing differentiable rendering on the three-dimensional model of the multi-target based on the first preset pose to obtain first high-dimensional data includes:
[0022] Based on the preset poses of the multiple targets, transforming the three-dimensional models of the multiple targets from an object coordinate system to a camera coordinate system according to learnable pose transformation parameters, wherein the pose transformation parameters are used to represent the real-time poses of the targets;
[0023] Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model to obtain the first high-dimensional data.
[0024] Preferably, after the step of performing differentiable rendering on the three-dimensional model of the multi-object to obtain first high-dimensional data, the method further comprises:
[0025] Calculating occlusion perception data corresponding to occlusion information of the target using the depth information and contour information of the first high-dimensional data;
[0026] The step of updating the first preset pose to obtain a first target pose based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data includes:
[0027] The gradient descent method is used to iteratively minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable pose transformation parameters.
[0028] Preferably, the step of performing differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data includes:
[0029] The converted three-dimensional model is subjected to the differentiable rendering to obtain first high-dimensional data having a variable contour, wherein a range of the variable contour is determined based on a degree of difference between the first pixel data and a preset posture.
[0030] Preferably, performing the differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data includes:
[0031] Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain second RGBXY derivatives of the first high-dimensional segmented pixel data, where both the first RGBXY derivative and the second RGBXY derivative include information about color change and spatial position change.
[0032] Preferably, updating the first preset pose to obtain a first target pose based on minimizing an error between the first high-dimensional segmented pixel data and the first high-dimensional data includes:
[0033] Based on optimal transmission, minimize the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data, update the first preset pose to obtain the first target pose.
[0034] Preferably, the target includes a robot; the posture conversion parameters include joint angles of the robot;
[0035] The step of converting the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose conversion parameters based on the preset poses of the multiple targets includes:
[0036] Based on the preset posture and preset joint angles of the robot, the three-dimensional model of each component of the robot is converted from the object coordinate system of each component to the camera coordinate system by calculating the robot forward kinematics matrix;
[0037] The step of performing differentiable rendering on the converted three-dimensional model based on preset camera parameters to obtain the first high-dimensional data includes:
[0038] Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional robot model to obtain the first high-dimensional data.
[0039] Preferably, the preset posture includes preset translation parameters and preset rotation parameters of the multiple targets;
[0040] Before the step of transforming the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose transformation parameters, the method further includes:
[0041] Determine preset translation parameters using point cloud data of each target;
[0042] Utilizing preset image metrics and / or features and / or pre-trained models, calculating similarities between templates in the template set and a reference image to determine the most matching target template, and determining preset rotation parameters based on the rotation parameters corresponding to the target template;
[0043] The template set includes rotation parameters and corresponding images, and is generated by performing a plurality of uniform samplings on a selected rotation axis.
[0044] Preferably, before the step of calculating the similarity between the templates in the template set and the reference image to determine the most matching target template, the method further comprises:
[0045] Performing several uniform samplings on the selected rotation axis to generate a preset number of sampling data;
[0046] Based on preset camera parameters, a three-dimensional model of the target, initial translation parameters and a preset number of sampling data, a set of images of different poses of the three-dimensional model is rendered and generated to constitute the template set.
[0047] Preferably, before the step of performing differentiable rendering on the three-dimensional model of the multi-object to obtain the first high-dimensional data, the method further comprises:
[0048] Acquire a plurality of replica models of a three-dimensional model of a batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters;
[0049] Calculating occlusion perception data corresponding to occlusion information of the batch optimized objects using high-dimensional data obtained by rendering the replica models of the batch optimized objects;
[0050] The loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object is calculated to find the target replica model with the minimum loss value. The target replica model is used to determine the preset pose for this round of pose estimation.
[0051] Preferably, the step of acquiring real-time video data includes:
[0052] Use multiple cameras to capture images from different angles to reduce pose ambiguity caused by mutual occlusion and self-occlusion of the target.
[0053] The real-time target detection data of each camera are matched with each other by using three-dimensional point cloud reprojection to obtain the real-time video data.
[0054] Preferably, the step of using three-dimensional point cloud reprojection to match the real-time target detection data of each camera includes:
[0055] Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects;
[0056] Determining camera parameters, where the camera parameters include intrinsic parameters of the first camera and the second camera, and extrinsic parameters for transforming a field of view of the first camera and a field of view of the second camera;
[0057] Based on the determined camera parameters, reprojecting the depth image or three-dimensional point cloud data of the first camera into the field of view of the second camera to obtain a corresponding reprojected two-dimensional point group;
[0058] comparing the mask information of the one or more objects in the image of the second camera with the reprojected two-dimensional point group to determine which object's mask information has the greatest overlap with the reprojected two-dimensional point group;
[0059] Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
[0060] The present disclosure also provides a method for controlling a robot, wherein:
[0061] Obtaining the pose information of the robot and the pose information of the target object using the method for estimating the multi-target pose as described above;
[0062] Inputting the posture information of the robot and the posture information of the target object into the control system of the robot;
[0063] Based on the input, the control system controls the interaction of the robot with the target object.
[0064] Those skilled in the art will appreciate that the robot's control system, upon obtaining the robot's pose information and the pose information of the target object, can control the robot to achieve precise movement based on this input pose information. For example, when grasping a target object, this input pose information enables the control system to calculate the optimal path and movement strategy to minimize the risk of collision or slippage.
[0065] The present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for estimating multi-target poses or the method for controlling a robot as described above.
[0066] The present disclosure also provides a system for estimating multi-target poses, the system comprising:
[0067] One or more processors, wherein the one or more processors implement the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0068] Preferably, the system further comprises one or more cameras for acquiring video data and / or image data of the estimated target.
[0069] Preferably, the system further comprises a robot;
[0070] The system obtains the pose information of the robot and the pose information of the target object by using the method for estimating the multi-target poses as described above;
[0071] The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
[0072] The present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for estimating multi-target poses or the method for controlling a robot as described above is implemented.
[0073] The present disclosure also provides a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, implement the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0074] The present disclosure also provides a computer program, which, when executed by a processor, can execute the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0075] The positive progress of this disclosure is:
[0076] The method for estimating the pose of multiple targets provided by the present invention takes a real-time video stream as input and processes each latest image frame of the real-time video stream in sequence to achieve real-time dynamic tracking; by performing image segmentation on each image frame to identify each target therein, a high-dimensional segmented pixel array including the pixel data of each target is obtained, and the high-dimensional array is compared with the high-dimensional array obtained by differentiable rendering of known three-dimensional models of the multiple targets, and the pose is estimated by minimizing the error, thereby achieving pose estimation of multiple targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. For those skilled in the art, it is possible to apply this specification to other similar scenarios based on these drawings without inventive effort.
[0078] FIG1 is a first flow chart of a method for estimating multi-target poses provided in Example 1 of the present disclosure;
[0079] FIG2 is a second flow chart of a method for estimating multi-target poses provided in Example 1 of the present disclosure;
[0080] FIG3 is a schematic diagram of static pose calibration for object pose estimation provided in Example 1 of the present disclosure.
[0081] FIG4 is a schematic diagram of real-time pose tracking for object pose estimation provided in Example 1 of the present disclosure.
[0082] FIG5 is a schematic diagram of static pose calibration of a robot and an object pose estimation provided in Example 1 of the present disclosure.
[0083] FIG6 is a schematic diagram of static pose calibration of a robot and an object pose estimation provided in Example 1 of the present disclosure.
[0084] FIG7 is a schematic diagram of a process for retrieving rough initial rotation parameters provided in Example 1 of the present disclosure.
[0085] FIG8 is a visualization result of different initial poses with / without initial translation parameters and initial rotation parameters at the beginning of pose estimation provided by Example 1 of the present disclosure.
[0086] FIG9 is a schematic diagram of the large-angle rotation problem in the optimization process provided in Example 1 of the present disclosure.
[0087] FIG10 is a visualization of an example of a teapot grid replica used in the batch optimization strategy provided in Example 1 of the present disclosure.
[0088] 11a-c are schematic diagrams of the configuration of multiple cameras and detection results at different viewing angles provided in Example 1 of the present disclosure.
[0089] 12a-c are schematic diagrams of reprojected 3D point clouds from camera view 1 to view 2 provided in embodiment 1 of the present disclosure.
[0090] FIG13 is a schematic diagram showing the effect of the interaction between the robot and the object during the grasping operation provided in Example 1 of the present disclosure.
[0091] FIG14 is a schematic structural diagram of an electronic device provided in Example 4 of the present disclosure. DETAILED DESCRIPTION
[0092] The present disclosure is further illustrated below by way of examples, but the present disclosure is not limited to the scope of the examples.
[0093] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places herein does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0094] It should be understood that the terms "device," "system," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0095] As used herein, unless the context clearly indicates otherwise, the terms "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include additional steps or elements.
[0096] The definition of inclusion herein, such as the terms “having”, “may have”, “include” or “may include” as used herein, indicates the existence of the corresponding functions, operations, elements, etc. herein, and does not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that the terms “including” or “having” as used herein indicate the existence of the features, numbers, steps, operations, elements, components or their combination described in the specification, and do not exclude the existence or addition of one or more other features, numbers, steps, operations, elements, components or their combination.
[0097] Flowcharts are used herein to illustrate the operations performed by the systems according to the embodiments of the present invention. It should be understood that the preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0098] Example 1
[0099] Please refer to Figure 1, which is a first flow chart of the method for estimating multiple target poses in this embodiment. Specifically, as shown in Figure 1, the method includes:
[0100] S11, acquiring real-time video data, where the real-time video data includes first image frames of multiple targets; by taking the real-time video stream as input, the real-time video stream can be divided into a plurality of real-time image frames;
[0101] S12. Based on the first image frame, obtain first high-dimensional segmented pixel data, the first high-dimensional segmented pixel data including first pixel data of each target, the first pixel data being a three-dimensional or higher-dimensional array; image segmentation can be performed on each image frame to identify each target therein, thereby obtaining a high-dimensional segmented pixel array including pixel data of each target.
[0102] S13. Based on the first preset pose, perform differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array;
[0103] S14. Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, update the first preset posture to obtain a first target posture.
[0104] After processing the first image frame, by sequentially processing each image frame of the real-time video stream, dynamic real-time estimation of the poses of multiple targets can be achieved. Therefore, the real-time video data also includes at least the second image frames of the multiple targets;
[0105] Please refer to FIG2 , which is a second flow chart of the method for estimating multiple target poses in this embodiment. Specifically, as shown in FIG2 , after step S14 , the method further includes:
[0106] S15. Obtaining second high-dimensional segmented pixel data based on the second image frame;
[0107] S16. Based on the first target pose, perform differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array;
[0108] S17. Based on minimizing the error between the second high-dimensional segmented pixel data and the second high-dimensional data, updating the first target pose to obtain a second target pose;
[0109] S18. Based on other image frames of the real-time video data, update the posture corresponding to the previous image frame to obtain other target postures.
[0110] In some examples, the following method steps can be implemented by a program or a custom circuit, or a combination of a custom circuit and a program. For example, the following method steps can be performed by a combination of one or more GPUs, one or more CPUs, or any other technically feasible one or more processors such as an ASIC chip. In some examples, one or more processors such as an ASIC chip can include a memory. Further, those skilled in the art will understand that the scope of protection of the present disclosure can include any system and device that can perform the following method steps.
[0111] In an optional embodiment, step S12 includes:
[0112] S121, using a segmentation algorithm to segment each object in the first image frame to obtain a mask array;
[0113] S122 . Obtain first high-dimensional segmented pixel data based on the mask array and the channel data of the first image frame.
[0114] Image segmentation
[0115] The inventors have found that image segmentation algorithms involve dividing an image (RGB or RGB depth) into multiple fragments (sets of pixels) to simplify the representation of the image into more meaningful and easier to analyze content. Typically, image segmentation algorithms assign a label to each pixel in the image, and pixels with the same label share certain features (color, edge, texture, category, object instance, etc.). According to different classification criteria, image segmentation algorithms can be divided into categories such as semantic segmentation, instance segmentation, and panoramic segmentation. Among them, in some embodiments of the present disclosure, for example, an instance segmentation algorithm can be used, in which each fragment of the image uniquely represents a specific instance of an object. This method not only classifies the pixels of the image into different object classes, but also distinguishes different instances of the same class. For example, in a street scene, instance segmentation will identify and separate each individual car, pedestrian, or any other important object, treating each object as a unique entity, even though they may belong to the same category.
[0116] In another optional embodiment, step S13 includes:
[0117] S131. Based on the preset poses of multiple targets, the three-dimensional models of the multiple targets are transformed from the object coordinate system to the camera coordinate system according to the learnable pose transformation parameters. The pose transformation parameters are used to characterize the real-time pose of the targets. Object pose estimation is the process of determining the position and orientation (6D pose) of an object in space, which can be achieved by processing image data.
[0118] S132. Based on preset camera parameters, perform differentiable rendering on the converted 3D model to obtain first high-dimensional data. Unlike traditional non-differentiable rendering, differentiable rendering can calculate the gradient of the 3D target and propagate it through the image, bridging the gap between 2D and 3D image processing. The gradient is back-propagated based on the rendering output to optimize the 3D scene parameters (such as camera pose, 3D model pose or shape, and lighting conditions).
[0119] Specifically, taking the image frame as the input image, it can be expressed as The input image can be an RGB image or RGBD image containing N targets; the segmentation algorithm is used to segment the input image and output a mask array Then, combined with the channel data of the image frame, a high-dimensional array can be obtained. That is, high-dimensional segmented pixel data, where C represents the channel data of each image frame (such as contour, RGB and depth channels). ⊙ represents element-wise multiplication, which applies the mask to the input image I, which involves the array multiplication broadcast mechanism, that is, Repeat N times to get Then with I mask Perform the ⊙ operation, and then get the shape of A as 3(4) corresponds to the dimension I mask Superposition is performed to obtain the dimension corresponding to C.
[0120] The task of 6D object pose estimation is to estimate the rigid transformation of all targets. The rigid transformation can be expressed as Rigid transformation is used to transform the 3D object model M = {M n |n∈{1,2,…N}} maps from the object coordinate system to the camera coordinate system. Assume that the 3D object model M and the camera intrinsic parameters K are known. Each model M n Can be defined as A set of vertices in and a set of polygons describing the surface of the object. n A 4×4 rigid transformation matrix P n= [R, T; 0, 1], where R is a 3×3 rotation matrix and T is a 3×1 translation vector. The translation T specifies the origin of the object coordinates in the camera coordinate system. To ensure that R is physically valid, the rotation matrix R is represented by the quaternion q, which can be further converted into a 3×3 rotation matrix.
[0121] Then, the pose estimation process is as follows: For a given initial pose P′ and object model M, the object model can be transformed to obtain M P′ ={P′ n ·M n |n∈{1,2,…N}}, where P′ n is the initial pose of the nth object, M n is the 3D model of the nth object. Using a differentiable renderer And the camera intrinsic parameter K, the first high-dimensional data can be rendered Where H' and W' represent the height and width of the rendered image, respectively, and C represents the number of channels (such as alpha, RGB, depth). It should match the structure of the reference array A, which is the high-dimensional array generated by the segmentation algorithm and encodes the class identification information.
[0122] In an optional implementation, step S132 specifically includes:
[0123] The converted three-dimensional model is subjected to differentiable rendering to obtain first high-dimensional data having a variable contour, wherein the range of the variable contour is determined based on the degree of difference between the first pixel data and a preset posture.
[0124] In another optional implementation, step S132 specifically includes:
[0125] Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain second RGBXY derivatives of the first high-dimensional segmented pixel data, where both the first RGBXY derivative and the second RGBXY derivative include information about color change and spatial position change.
[0126] Furthermore, step S14 may include:
[0127] Based on the optimal transmission, the pixel level difference between the second RGBXY derivative of the first high-dimensional segmented pixel data and the first RGBXY derivative of the first high-dimensional data is minimized; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmented pixel data and the first RGBXY derivative of the first high-dimensional data, the first preset posture is updated to obtain the first target posture.
[0128] Differentiable Rendering
[0129] The inventors have discovered that, in some embodiments of the present disclosure, differentiable rendering can be used to calculate the gradients of a 3D target object and propagate them across images. This differentiable rendering approach bridges 2D and 3D image processing methods. By backpropagating the gradients of the rendered output, it can optimize 3D scene data (e.g., camera pose, 3D model pose or shape, lighting conditions, etc.).
[0130] The document Liu, S., Li, T., Chen, W., & Li, H. Soft rasterizer: A differentiable renderer for image-based 3d reasoning. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) (pp. 7708-7717). (2019) records the loss calculation by rendering a single real image of a geometric primitive, and all the disclosed contents are incorporated into the present disclosure. In some embodiments of the present invention, the smooth edge contour tensor obtained by rendering has a variable smooth edge range, and the smooth edge range can be adjusted to be larger or smaller according to the difference between the real object pose and the virtual digital twin pose. When there is a large difference between the real object pose and the virtual digital twin pose, a larger smooth range can be used to allow the gradient flow of the digital twin when it is away from the real target object. When the difference between the two is small, a smaller smooth edge can be used to perform a higher precision pose estimation.
[0131] Occlusion in Rendered Array
[0132] The inventors found that since the reference high-dimensional array A only captures the visual information after occlusion, the high-dimensional rendering model Contains complete visual information of each target and does not consider the occlusion between objects. Direct comparison will lead to pose estimation errors. Therefore, it is necessary to render the model in a high-dimensional After step S13, the method further includes:
[0133] S21, using the depth information and contour information of the first high-dimensional data to calculate the occlusion perception data corresponding to the occlusion information of the target; high-dimensional rendering model The occlusion information can be calculated by using the z-buffer (depth buffer) to determine the visibility of objects at each pixel. A set of binary object masks can be used in the process. and the corresponding depth map Where n∈N is the index of the object in the array. Specifically, in order to treat non-object areas as far away, a large constant δ is added to the pixels of the depth map to perform depth adjustment: D′(x,y,n)=D(x,y,n)+(1-Mask(x,y,n))×δ
[0134] For each pixel (x,y), the frontmost object is determined by finding the object with the smallest depth among all objects for visibility measurement:
[0135] The occlusion mask calculation formula for each pixel of each object is: F(x,y,n)=1[γ(x,y)=n]
[0136] Here, 1 is an indicator function that returns 1 if the condition is true, otherwise it returns 0.
[0137] Multiplying the occlusion mask with the initial object mask gives the final occlusion mask for each object: O(x,y,n)=F n (x,y)×Mask(x,y,n)
[0138] Using the mask O(x,y,n), we can calculate an array Taking into account
[0139] RGBXY Derivatives
[0140] The inventors discovered that RGBXY derivatives are a specific type of gradient used in differentiable rendering that accounts for both color (RGB) and spatial (XY) variations in an image. Unlike traditional per-pixel derivatives that are calculated based solely on color differences, RGBXY derivatives also incorporate spatial variations, thereby improving the robustness of gradient-based optimization. These derivatives can more accurately track object transformations (such as position and rotation) even when there are significant differences between the initial and target poses.
[0141] Step S14 includes:
[0142] S141. Iterate using the gradient descent method to minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable pose transformation parameters.
[0143] The pose transformation parameter P is set as a learnable parameter, and the object pose is estimated by minimizing the error between the real high-dimensional array and the occlusion-aware rendering array. The optimization goal is:
[0144] Among them, Loss is a metric used to evaluate the difference between each object in the captured image I and its corresponding digital twin in the rendered image. The loss function is defined as a weighted combination of the error of each modality (such as mean squared error (MSE), mean absolute error, intersection over union (IoU), DICE coefficient, etc.) and the optimal transmission error of the RGBXY points of the occlusion-aware image: Loss = λ c L c +λ d L d +λ mask L mask +λ ot L ot
[0145] Among them, A * (y,x,n) represents the pixel value of the mode* of A, which can be RGB color, depth and contour occlusion; Q i (n) represents point i in a set of RGBXY points (including only points with non-zero RGB color values, where XY represents spatial coordinates) in a reference image taken from target n; represents a point σ(i) in a set of RGBXY points in the occlusion-adjusted rendered image of target n; the matching function σ(·) defines a one-to-one correspondence between points in the captured reference image and points in the occlusion-adjusted rendered image, minimizing the total distance between the matched images; λ c ,λ d ,λ mask andλ ot Represents the weight of the balance loss. The above loss function is fully differentiable and the gradient descent method can be used to optimize the pose transformation parameter P.
[0146] Optimal Transport
[0147] The inventors discovered that using a mathematical framework called optimal transport, they can find the most efficient way to transform one probability distribution into another by minimizing a cost function. In computer vision and rendering, optimal transport is used to match pixels or features between two images (such as a rendered image and a reference image) to ensure that corresponding features are correctly aligned, even when there are significant differences in the positioning of the objects.
[0148] The document Xing, J., Luan, F., Yan, LQ, Hu, X., Qian, H., & Xu, K. Differentiable rendering using RGBXY derivatives and optimal transport. ACM Transactions on Graphics (TOG), 41(6), 1-13. (2022) records an innovative method for differentiable rendering using RGBXY derivatives and optimal transport, which solves the significant limitations of traditional pixel derivative methods, and the entire disclosure thereof is incorporated into the present disclosure. It utilizes 5DRGBXY derivatives to capture changes in color and spatial position, improving the robustness of the optimization task, especially for remote object transformations. The inventors found that some embodiments of the present disclosure can be based on this technology to integrate RGBXY derivatives and optimal transport into the real-time differentiable rendering system of this embodiment. By doing so, significant advantages are obtained in processing dynamic environments with continuous video input because it improves the accuracy of object pose estimation. RGBXY derivatives take into account color and spatial changes, providing more robust gradient-based optimization for large object transformations. Optimal transfer minimizes a cost function that quantifies the difference between the RGBXY features of the rendered and target images, ensuring efficient alignment. By minimizing this function, optimal transfer finds the most effective way to match features even when there are significant differences in object positioning. The combination of RGBXY optimal transfer improves silhouette matching across frames, making it more robust in dynamic multi-object environments.
[0149] The inventors discovered that the above-mentioned target pose estimation process can be implemented using differential rendering libraries such as SoftRas, Pytorch3D, and nvdiffrast. The document Liu, S., Li, T., Chen, W., & Li, H. Soft rasterizer: A differentiable renderer for image-based 3D reasoning. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) (pp. 7708-7717). (2019) describes the use of the SoftRas differential rendering library to implement pose estimation, and the entire disclosure thereof is incorporated into this disclosure. Reference Ravi, N., Reizenstein, J., Novotny, D., Gordon, T., Lo, WY, Johnson, J., & Gkioxari, G. Accelerating 3d deep learning with pytorch3d.arXiv:2007.08501 (2020) records the use of Pytorch3D differential rendering library to implement pose estimation, and all the disclosed contents are incorporated into this disclosure. References Laine, S., Hellsten, J., Karras, T., Seol, Y., Lehtinen, J., & Aila, T. Modular primitives for high-performance differentiable rendering. ACM Transactions on Graphics (TOG), 39(6), 1-14. (2020). and J. Tremblay, B. Wen, V. Blukis, B. Sundaralingam, S. Tyree, and S. Birchfield. Diff-DOPE: Differentiable Deep Object Pose Estimation. arXiv: 2310.00463 (2023) record the use of the nvdiffrast differential rendering library to implement pose estimation, and all the disclosed contents are incorporated into this disclosure.
[0150] Compared to the aforementioned rendering techniques, which focus on estimating 6D object poses from 2D static reference images, some embodiments of the present disclosure can address the problem of real-time dynamic tracking of the poses of multiple target 6D objects. Therefore, the methods of some embodiments of the present disclosure can have two processing stages: 1) static pose calibration of the initial digital twin model; and 2) dynamic pose tracking of the digital twin model. The second stage is achieved by continuously optimizing the aforementioned loss function by continuously updating the image I using a real camera, so that the digital twin of the rendered image follows the target of the captured image in 3D space.
[0151] The following examples use 3D-printed teapots and teacups, depth cameras, and differential rendering libraries. The method of this embodiment achieves real-time tracking of teapots and teacups at a speed of approximately 60 FPS. Figures 3 and 4 show the optimization process of the 6D pose of the digital twin in the first and second stages. Figure 3 is a static calibration process of the pose of the digital twin. The leftmost figure illustrates the initial pose of the real teapot, teacup, and its digital twin (the diagonal area). The model of the digital twin is rendered by a known three-dimensional model, camera parameters, and the current estimated pose. From left to right, the method of this embodiment can calibrate the pose of the digital twin to match the real target. Figure 4 is a real-time pose tracking process of the digital twin. From left to right, while changing the physical pose of the real teapot and the real teacup, the method of this embodiment can continuously track the pose of the target and achieve real-time digital twin reconstruction.
[0152] The target of this embodiment may also include a robot; the posture conversion parameter includes the joint angle of the robot; step S131 may specifically include:
[0153] Based on the preset posture and preset joint angles of the robot, the three-dimensional model of each component of the robot is converted from the object coordinate system of each component to the camera coordinate system by calculating the robot's forward kinematics matrix;
[0154] Step S132 may specifically include:
[0155] Based on the preset camera parameters, the converted robot three-dimensional model is subjected to differentiable rendering to obtain the first high-dimensional data.
[0156] Given an input image I (RGB or RGBD), the goal of robot pose estimation is to recover the state of a robot with known joints in a 3D scene. The state of an articulated robot is defined by two key components: (i) the 6D pose P of the robot base base , i.e., the 3D translation and 3D rotation relative to the origin of the camera coordinates, and (ii) the joint angle θ of the robot joints. For an articulated robot with a known structure, its structure is usually represented by links and joint angles, where the joint angles determine the pose of each robot link relative to the origin coordinates of the robot base. By using the 6D pose P of the basebase and joint angles θ, the mesh of each robot link can be transformed from its initial pose to appropriate poses to represent the robot in various configurations.
[0157] This embodiment uses an optimization method similar to that used for object pose estimation, but takes into account the joint angle θ. Given the initial 6D pose P′ of the robot base base and the initial joint angle θ′ and the robot model M with N links robot ={l n |n∈{1,2,…N}}, first calculate the forward kinematics matrix of the initial joint angle θ′ of each link Then, the transformed robot model can be obtained That is, there is a configuration P′ base and a ready-to-render robot mesh with preset joint angles θ′. Depending on the type of robot segmentation algorithm used, a robot mask can be obtained, which can be a single-channel mask combining all links It can also be a multi-channel mask where each channel represents a separate link This selection generates a different reference array A=I⊙I mask . Using a differentiable renderer based on the shape of A and the camera intrinsic parameters K to render a high-dimensional model Where H′, W′ and C represent the height, width and channels (such as RGB, depth) of the rendered image, respectively.
[0158] Forward Kinematics
[0159] Forward kinematics is a sequence of physical transformations that describes the physical transformations available to each component in the robot. The physical pose of a robot component is determined not only by the joint angles or translation phases (relative pose to its neighbors) of its immediate neighbors, but also by the relative pose of that neighbor with respect to its own neighbors. Therefore, it is more intuitive to describe the physical transformation of a robot component as a composite transformation of this sequence of transformations. In computer graphics, this composite transformation can be simply described as the multiplication of the component's geometric point cloud coordinate grid by a 4x4 matrix. Typically, a unique 4x4 matrix is provided for each independent component of the robot to describe the complete forward kinematics.
[0160] Then, the above-mentioned pose estimation method can be used to calculate the corresponding occlusion perception data In addition, the forward kinematics of the robot is differentiable. Therefore, we can use the parameter P baseand joint angles θ are set as learnable parameters. By minimizing the error between the true array and the block perception data, the correct robot 6D pose and / or joint angles can be estimated:
[0161] Among them, Loss is the same indicator used in the object pose estimation part.
[0162] Similar to the multi-target pose estimation mentioned above, robot pose estimation can optimize multiple targets or robot arms (or both simultaneously) and also consists of two stages: static pose calibration and real-time pose tracking. Therefore, given an input image of a robot arm and a target object containing a known 3D model, the robot and object pose estimation method of this embodiment can first calibrate their corresponding digital twins to the correct pose and continuously track their poses in real time as the robot arm and target object move. The following is an example using a differential rendering library with a Universal Robots UR5 model and a teapot. First, a series of images are rendered as the poses of the target robot and the target object. Then, a robot digital twin model is constructed based on the URDF file of the UR5 robot arm, and an attempt is made to optimize its pose using differentiable rendering to match the pre-rendered robot image and the teapot.
[0163] Unified Robotics Description Format (URDF)
[0164] The Unified Robot Description Format (URDF) is an example of a data format that can be used to describe the physical geometry of a robot and its components (in the form of a point cloud mesh), as well as the degrees of freedom available on the robot to implement its physical structure transformations (such as joint angle rotations and physical translations). The mesh components in a URDF file may include a visual mesh (a large number of triangle mesh polygons used for visualization) and a collision mesh (a small number of triangle mesh polygons used for efficient collision detection).
[0165] Figures 5 and 6 show the calibration and tracking process of the pose of the digital twins of the robot and teapot. Figure 5 shows the static calibration process of the pose of the digital twins of the robot and teapot. The leftmost figure shows the initial pose of the robot and teapot (color) and their digital twin (white ghost). The model of the digital twin is rendered by a known three-dimensional model, camera parameters and the current estimated pose. From left to right, the method of this embodiment can calibrate the pose of the digital twin to match the real target. Figure 6 shows the real-time pose tracking process of the digital twin. From left to right, the physical poses of the robot and the teacup are changed at the same time. The method of this embodiment can continuously track the pose of the target and realize real-time digital twin reconstruction.
[0166] The pose estimation method of this embodiment estimates the poses of multiple objects and robots in an iterative manner. The iterative nature of such algorithms typically involves refining the setting of the initial pose through multiple optimization steps to achieve convergence with the true pose. However, if the initial pose is too far from the true pose, the optimization process may be limited by local minima or take longer to converge, resulting in inaccurate pose estimation and high system latency. On the other hand, selecting a suitable initial pose can significantly accelerate convergence, reduce the computational burden, and improve overall estimation accuracy. Therefore, selecting a suitable initial pose is crucial to ensuring the accuracy and efficiency of the estimation process.
[0167] In an optional embodiment, the preset pose includes preset translation parameters and preset rotation parameters of the multiple targets; before step S12, the method further includes:
[0168] S31, determining preset translation parameters using the point cloud data of each target;
[0169] S32. Using preset image indicators and / or features and / or pre-trained models, calculate the similarity between the templates in the template set and the reference image to determine the most matching target template, and determine the preset rotation parameters according to the rotation parameters corresponding to the target template; the template set includes the rotation parameters and the corresponding images, and the template set is generated by performing several uniform samplings on the selected rotation axis. References Nguyen, VN, Hu, Y., Xiao, Y., Salzmann, M., & Lepetit, V. Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (CVPR) (pp. 6771-6780). (2022) describes a pose estimation method based on template matching, and all the contents disclosed therein are incorporated into the present disclosure. Compared to generating a template by positioning a camera on a hemispherical grid and fixing the CAD model of the object at the center of the sphere, the method of some embodiments of the present disclosure can use real internal parameters and external parameters to form a fixed virtual camera and rotate the CAD model of the object in three-dimensional space using a uniformly sampled rotation matrix.
[0170] Specifically, before step S32, the method further includes:
[0171] Performing uniform sampling several times on the selected rotation axis to generate a preset number of sampling data;
[0172] Based on the preset camera parameters, the target's three-dimensional model, the initial translation parameters, and the preset number of sampling data, a set of images of different poses of the three-dimensional model is rendered to form a template set.
[0173] Specifically, for a given input image I (RGBD) and a CAD model M of the target, the goal of roughly estimating the initial pose is to provide a better starting point P close to the global minimum P. i , to avoid being limited by local minima and reduce the time cost of the optimization process. The rough initial object pose is composed of a 4×4 rigid transformation matrix P i =[R i ,T i ; 0, 1] means, where R i is a 3×3 rotation matrix, T i is a 3×1 translation vector. i Represented as quaternion q i , which can be easily converted into a 3×3 rotation matrix.
[0174] Initial translation parameter T i It can be obtained from the depth map D of a single camera view: first, the depth map is back-projected onto the point cloud in the camera coordinate system. The depth map is represented as D(u,v), where (u,v) is the pixel coordinate and D(u,v) is the depth value (distance from the camera). The intrinsic parameters of the camera are represented by the camera matrix K, which contains the focal length f x ,f y and the principal point (c x ,c y ). D(u,v) can be converted into 3D coordinates (x,y,z) using the following formula: z=D(u,v)
[0175] By applying this transformation to each pixel in the depth map, a complete point cloud can be obtained. The 3D point cloud of the corresponding object is extracted using the object segmentation mask. The rough initial translation parameter T i This can be selected from the target's 3D point cloud, either as a specific point or as the average position of all points in the point cloud, thus providing a reasonable estimate of the object's initial translation parameters.
[0176] In order to obtain an initial rotation parameter R close to the true value i , this embodiment adopts a similarity matching method between a pre-generated template set and a reference image. The template set consists of rendering target images with different rotation poses, where the rotation quaternion q b As the key value, the corresponding rendered image Rb As numerical values. Figure 7 illustrates how to retrieve the rough initial rotation parameters. The following is a detailed description of how to generate the target template mask and match it with the target reference mask:
[0177] First, use Euler angles to uniformly sample in three-dimensional space to obtain a set of rotation quaternions q b , where the number of variations along the x-axis, y-axis, and z-axis is denoted as var i , i∈{x,y,z}. These variants are uniformly distributed along each axis in the range [0, 360) degrees, and the sampling interval Δθ is expressed as where i∈{x,y,z}. Create the Cartesian product of these rotation angles to form a comprehensive 3D mesh containing all possible rotation combinations. The Euler angles are converted to a rotation matrix according to the following formula: R = R z (γ)R y (β)R x (α)
[0178] Among them, R x ,R y and R z are the rotation matrices about the x-axis, y-axis, and z-axis respectively, and α, β, and γ represent the corresponding Euler angles respectively.
[0179] These matrices are then converted to quaternions q b The total number of possible rotations is J = var x ×var y ×var z By incorporating prior knowledge about the geometry of the target object, different var i Specifically, for objects with poor directional symmetry, it is necessary to set a larger number of variants along the corresponding axis. As a rule of thumb, for objects lacking symmetry along a particular coordinate axis, the number of variants along that axis is typically between 6 and 12, depending on the shape and volume of the object's unique geometric features.
[0180] Figure 7 is a schematic diagram of the process of retrieving rough initial rotation parameters, which includes four steps: obtaining reference mask, generating template, feature extraction and similarity matching. Based on the intrinsic parameters of the camera, the CAD model of the target object, and the initial translation parameters and the rotation list Generate a set of images of the CAD model in different poses using existing rendering tools. b It can include masks, RGB images, or RGB-D images with corresponding masks. A template set is constructed based on the above data, where the rotation quaternion q b As the key value, the corresponding rendered image R bAs a numerical value. Both the rendered template and the reference mask are cropped around the object center and resized for better matching.
[0181] The similarity between the rendered template and the reference mask can be calculated using various metrics, such as traditional image similarity metrics (such as mean square error (MSE), intersection over union (IoU), DICE coefficient), hand-crafted features (such as SIFT, HOG) or learning-based models (such as ResNet, VGG). In the method of this embodiment, a pre-trained VGG network (without the classification layer) is used to obtain the image from R b and R i These feature maps are represented as F b and F r , and evaluate the cosine similarity between them by comparison to determine the highest similarity at the feature level. The goal of the rough initial rotation parameter estimation is to find the best match R with the reference mask r Template image R i Then, the rough initial rotation parameters are selected as the rotation quaternion q i FIG8 shows the visualization results of different initial poses with / without initial translation parameters and initial rotation parameters at the beginning of pose estimation, demonstrating the effectiveness of the initial 6D pose estimation strategy of this embodiment.
[0182] It should be noted that for multiple cameras, feature rendering and matching can be performed separately, but the sum of the highest similarity values is used to determine the shared coarse initial rotation parameter input. At the same time, when the camera is fixed, the camera parameters are usually unchanged, so F can be obtained in a one-time offline preprocessing process. b .
[0183] In addition, during the iterative optimization process, when the actual object moves quickly and the current object pose Far from reaching the expected global minimum P, the gradient descent method may lead to Stuck in local minima. One of the prominent problems is large-angle rotation, which may cause optimization failure for objects that lack geometric symmetry. Figure 9 is a schematic diagram of the large-angle rotation problem during the optimization process. The leftmost image is the reference image used for pose estimation. The right figure shows the relationship between the rotation angle and the loss value when the translation parameters are perfect. The white ghost shows the pose of the digital twin. From the initial pose (0) to the optimal pose (182), multiple local minimum traps will appear in the optimization process, and the gradient descent method is difficult to solve this problem. To solve this problem, this embodiment uses a batch optimization strategy, that is, optimizing a batch of mesh copies with different initial poses instead of a single mesh.
[0184] Therefore, before step S13, the method further includes:
[0185] S41, obtaining a plurality of replica models of the three-dimensional model of the batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters;
[0186] S42, calculating occlusion perception data corresponding to occlusion information of the batch optimized objects using high-dimensional data obtained by rendering the replica models of the batch optimized objects;
[0187] S43. Calculate the loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object to find the target replica model with the minimum loss value. The target replica model is used to determine the preset pose for this round of pose estimation.
[0188] For each 3D object model in batch optimization mode, it is copied to the original mesh M B To create a batch of mesh copies Where J is the batch size of each 3D model. Each replica is assigned a corresponding initial pose compensation It comes from the collection Set P′ Δ Including translation compensation matrix and the rotation compensation matrix (also expressed as quaternion ).parameter and It can be a fixed constant or a learnable parameter based on different configurations. Typically, the translation compensation parameter Initialize at the geometric center of the object or the origin of the object coordinate system, and rotate the compensation parameters Similar to the uniform sampling from the three-dimensional space introduced in the template generation section, it is also possible to select from the top K poses in the template set that are most similar to the current reference image. In 6D object pose estimation, since rotation estimation is more challenging than translation estimation, the batch optimization strategy is mainly to improve the accuracy of rotation estimation. The role of translation compensation is to help the rotation batch optimization move in the local space and promote optimization convergence. Based on the different learnability of translation and rotation offsets, this embodiment provides P′ Δ Three different configurations:
[0189] 1) Shared translation, fixed rotation: For all batch sizes J, all replicas have the same learnable translation compensation The rotation compensation remains unchanged. The posture compensation is expressed as:
[0190] In this case, translation compensation can be viewed as a learnable rotation center for rotating mesh replicas.
[0191] 2) Independent translation, fixed rotation: Each replica has its own set of independently learnable translation parameters, and the rotation compensation remains unchanged. The posture compensation is expressed as:
[0192] In this case, each rotated replica is given an additional translational degree of freedom, enabling it to explore a larger local space.
[0193] 3) Fully learnable translation and rotation: and They can all be learned independently. Posture compensation is expressed as:
[0194] In this case, each mesh copy can be treated as an independent mesh that can be freely translated and rotated in 3D space.
[0195] By selecting different configurations based on the geometric characteristics of the target object, a balance can be achieved between computational cost and computational accuracy. In some embodiments, configuration 1) can adapt to objects with good rotational symmetry and converge faster; configurations 2) and 3) can adapt to objects with complex shapes and no symmetry, with a higher upper limit of accuracy but slower convergence. Figure 10b is a visualization of an example of a teapot mesh replica used in the batch optimization strategy corresponding to the teapot in Figure 10a, where 101 is the spout and 102 is the handle. The batch processing object consists of 6 mesh replicas along the Y axis.
[0196] In the context of batch optimization pose estimation, objects optimized in batch mode are called "batch optimized objects", and other objects in the scene that are not optimized in batch mode are called "non-batch optimized objects". Dynamic optimization is performed on the replica with the smallest loss value in the batch to escape the local minimum trap. First, using the global initial pose and compensation posture P′ Δ The initial copy model M of the batch object dup To perform the conversion, the conversion formula is:
[0197] And convert the model of non-batch objects, the conversion formula is:
[0198] Where N is the number of non-batch optimized objects. Then, using a differentiable renderer and the camera intrinsic parameter K, a high-dimensional rendering model of a batch of objects can be generated And another high-dimensional rendering model for non-batch objects
[0199] However, when there are multiple objects in the scene, occlusion must be accounted for in the loss computation. Since the pose of an object with replicas may correspond to the pose of any replica in the batch and is unknown, directly computing the occlusion relationship between objects in the batch and other scene objects is difficult. This complexity stems from the need to enumerate all possible combinations of all replicas of all batched objects and all non-batched objects in the scene to accurately evaluate mutual occlusion.
[0200] To solve the above problem, in this embodiment, only one target object is set to batch optimization state in one optimization iteration, and occlusion is calculated separately: First, the occlusion perception high-dimensional data of all copies of the batch optimized object occluded by all other non-batch optimized objects is calculated. Then, the loss between the occlusion-aware high-dimensional data and the reference image of the batch optimized object is calculated, and the copy with the smallest loss value is selected as the best-fit copy. The pose of the best-fit copy can then be used to determine the pose of the batch optimized object, and all non-batch optimized objects can be calculated. The occlusion perception data. The occlusion calculation can refer to the description in the object pose estimation section.
[0201] Specifically, the loss function of a batch of objects can be expressed as:
[0202] L dup The minimum loss in and the index of the replica corresponding to the minimum loss can be expressed as:
[0203] in, Represents the loss value of each batch optimization (batch) j.
[0204] Then, calculate the loss for the non-batch optimized object:
[0205] Optimize the learnable parameter P Δ ,P B ,P NB The total loss of back propagation can be expressed as Loss total =L best +L batch +L non-batch P Δ ,P B ,P NB =argmin Loss total
[0206] By this method, batch objects with occlusion information can be effectively incorporated into the optimization process. The poses of the batch objects are finally determined as
[0207] In addition, single-camera systems are often subject to limitations such as occlusion and limited field of view, which can seriously affect the accuracy of pose estimation. Using a multi-camera system can alleviate the above problems because a multi-camera system can obtain different perspectives of an object and provide complementary information for pose estimation of multiple objects. Therefore, step S11 may include:
[0208] S111. Use several cameras to capture images from different angles to reduce the pose ambiguity caused by mutual occlusion and self-occlusion of the target.
[0209] S112. Use 3D point cloud reprojection to match the real-time target detection data of each camera.
[0210] Specifically, step S112 includes:
[0211] Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects;
[0212] Determining camera parameters, where the camera parameters include intrinsic parameters of the first camera and the second camera, and parameters for transforming the field of view of the first camera and the field of view of the second camera;
[0213] Based on the determined camera parameters, reprojecting the image of the first camera into the field of view of the second camera to obtain a reprojected three-dimensional point cloud;
[0214] comparing the mask information of the one or more objects in the image of the second camera with the reprojected three-dimensional point cloud to determine which object's mask information has the greatest overlap with the reprojected three-dimensional point cloud;
[0215] Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
[0216] When using a multi-camera system for object detection, the key is to solve the problem of ID mismatch. When multiple cameras detect the same object, the detection model may assign different IDs to the same object, and the problem of ID mismatch will occur. Since each camera provides an independent perspective, the model may not be able to identify whether the detection results under different perspectives correspond to the same object. For RGB images or RGBD images from a single camera view, existing object detection and segmentation methods output a dictionary containing the detected object information, including object ID, category, bounding box and mask. In order to solve the problem of ID mismatch, it is necessary to use this information to establish the correspondence between object IDs in different camera views to ensure that the same object can be consistently identified in each view. Therefore, it is usually necessary to integrate multi-view object information association technologies, such as triangulation, feature matching, etc. In this embodiment, an object ID matching algorithm based on point cloud or depth reprojection is utilized, which includes the following steps:
[0217] (1) Camera calibration
[0218] Intrinsic calibration involves determining the internal parameters of the camera, such as focal length (f x ,f y ), main point (c x ,c y ) and distortion coefficients k1, k2, p1, and p2. These parameters define how the camera projects real-world 3D image points onto the 2D image plane. This information can typically be loaded from the camera driver API or calibrated using a checkerboard-like pattern.
[0219] The extrinsic matrix determines the position and orientation of the camera relative to the world coordinate system, which is usually determined by a calibration plate (ChArUco calibration plate). The extrinsic parameters include the rotation matrix R and the translation vector T, which describe the transformation from the world coordinate system to the camera coordinate system. Given a 3D point (X, Y, Z), its corresponding 2D image point (u, v) can be converted using the following formula:
[0220] Among them, (X c ,Y c ,Z c ) is the coordinate of the point in the camera coordinate system. By solving the above formula using the corner points of the ChArUco calibration target, we can estimate the rotation matrix and translation vector, thereby accurately determining the position and orientation of the camera in the world coordinate system.
[0221] For the external calibration of multiple cameras, after the intrinsic calibration, the ChArUco plate is placed in the field of view of all cameras as a common reference object. Each camera captures images of the calibration plate, and from these images the pose (in terms of rotation and translation) of each camera relative to the ChArUco calibration plate is estimated. Once the pose of each camera relative to the ChArUco calibration plate is determined, a common world coordinate system can be established by selecting one camera as a reference. The external parameters of the other cameras are then transformed into the coordinate system of the reference camera, ensuring that all cameras are aligned in a unified spatial frame.
[0222] (2) Point cloud reprojection and ID matching
[0223] Taking two cameras as an example, the process of converting the depth image from camera 1 to a point cloud involves back-projecting the depth values into a 3D coordinate system using the intrinsic parameters of camera 1. By applying an extrinsic transformation between the two cameras, the point cloud is transformed into the coordinate system of camera 2. The 3D points are then re-projected into the image plane of camera 2 using the intrinsic parameters of camera 2, generating the depth information captured by camera 1 in the view of camera 2.
[0224] ID matching is performed after the points from camera 1 are reprojected onto the image plane of camera 2. For each object detected in view 1, the point cloud corresponding to the target object's surface is extracted using the mask information, and the corresponding reprojected points in the camera 2 view are calculated. The reprojected points are then compared with the mask of the object detected in view 2. By determining the ID of the mask in view 2 that has the highest overlap with the reprojected point cloud, a correspondence between the object IDs in the two camera views is established, ensuring consistent identification of the same object in both views.
[0225] Figures 11a-c show the multi-camera setup and the detection results from different viewpoints. Figures 12a-c show schematic diagrams of the reprojected 3D point cloud from camera view 1 to view 2, and Figures 12a-c are mask images. The white areas in the image represent different objects, and the black points in the image are the reprojected points of the point cloud of the surface of objects with different IDs in view 1 in the view of camera 2. By checking the overlap rate with the occlusion mask detected in view 2, the correspondence between the object IDs in the two camera views can be determined. In this example, ID (Identity) 1 (view 1) matches ID 4 (view 2), ID 2 (view 1) matches ID 1 (view 2), and ID 3 (view 1) matches ID 2 (view 2).
[0226] For each object, the ID is known in each camera view, which enables reference images of the object to be obtained from multiple angles. In each camera view, the digital twin can be rendered using shared learnable parameters (representing the object pose), and these rendered images are compared with the corresponding reference images to calculate the loss value of the view. By adding the loss values of all camera views, complementary information from multiple angles is utilized. Gradient descent is then used to optimize the shared learnable parameters, making it possible to accurately determine the pose of the object by combining data from all camera views.
[0227] Real-time pose estimation, as demonstrated in this embodiment, is crucial for achieving precise, adaptive robotic manipulation, particularly in dynamic environments where objects may move or change in spatial relationship to each other or to the robot. For tasks such as pick-and-place, assembly, and bundling, the robot must accurately estimate the pose of individual objects and their interactions with the robot. This becomes particularly important when managing multiple objects or handling complex operations such as attaching or detaching objects to or from a robot end-effector.
[0228] Real-time pose estimation systems provide direct input to the robot's control system, enabling it to execute precise movements. For example, when grasping an object, pose data enables the control system to calculate the optimal approach and manipulation strategy to minimize the risk of collision or slippage. Furthermore, when an object is attached to the robot, the system can utilize more advanced manipulation strategies, such as treating the object and robot as a single entity for collision avoidance and path planning. For example, the system can combine the relative poses of several objects to implement a more optimized control strategy. When the robot grasps an object, it treats the robot and object as a single entity for path planning and collision avoidance.
[0229] In more complex situations, such as tying multiple objects, the collective pose of the group must be managed while keeping the individual objects relatively stable. Real-time feedback from the robot, such as tactile and proprioceptive data, further enhances this process, enabling the system to detect and compensate for unexpected events such as slippage, deformation, or external interactions. By combining visual pose tracking with feedback from the robot's sensors, the system can dynamically adjust its control strategy to ensure robustness and precision in object-robot interactions. This integration of multimodal data is critical to improving the overall accuracy of manipulation tasks in complex and unpredictable environments.
[0230] Figure 13 illustrates the interaction between the digital robot and the object during the grasping operation. The dotted teacup represents the object as a separate pose-optimized object, while the solid line represents the object as it is fixed to the manipulator for pose estimation. From (a) to (b), the manipulator gradually approaches the handle of the teacup. In (c), the manipulator successfully grasps the teacup and secures the object to the manipulator. From (d) to (e), the pose of the teacup is continuously updated based on its relative position to the manipulator. In (f), the manipulator releases the teacup, freeing it from the manipulator.
[0231] The method for estimating multi-target poses provided in this embodiment has the following advantages:
[0232] 1. We create digital twins of objects from their high-dimensional data and simultaneously optimize their 6D poses at each iteration, taking into account physical occlusions between objects. This process is enhanced by integrating RGBXY derivatives and optimal transport. RGBXY derivatives capture color and spatial variations, enabling robust gradient-based optimization suitable for large object transformations, while optimal transport minimizes the cost function, effectively matching RGBXY features even when the objects have significant differences in positioning. As a result, our pose estimation method accurately reflects the real-time pose of all target objects with high fidelity.
[0233] 2. The initial pose estimation method and batch optimization strategy ensure fast convergence of the optimization while getting rid of the local minimum trap in the pose optimization process.
[0234] 3. Supports multi-camera stream input, which can alleviate the pose ambiguity problem caused by occlusion and can handle complex scene-level pose estimation.
[0235] 4. Integrating robot pose estimation and object pose estimation into a unified paradigm can be applied to real-time digital twin reconstruction, realizing a closed-loop and real-time feedback method for robot control.
[0236] Example 2
[0237] This embodiment provides a robot control method, wherein:
[0238] Using the method for estimating multi-target poses in Example 1 to obtain pose information of the robot and pose information of the target object;
[0239] Inputting the posture information of the robot and the posture information of the target object into the control system of the robot;
[0240] Based on the input, the control system controls the interaction of the robot with the target object.
[0241] The robot control method provided in this embodiment can reflect the real-time posture of all target objects with high fidelity by utilizing the above-mentioned method for estimating multi-target postures, and can enable the control system to accurately control the interaction between the robot and the object with the optimal strategy, integrating the robot posture estimation and the object posture estimation into a unified paradigm, thereby realizing a closed-loop and real-time feedback method for robot control.
[0242] Example 3
[0243] This embodiment provides a system for estimating the poses of multiple targets, the system comprising: one or more processors, the one or more processors implementing the method for estimating the poses of multiple targets in embodiment 1 or the method for controlling the robot in embodiment 2.
[0244] For example, the one or more processors may include one or more GPUs, one or more CPUs, or any other technically feasible combination of one or more processors such as ASIC chips. In some examples, the one or more processors such as ASIC chips may include memory.
[0245] Preferably, the system further comprises one or more cameras for acquiring video data and / or image data of the estimated target.
[0246] Preferably, the system further comprises a robot;
[0247] The system uses the method for estimating multi-target poses in Example 1 to obtain pose information of the robot and pose information of the target object;
[0248] The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
[0249] As for the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only illustrative, wherein the modules described as separate components may or may not be physically separated, and the modules displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. Ordinary technicians in this field can understand and implement it without expending creative work.
[0250] Example 4
[0251] Figure 14 is a schematic diagram of the structure of an electronic device provided in Example 4 of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method of Example 1 or Example 2 is implemented. The electronic device 30 shown in Figure 14 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.
[0252] As shown in FIG14 , the electronic device 30 may be a general-purpose computing device, such as a server device. Components of the electronic device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, and a bus 33 connecting various system components (including the memory 32 and the processor 31).
[0253] The bus 33 includes a data bus, an address bus, and a control bus.
[0254] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .
[0255] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0256] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the method of embodiment 1 or embodiment 2 of the present disclosure.
[0257] The electronic device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown, the network adapter 36 communicates with other modules of the model-generating device 30 via a bus 33. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0258] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0259] Example 5
[0260] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method of embodiment 1 or embodiment 2 is implemented.
[0261] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0262] In a possible implementation manner, the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the method of embodiment 1 or embodiment 2.
[0263] The program code for executing the present disclosure may be written in any combination of one or more programming languages, and may be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0264] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.
Claims
1. A method for estimating multi-target poses, the method comprising: Acquire real-time video data, wherein the real-time video data includes first image frames of the multiple targets; Based on the first image frame, obtaining first high-dimensional segmented pixel data, wherein the first high-dimensional segmented pixel data includes first pixel data of each target, and the first pixel data is a three-dimensional or higher-dimensional array; Based on a first preset posture, performing differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array; Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, the first preset posture is updated to obtain a first target posture.
2. The method for estimating multiple target poses as claimed in claim 1, wherein: The real-time video data also includes at least a second image frame of the multiple targets; After the step of obtaining the target posture, the method further includes: Based on the second image frame, obtaining second high-dimensional segmented pixel data; Based on the first target pose, performing differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array; Based on minimizing the error between the second high-dimensional segmented pixel data and the second high-dimensional data, updating the first target pose to obtain a second target pose; Based on other image frames of the real-time video data, the posture corresponding to the previous image frame is updated to obtain other target postures.
3. The method for estimating multiple target poses as claimed in claim 1, wherein: The step of obtaining first high-dimensional segmented pixel data based on the first image frame comprises: Using a segmentation algorithm to segment each object in the first image frame to obtain a mask array; The first high-dimensional segmented pixel data is obtained based on the mask array and the channel data of the first image frame.
4. The method for estimating multiple target poses as claimed in claim 1, wherein: The step of performing differentiable rendering on the three-dimensional model of the multi-target based on the first preset posture to obtain first high-dimensional data includes: Based on the preset poses of the multiple targets, transforming the three-dimensional models of the multiple targets from an object coordinate system to a camera coordinate system according to learnable pose transformation parameters, wherein the pose transformation parameters are used to characterize the real-time poses of the targets; Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model to obtain the first high-dimensional data.
5. The method for estimating multiple target poses as claimed in claim 4, wherein: After the step of performing differentiable rendering on the three-dimensional model of the multiple targets to obtain the first high-dimensional data, the method further includes: Calculating occlusion perception data corresponding to the occlusion information of the target using the depth information and contour information of the first high-dimensional data; The step of updating the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises: The gradient descent method is used to iteratively minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable posture transformation parameters.
6. The method for estimating multiple target poses as claimed in claim 4, wherein: The step of performing differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data further includes: The converted three-dimensional model is subjected to the differentiable rendering to obtain first high-dimensional data having a variable contour, wherein a range of the variable contour is determined based on a degree of difference between the first pixel data and a preset posture.
7. The method for estimating multiple target poses as claimed in claim 4, wherein: Performing the differentiable rendering on the converted three-dimensional model to obtain first high-dimensional data includes: Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain a second RGBXY derivative of the first high-dimensional segmented pixel data, wherein the first RGBXY derivative and the second RGBXY derivative both include information of color change and spatial position change.
8. The method for estimating multiple target poses as claimed in claim 7, wherein: The updating of the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises: Based on optimal transmission, the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data is minimized; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data, the first preset posture is updated to obtain the first target posture.
9. The method for estimating multiple target positions and postures as claimed in claim 4, wherein: The target includes a robot; the posture conversion parameters include joint angles of the robot; The step of converting the three-dimensional model of the multiple targets to camera coordinates based on the preset poses of the multiple targets according to the learnable pose conversion parameters comprises: Based on the preset posture and preset joint angles of the robot, the three-dimensional model of each component of the robot is converted from the object coordinate system of each component to the camera coordinate system by calculating the forward kinematics matrix of the robot; The step of performing differentiable rendering on the converted three-dimensional model based on preset camera parameters to obtain the first high-dimensional data includes: Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model of the robot to obtain the first high-dimensional data.
10. The method for estimating multiple target positions and postures according to claim 4, wherein: The preset poses include preset translation parameters and preset rotation parameters of the multiple targets; Before the step of converting the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose conversion parameters, the method further includes: Determine preset translation parameters using point cloud data of each target; Using preset image indicators and / or features and / or pre-trained models, calculate the similarity between the templates in the template set and the reference image to determine the most matching target template, and determine the preset rotation parameters according to the rotation parameters corresponding to the target template; The template set includes rotation parameters and corresponding images, and is generated by performing a plurality of uniform samplings on a selected rotation axis.
11. The method for estimating multiple target positions and postures as claimed in claim 10, wherein: Before calculating the similarity between the templates in the template set and the reference image to determine the most matching target template, the method further includes: Performing uniform sampling several times on the selected rotation axis to generate sampling data of a preset number of samplings; Based on preset camera parameters, a three-dimensional model of the target, initial translation parameters and a preset number of sampling data, a set of images of different positions and postures of the three-dimensional model is rendered to form the template set.
12. The method for estimating multiple target positions and postures as claimed in claim 10, wherein: Before the step of performing differentiable rendering on the three-dimensional model of the multiple objects to obtain the first high-dimensional data, the method further includes: Acquire a plurality of replica models of the three-dimensional model of the batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters; Calculate the occlusion perception data corresponding to the occlusion information of the batch optimized objects using the high-dimensional data obtained by rendering the replica model of the batch optimized objects; The loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object is calculated to find the target replica model with the minimum loss value. The target replica model is used to determine the preset pose for this round of pose estimation.
13. The method for estimating multiple target positions and postures according to claim 1, wherein: The step of acquiring real-time video data comprises: Use several cameras to perform surround shooting from different angles to reduce the pose ambiguity caused by mutual occlusion and self-occlusion of the target; The real-time target detection data of each camera are matched with each other by using three-dimensional point cloud reprojection to obtain the real-time video data.
14. The method for estimating multiple target positions and postures as claimed in claim 13, wherein: The step of using the three-dimensional point cloud reprojection to match the real-time target detection data of each camera includes: Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects; Determining camera parameters, the camera parameters including intrinsic parameters of the first camera and the second camera, and extrinsic parameters for transforming a field of view of the first camera and a field of view of the second camera; Based on the determined camera parameters, reprojecting the depth image or three-dimensional point cloud data of the first camera into the field of view of the second camera to obtain a corresponding reprojected two-dimensional point group; comparing the mask information of the one or more objects in the image of the second camera with the reprojected two-dimensional point group to determine which object's mask information has the greatest overlap with the reprojected two-dimensional point group; Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
15. A method for controlling a robot, the method comprising: Acquire the posture information of the robot and the posture information of the target object by using the method for estimating multi-target postures as described in any one of claims 1 to 14; Inputting the posture information of the robot and the posture information of the target object into the control system of the robot; Based on the input, the control system controls the interaction of the robot with the target object.
16. A computer program product, comprising a computer program, wherein when the computer program is executed by one or more processors, the computer program performs the method for estimating multi-target poses as described in any one of claims 1 to 14 or the method for controlling a robot as described in claim 15.
17. A system for estimating multiple target positions, the system comprising: One or more processors, wherein the one or more processors implement the method for estimating multi-target poses as described in any one of claims 1 to 14 or the robot control method as described in claim 15.
18. The system for estimating multiple target poses according to claim 17, wherein: The system also includes one or more cameras for acquiring video data and / or image data of the estimated target.
19. The system for estimating multiple target poses according to claim 17, wherein: The system also includes a robot; The system obtains the posture information of the robot and the posture information of the target object by using the method for estimating multi-target postures according to any one of claims 1 to 14; The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for estimating multi-target poses as described in any one of claims 1 to 14 or the method for controlling a robot as described in claim 15 is performed.
21. A computer-readable medium having computer instructions stored thereon, wherein the computer instructions, when executed by a processor, implement the method for estimating multi-target poses as described in any one of claims 1 to 12 or the method for controlling a robot as described in claim 15.
22. A computer program, which, when executed by one or more processors, can execute the method for estimating multi-target poses as described in any one of claims 1 to 14 or the robot control method as described in claim 15.
Citation Information
Patent Citations
Network training method and device and attitude prediction method and device
CN111783986A
Robot grabbing method and system based on 6D pose estimation
CN115641322A
Object pose estimation method based on template matching and probability distribution
CN115761734A
Three-dimensional model generation method and device, computer equipment and storage medium
CN116824092A
Systems and methods for six-degree of freedom pose estimation of deformable objects
US20220343537A1