Method for estimating multi-target poses, robot control method, system and product
Through the segmentation and differentiable rendering of real-time video data, the preset pose is updated, which solves the problem of multi-objective pose estimation in complex environments, and realizes high-precision real-time pose estimation and robot control.
Patent Information
- Application Number
- CN202510090897.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively estimate the position of multiple targets in complex environments, and cannot meet the needs of high-precision control of robots.
By acquiring real-time video data, high-dimensional pixel data are obtained by segmenting, and preset poses are updated using differentiable rendering and minimized error methods to achieve real-time estimation of multi-objective poses.
It realizes high-precision real-time estimation of multi-objective poses, supports high-precision control of robots in complex environments, and improves the accuracy of dynamic tracking and pose calibration.
Smart Images

Figure CN120047534A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of computer image processing, and in particular to a method for estimating multi-target postures, a robot control method, system and product. Background Art
[0002] Taking US Patent No. US10565747B2 and US Patent No. US10565747B2 as examples, US10565747B2 records a solution for using a particle filter and an autoencoder to estimate the pose of an object in a static image, and US10565747B2 records the establishment of a differentiable rendering pipeline for a given static two-dimensional image data input. However, the prior art lacks a method for estimating the pose of multiple targets in a complex environment, and cannot meet the requirements for high-precision control of robots. Summary of the invention
[0003] The technical problem to be solved by the present disclosure is to overcome the above-mentioned defects in the prior art and to provide a method for estimating multi-target postures, a robot control method, system and product.
[0004] The present invention solves the above technical problems through the following technical solutions:
[0005] The present disclosure provides a method for estimating multi-target poses, the method comprising:
[0006] Acquire real-time video data, wherein the real-time video data includes first image frames of the multiple targets;
[0007] Based on the first image frame, obtaining first high-dimensional segmented pixel data, wherein the first high-dimensional segmented pixel data includes first pixel data of each target, and the first pixel data is a three-dimensional or higher-dimensional array;
[0008] Based on a first preset posture, performing differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array;
[0009] Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, the first preset posture is updated to obtain a first target posture.
[0010] Preferably, the real-time video data further includes at least a second image frame of the multiple targets;
[0011] After the step of obtaining the target posture, the method further includes:
[0012] Based on the second image frame, obtaining second high-dimensional segmented pixel data;
[0013] Based on the first target pose, performing differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array;
[0014] Based on minimizing the error between the second high-dimensional segmented pixel data and the second high-dimensional data, updating the first target pose to obtain a second target pose;
[0015] Based on other image frames of the real-time video data, the posture corresponding to the previous image frame is updated to obtain other target postures.
[0016] Preferably, the step of obtaining first high-dimensional segmented pixel data based on the first image frame comprises:
[0017] Using a segmentation algorithm to segment each object in the first image frame to obtain a mask array;
[0018] The first high-dimensional segmented pixel data is obtained based on the mask array and the channel data of the first image frame.
[0019] Preferably, the step of performing differentiable rendering on the three-dimensional model of the multi-target based on the first preset posture to obtain the first high-dimensional data comprises:
[0020] Based on the preset poses of the multiple targets, transforming the three-dimensional models of the multiple targets from an object coordinate system to a camera coordinate system according to learnable pose transformation parameters, wherein the pose transformation parameters are used to characterize the real-time poses of the targets;
[0021] Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model to obtain the first high-dimensional data.
[0022] Preferably, after the step of performing differentiable rendering on the three-dimensional model of the multi-objective to obtain the first high-dimensional data, the method further comprises:
[0023] Calculating occlusion perception data corresponding to the occlusion information of the target using the depth information and contour information of the first high-dimensional data;
[0024] The step of updating the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises:
[0025] The gradient descent method is used to iteratively minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable posture transformation parameters.
[0026] Preferably, the step of performing differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data comprises:
[0027] The converted three-dimensional model is subjected to the differentiable rendering to obtain first high-dimensional data having a variable contour, wherein a range of the variable contour is determined based on a degree of difference between the first pixel data and a preset posture.
[0028] Preferably, performing the differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data comprises:
[0029] Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain a second RGBXY derivative of the first high-dimensional segmented pixel data, wherein the first RGBXY derivative and the second RGBXY derivative both include information of color change and spatial position change.
[0030] Preferably, updating the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises:
[0031] Based on optimal transmission, the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data is minimized; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data, the first preset posture is updated to obtain the first target posture.
[0032] Preferably, the target includes a robot; the posture conversion parameters include joint angles of the robot;
[0033] The step of converting the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose conversion parameters based on the preset poses of the multiple targets comprises:
[0034] Based on the preset posture and preset joint angles of the robot, the three-dimensional model of each component of the robot is converted from the object coordinate system of each component to the camera coordinate system by calculating the forward kinematics matrix of the robot;
[0035] The step of performing differentiable rendering on the converted three-dimensional model based on preset camera parameters to obtain the first high-dimensional data includes:
[0036] Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model of the robot to obtain the first high-dimensional data.
[0037] Preferably, the preset posture includes preset translation parameters and preset rotation parameters of the multiple targets;
[0038] Before the step of transforming the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose transformation parameters, the method further comprises:
[0039] Determine preset translation parameters using point cloud data of each target;
[0040] Using preset image indicators and / or features and / or pre-trained models, calculate the similarity between the templates in the template set and the reference image to determine the most matching target template, and determine the preset rotation parameters according to the rotation parameters corresponding to the target template;
[0041] The template set includes rotation parameters and corresponding images, and is generated by performing a plurality of uniform samplings on a selected rotation axis.
[0042] Preferably, before the step of calculating the similarity between the templates in the template set and the reference image to determine the most matching target template, the method further comprises:
[0043] Performing uniform sampling several times on the selected rotation axis to generate sampling data of a preset number of samplings;
[0044] Based on preset camera parameters, a three-dimensional model of the target, initial translation parameters and a preset number of sampling data, a set of images of different positions and postures of the three-dimensional model is rendered to form the template set.
[0045] Preferably, before the step of performing differentiable rendering on the three-dimensional model of the multi-objective to obtain the first high-dimensional data, the method further comprises:
[0046] Acquire a plurality of replica models of the three-dimensional model of the batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters;
[0047] Calculate the occlusion perception data corresponding to the occlusion information of the batch optimized objects using the high-dimensional data obtained by rendering the replica model of the batch optimized objects;
[0048] The loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object is calculated to find the target replica model with the minimum loss value. The target replica model is used to determine the preset pose for this round of pose estimation.
[0049] Preferably, the step of acquiring real-time video data includes:
[0050] Use several cameras to perform surround shooting from different angles to reduce the pose ambiguity caused by mutual occlusion and self-occlusion of the target;
[0051] The real-time target detection data of each camera are matched with each other by using three-dimensional point cloud reprojection to obtain the real-time video data.
[0052] Preferably, the step of using three-dimensional point cloud reprojection to match the real-time target detection data of each camera includes:
[0053] Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects;
[0054] Determining camera parameters, the camera parameters including intrinsic parameters of the first camera and the second camera, and extrinsic parameters for transforming a field of view of the first camera and a field of view of the second camera;
[0055] Based on the determined camera parameters, reprojecting the depth image or three-dimensional point cloud data of the first camera into the field of view of the second camera to obtain a corresponding reprojected two-dimensional point group;
[0056] comparing the mask information of the one or more objects in the image of the second camera with the reprojected two-dimensional point group to determine which object's mask information has the greatest overlap with the reprojected two-dimensional point group;
[0057] Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
[0058] The present disclosure also provides a robot control method, the method:
[0059] The method for estimating multi-target poses as described above is used to obtain the pose information of the robot and the pose information of the target object;
[0060] Inputting the posture information of the robot and the posture information of the target object into the control system of the robot;
[0061] Based on the input, the control system controls the interaction of the robot with the target object.
[0062] Those skilled in the art will appreciate that the robot control system can control the robot to achieve precise movement based on the obtained posture information of the robot and the posture information of the target object. For example, when grasping the target object, the input posture information can enable the control system to calculate the optimal path and movement strategy, thereby minimizing the risk of collision or sliding.
[0063] The present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0064] The present disclosure also provides a system for estimating multi-target poses, the system comprising:
[0065] One or more processors, wherein the one or more processors implement the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0066] Preferably, the system further comprises one or more cameras for acquiring video data and / or image data of the estimated target.
[0067] Preferably, the system further comprises a robot;
[0068] The system uses the method for estimating multi-target poses as described above to obtain pose information of the robot and pose information of the target object;
[0069] The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
[0070] The present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for estimating multi-target postures as described above or the method for controlling a robot as described above is implemented.
[0071] The present disclosure also provides a computer-readable medium having computer instructions stored thereon, which, when executed by a processor, implement the method for estimating multi-target poses as described above or the method for controlling a robot as described above.
[0072] The positive and progressive effects of this disclosure are:
[0073] The method for estimating the posture of multiple targets provided by the present invention takes a real-time video stream as input, processes each latest image frame of the real-time video stream in sequence, and realizes real-time dynamic tracking; by performing image segmentation on each image frame to identify each target therein, a high-dimensional segmented pixel array including pixel data of each target is obtained, and the high-dimensional array is compared with a high-dimensional array obtained by differentiable rendering of known three-dimensional models of the multiple targets, and the posture is estimated by minimizing the error, thereby realizing posture estimation of multiple targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of this specification. For ordinary technicians in this field, without paying creative work, this specification can also be applied to other similar scenarios based on these drawings.
[0075] Figure 1 A first flow chart of a method for estimating multi-target poses provided in Embodiment 1 of the present disclosure;
[0076] Figure 2 A second flow chart of a method for estimating multi-target poses provided in Embodiment 1 of the present disclosure;
[0077] Figure 3 A schematic diagram of static pose calibration for object pose estimation provided in Example 1 of the present disclosure.
[0078] Figure 4 A schematic diagram of real-time pose tracking for object pose estimation provided in Example 1 of the present disclosure.
[0079] Figure 5 A schematic diagram of static pose calibration of a robot and an object pose estimation provided in Example 1 of the present disclosure.
[0080] Figure 6 A schematic diagram of static pose calibration of a robot and an object pose estimation provided in Example 1 of the present disclosure.
[0081] Figure 7 A schematic diagram of the process of retrieving rough initial rotation parameters provided in Embodiment 1 of the present disclosure.
[0082] Figure 8 The visualization results of different initial poses with / without initial translation parameters and initial rotation parameters at the beginning of pose estimation provided in Example 1 of the present disclosure.
[0083] Fig. 9 A schematic diagram of the large-angle rotation problem in the optimization process provided in Example 1 of the present disclosure.
[0084] Fig.10 A visualization effect of an example of a teapot grid replica used in the batch optimization strategy provided in Example 1 of the present disclosure.
[0085] Fig.11 A schematic diagram of the settings of multiple cameras and detection results at different viewing angles provided in Example 1 of the present disclosure.
[0086] Fig.12a -c is the camera view provided in Example 1 of the present disclosure Figure 1 To view Figure 2Schematic diagram of the reprojected 3D point cloud.
[0087] Fig.13 A schematic diagram of the effect of the interaction between the robot and the object during the grasping operation provided in Example 1 of the present disclosure.
[0088] Fig.14 A schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present disclosure. DETAILED DESCRIPTION
[0089] The present disclosure is further described below by way of examples, but the present disclosure is not limited to the scope of the examples.
[0090] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various places in the text does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0091] It should be understood that the "device", "system", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0092] As shown in this document, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural, unless the context clearly indicates an exception. Generally speaking, the terms "include" and "comprise" only indicate that the steps and elements that have been clearly identified are included, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0093] The definitions of inclusion in this document, such as the terms "having", "may have", "include" or "may include" used herein, indicate the existence of the corresponding functions, operations, elements, etc. herein, and do not limit the existence of one or more other functions, operations, elements, etc. In addition, it should be understood that the terms "including" or "having" used herein indicate the existence of the features, numbers, steps, operations, elements, components, or a combination thereof described in the specification, without excluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or a combination thereof.
[0094] Flowcharts are used herein to illustrate the operations performed by the systems of the embodiments of this invention. It should be understood that the preceding or following operations are not necessarily performed precisely in order. Instead, the various steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.
[0095] Example 1
[0096] Please refer to Figure 1 , which is a first flow chart of the method for estimating multi-target poses in this embodiment. Specifically, Figure 1 As shown, the method includes:
[0097] S11, acquiring real-time video data, the real-time video data including first image frames of multiple targets; by taking the real-time video stream as input, the real-time video stream can be divided into a plurality of real-time image frames;
[0098] S12. Based on the first image frame, first high-dimensional segmented pixel data is obtained, where the first high-dimensional segmented pixel data includes first pixel data of each target, and the first pixel data is a three-dimensional or higher-dimensional array; each image frame can be segmented to identify each target therein, thereby obtaining a high-dimensional segmented pixel array including pixel data of each target.
[0099] S13, based on the first preset posture, performing differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array;
[0100] S14. Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, update the first preset posture to obtain a first target posture.
[0101] After processing the first image frame, by sequentially processing each image frame of the real-time video stream, dynamic real-time estimation of the positions and postures of multiple targets can be achieved. Therefore, the real-time video data also includes at least the second image frame of the multiple targets;
[0102] Please refer to Figure 2 , which is a second flow chart of the method for estimating multi-target poses in this embodiment. Specifically, Figure 2 As shown, after step S14, the method further includes:
[0103] S15, obtaining second high-dimensional segmented pixel data based on the second image frame;
[0104] S16, based on the first target posture, performing differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array;
[0105] S17, updating the first target pose to obtain a second target pose based on minimizing an error between the second high-dimensional segmented pixel data and the second high-dimensional data;
[0106] S18. Based on other image frames of the real-time video data, update the posture corresponding to the previous image frame to obtain other target postures.
[0107] In some examples, the method steps described below can be implemented by a program or a custom circuit, or a combination of a custom circuit and a program. For example, the method steps below can be performed by a combination of one or more GPUs, one or more CPUs, or any other technically feasible one or more processors such as ASIC chips. In some examples, one or more processors such as ASIC chips may include a memory. Further, those skilled in the art will appreciate that the scope of protection of the present disclosure may include any system and device that can perform the following method steps.
[0108] In an optional implementation, step S12 includes:
[0109] S121, segmenting each target in the first image frame using a segmentation algorithm to obtain a mask array;
[0110] S122. Obtain first high-dimensional segmented pixel data based on the mask array and the channel data of the first image frame.
[0111] Image Segmentation
[0112] The inventors have found that image segmentation algorithms involve dividing an image (RGB or RGB depth) into multiple fragments (pixel sets) to simplify the representation of the image into more meaningful and easier to analyze content. Typically, image segmentation algorithms assign a label to each pixel in the image, and pixels with the same label share certain features (color, edge, texture, category, object instance, etc.). According to different classification criteria, image segmentation algorithms can be divided into categories such as semantic segmentation, instance segmentation, and panoramic segmentation. Among them, in some embodiments of the present disclosure, for example, an instance segmentation algorithm can be used, in which each fragment of the image uniquely represents a specific instance of an object. This method not only classifies the pixels of the image into different object classes, but also distinguishes different instances of the same class. For example, in a street scene, instance segmentation will identify and separate each individual car, pedestrian, or any other important object, treating each object as a unique entity, even though they may belong to the same category.
[0113] In another optional implementation, step S13 includes:
[0114] S131. Based on the preset poses of multiple targets, the three-dimensional models of the multiple targets are transformed from the object coordinate system to the camera coordinate system according to the learnable pose transformation parameters, and the pose transformation parameters are used to characterize the real-time pose of the target; object pose estimation is the process of determining the position and orientation (6D pose) of an object in space, which can be achieved by processing image data.
[0115] S132, based on the preset camera parameters, performing differentiable rendering on the converted 3D model to obtain the first high-dimensional data. Differentiable rendering can calculate the gradient of the 3D target and propagate through the image, and bridge the gap between 2D image processing and 3D image processing, and optimize the 3D scene parameters (camera pose, pose or shape of the 3D model, lighting conditions, etc.) according to the back propagation gradient of the rendering output.
[0116] Specifically, taking the image frame as the input image, it can be expressed as The input image can be an RGB image or RGBD image containing N targets; the segmentation algorithm can be used to segment the input image and output a mask array Then, combined with the channel data of the image frame, a high-dimensional array can be obtained. That is, high-dimensional segmented pixel data, where C represents the channel data of each image frame (such as contour, RGB and depth channels). ⊙ represents element-wise multiplication, which applies the mask to the input image I, which involves the array multiplication broadcast mechanism, that is, Repeat N times to get And I mask Perform the ⊙ operation, and then get the shape of A as 3(4) corresponds to the dimension I mask Superposition is performed to obtain the dimension corresponding to C.
[0117] The task of 6D object pose estimation is to estimate the rigid transformation of all targets. The rigid transformation can be expressed as Rigid transformation is used to transform the 3D object model M = {M n |n∈{1,2,…N}} maps from the object coordinate system to the camera coordinate system. Assume that the 3D object model M and the camera intrinsic parameters K are known. Each model M n It can be defined as A set of vertices in and a set of polygons describing the surface of the object. n A 4×4 rigid transformation matrix P n= [R, T; 0, 1], where R is a 3×3 rotation matrix and T is a 3×1 translation vector. The translation T specifies the origin of the object coordinates in the camera coordinate system. To ensure that R is physically valid, the rotation matrix R is represented by a quaternion q, which can be further converted into a 3×3 rotation matrix.
[0118] Then, the pose estimation process is as follows: For a given initial pose P′ and object model M, the object model can be transformed to obtain M P′ = {P′ n ·M n |n∈{1,2,…N}}, where P′ n is the initial pose of the nth object, M n is the 3D model of the nth object. Using a differentiable renderer And the camera intrinsic parameter K, can be rendered to get the first high-dimensional data Where H' and W' represent the height and width of the high-resolution rendered image, respectively, and C represents the number of channels (such as alpha, RGB, depth). It should match the structure of the reference array A, which is the high-dimensional array generated by the segmentation algorithm and encodes the class identification information.
[0119] In an optional implementation, step S132 specifically includes:
[0120] The converted three-dimensional model is subjected to differentiable rendering to obtain first high-dimensional data having a variable contour, wherein a range of the variable contour is determined based on a degree of difference between the first pixel data and a preset posture.
[0121] In another optional implementation, step S132 specifically includes:
[0122] Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain a second RGBXY derivative of the first high-dimensional segmented pixel data, wherein the first RGBXY derivative and the second RGBXY derivative both include information of color change and spatial position change.
[0123] Further, step S14 may include:
[0124] Based on optimal transmission, the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data is minimized; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data, the first preset posture is updated to obtain the first target posture.
[0125] Differentiable Rendering
[0126] The inventors have found that in some embodiments of the present disclosure, the gradient of a three-dimensional target object can be calculated by using a differentiable rendering method and can be propagated between images. The differentiable rendering method can connect the two-dimensional image processing method and the three-dimensional image processing method, and can optimize the three-dimensional scene data (for example, camera posture, three-dimensional model posture or shape, light conditions, etc.) by back-propagating the gradient of the rendered output.
[0127] The document Liu, S., Li, T., Chen, W., & Li, H. Soft rasterizer: A differentiable renderer for image-based 3d reasoning. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) (pp. 7708-7717). (2019) records the loss calculation by rendering a single real image of geometric primitives, and all the contents disclosed are incorporated into the present disclosure. In some embodiments of the present invention, the rendered smooth edge contour tensor has a variable smooth edge range, and the smooth edge range can be adjusted to be larger or smaller according to the difference between the real object pose and the virtual digital twin body pose. When there is a large difference between the real object pose and the virtual digital twin body pose, a larger smooth range can be used to allow the gradient flow of the digital twin body when it is away from the real target object. When the difference between the two is small, a reduced smooth edge can be used for more accurate pose estimation.
[0128] Occlusion in Rendered Array
[0129] The inventors found that since the reference high-dimensional array A only captures the visual information after occlusion, the high-dimensional rendering model Contains complete visual information of each target and does not consider the occlusion between objects. Direct comparison will lead to pose estimation errors. Therefore, it is necessary to render the model in a high-dimensional After step S13, the method further includes:
[0130] S21, using the depth information and contour information of the first high-dimensional data to calculate and obtain occlusion perception data corresponding to the occlusion information of the target; high-dimensional rendering model The occlusion information can be calculated by using a z-buffer (depth buffer) to determine the visibility of objects at each pixel. A set of binary object masks can be used in the process. and the corresponding depth map Where n∈N is the index of the object in the array. Specifically, the non-object area is first regarded as far away, and a large constant δ is added to the pixels in the non-object area of the depth map for depth adjustment:
[0131] D′(x,y,n)=D(x,y,n)+(1-Mask(x,y,n))×δ
[0132] For each pixel (x, y), the frontmost object is determined by finding the object with the smallest depth among all objects for visibility measurement:
[0133]
[0134] The occlusion mask for each pixel of each object is calculated as:
[0135] F(x,y,n)=1[γ(x,y)=n]
[0136] Where 1 is an indicator function which returns 1 if the condition is true, otherwise it returns 0.
[0137] Multiplying the occlusion mask with the initial object mask gives the final occlusion mask for each object:
[0138] O(x,y,n)=F n (x,y)×Mask(x,y,n)
[0139] Using the mask O(x,y,n), we can calculate an array Considering
[0140]
[0141] RGBXYDerivatives
[0142] The inventors discovered that RGBXY derivatives are a specific type of gradient used in differentiable rendering that takes into account both color (RGB) and spatial (XY) variations in the image. Unlike traditional per-pixel derivatives that are calculated only based on color differences, RGBXY derivatives also incorporate spatial variations, thereby improving the robustness of gradient-based optimization. These derivatives can more accurately track object transformations (such as position and rotation) even when there is a significant difference between the initial pose and the target pose.
[0143] Step S14 includes:
[0144] S141. Iterate using the gradient descent method to minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable posture transformation parameters.
[0145] The pose transformation parameter P is set as a learnable parameter to estimate the object pose by minimizing the error between the real high-dimensional array and the occlusion-aware rendering array. The optimization goal is:
[0146]
[0147] Among them, Loss is a metric used to evaluate the difference between each target in the captured image I and the corresponding digital twin in the rendered image. The loss function is defined as a weighted combination of the error of each modality (such as mean square error (MSE), mean absolute error, intersection over union (IoU), DICE coefficient, etc.) and the optimal transmission error of the RGBXY points of the occlusion-aware image:
[0148]
[0149] Loss = λ c L c +λ d L d +λ mask L mask +λ ot L ot
[0150] Among them, A * (y,x,n) represents the pixel value of the mode* of A, which can be RGB color, depth and contour occlusion; Q i (n) represents point i in a set of RGBXY points (including only points with non-zero RGB color values, XY represents spatial coordinates) in a reference image taken from target n; represents a point σ(i) in a set of RGBXY points in the occlusion-adjusted rendered image of target n; the matching function σ(·) defines a one-to-one correspondence between points in the captured reference image and points in the occlusion-adjusted rendered image, minimizing the total distance between the matched images; λ c ,λ d ,λ mask andλ ot Represents the weight of the balance loss. The above loss function is fully differentiable, and the gradient descent method can be used to optimize the pose transformation parameter P.
[0151] Optimal Transport
[0152] The inventors found that the mathematical framework of optimal transfer can find the most effective way to transform one probability distribution into another by minimizing the cost function. In the field of computer vision and rendering, optimal transfer is used to match pixels or features between two images (such as a rendered image and a reference image) to ensure that corresponding features are correctly aligned, even when there are significant differences in the positioning of objects.
[0153] The document Xing, J., Luan, F., Yan, LQ, Hu, X., Qian, H., & Xu, K. Differentiable rendering using RGBXY derivatives and optimal transport. ACM Transactions on Graphics (TOG), 41(6), 1-13. (2022) records an innovative method for differentiable rendering using RGBXY derivatives and optimal transport, which solves the significant limitations of traditional pixel derivative methods, and the entire content disclosed is incorporated into the present disclosure. It utilizes 5DRGBXY derivatives to capture changes in color and spatial position, improving the robustness of the optimization task, especially for remote object transformations. The inventors found that some embodiments of the present disclosure can integrate RGBXY derivatives and optimal transport into the real-time differentiable rendering system of the present embodiment based on this technology. By doing so, significant advantages are obtained in processing dynamic environments with continuous video input because it improves the accuracy of object pose estimation. RGBXY derivatives take into account color and spatial changes, providing more robust gradient-based optimization for large object transformations. Optimal transfer minimizes a cost function that quantifies the difference between the RGBXY features of the rendered and target images, ensuring efficient alignment. By minimizing this function, optimal transfer finds the most efficient way to match features even when there are significant differences in object positioning. The combination of RGBXY optimal transfer improves silhouette matching across frames, making it more robust in dynamic multi-object environments.
[0154] The inventors found that the above-mentioned target pose estimation process can be implemented using differential rendering libraries such as SoftRas, Pytorch3D and nvdiffrast. The document Liu, S., Li, T., Chen, W., & Li, H. Soft rasterizer: A differentiable renderer for image-based 3d reasoning. In Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) (pp. 7708-7717). (2019) records the use of the SoftRas differential rendering library to implement pose estimation, and all the contents disclosed are incorporated into the present disclosure. Reference Ravi, N., Reizenstein, J., Novotny, D., Gordon, T., Lo, WY, Johnson, J., & Gkioxari, G. Accelerating 3d deep learning with pytorch3d.arXiv:2007.08501 (2020) records the use of the Pytorch3D differential rendering library to implement pose estimation, and all of its disclosed contents are incorporated into the present disclosure. References Laine, S., Hellsten, J., Karras, T., Seol, Y., Lehtinen, J., & Aila, T. Modular primitives for high-performance differentiable rendering. ACM Transactions on Graphics (TOG), 39(6), 1-14. (2020). and J. Tremblay, B. Wen, V. Blukis, B. Sundaralingam, S. Tyree, and S. Birchfield. Diff-DOPE: Differentiable Deep Object Pose Estimation. arXiv: 2310.00463 (2023) record the use of the nvdiffrast differential rendering library to achieve pose estimation, and all the disclosed contents are incorporated into the present disclosure.
[0155] Compared with the above rendering technology which focuses on estimating the 6D object pose from a 2D static reference image, some embodiments of the present disclosure can solve the problem of real-time dynamic tracking of multiple target 6D object poses. Therefore, the method of some embodiments of the present disclosure can have two processing stages: 1) static pose calibration of the initial digital twin model; 2) dynamic pose tracking of the digital twin model. The second stage is achieved by continuously optimizing the above loss function by continuously updating the image I with a real camera, so that the digital twin of the rendered image moves in 3D space following the target of the captured image.
[0156] The following uses a 3D printed teapot and teacup, a depth camera, and a differential rendering library as an example. The method of this embodiment achieves real-time tracking of the teapot and teacup at a speed of about 60 FPS. Figure 3 and Figure 4 The optimization process of the 6D pose of the digital twin in the first and second stages is demonstrated. Figure 3 The static calibration process of the digital twin's pose. The leftmost figure shows the initial poses of the real teapot (blue), the real teacup (green) and its digital twin (white ghost). The digital twin model is rendered by the known 3D model, camera parameters and the current estimated pose. From left to right, the method of this embodiment can calibrate the pose of the digital twin to match the real target. Figure 4 This is the real-time posture tracking process of the digital twin. From left to right, the physical postures of the real teapot and the real teacup are changed at the same time. The method of this embodiment can continuously track the posture of the target and realize real-time digital twin reconstruction.
[0157] The target of this embodiment may also include a robot; the posture conversion parameter includes the joint angle of the robot; step S131 may specifically include:
[0158] Based on the preset posture and preset joint angles of the robot, the three-dimensional models of each component of the robot are converted from the object coordinate system of each component to the camera coordinate system by calculating the robot forward kinematics matrix;
[0159] Step S132 may specifically include:
[0160] Based on the preset camera parameters, the converted robot three-dimensional model is subjected to differentiable rendering to obtain the first high-dimensional data.
[0161] Given an input image I (RGB or RGBD), the goal of robot pose estimation is to recover the state of a robot with known joints in a 3D scene. The state of an articulated robot is defined by two key components: (i) the 6D pose P of the robot base base, i.e., the 3D translation and 3D rotation relative to the origin of the camera coordinates, and (ii) the joint angles θ of the robot joints. For an articulated robot with a known structure, its structure is usually represented by links and joint angles, where the joint angles determine the pose of each robot link relative to the origin coordinates of the robot base. By using the 6D pose P of the base base and joint angles θ, the mesh of each robot link can be transformed from the initial pose to the appropriate pose to represent the robot in various configurations.
[0162] This embodiment uses an optimization method similar to that of object pose estimation, but takes into account the joint angle θ. Given the initial 6D pose P of the robot base b ' ase and the initial joint angle θ′ and the robot model M with N links robot = {l n |n∈{1,2,…N}}, first calculate the forward kinematics matrix of the initial joint angle θ′ of each link Then, the transformed robot model can be obtained That is, P b ' ase and a ready-to-render robot mesh with preset joint angles θ′. Depending on the type of robot segmentation algorithm used, a robot mask can be obtained, which can be a single-channel mask combining all links It can also be a multi-channel mask where each channel represents a separate link This selection generates a different reference array A = I⊙I mask . Based on the shape of A, use a differentiable renderer and the camera intrinsic parameters K to render a high-dimensional model Where H′, W′ and C represent the height, width and channels (e.g., RGB, depth) of the rendered image, respectively.
[0163] Forward Kinematics
[0164] Forward kinematics is a sequence of physical transformations that describes the physical transformations available to each component in the robot. The physical pose of a robot component is determined not only by the joint angles or translation phases (relative pose to the neighbor) of its immediate neighbors, but also by the relative pose of that neighbor with respect to its own neighbors. Therefore, it is more intuitive to describe the physical transformation of a robot component as a composite transformation of a sequence of transformations. In computer graphics, such a composite transformation can be simply described as the multiplication of the component geometry point cloud coordinate grid with a 4x4 matrix. A unique 4x4 matrix is usually provided for each independent component of the robot to describe the complete forward kinematics.
[0165] Then, the above method of estimating pose can be used to calculate the corresponding occlusion perception data In addition, the forward kinematics of the robot is differentiable. Therefore, we can transform the parameter P base and joint angles θ or one of them is set as a learnable parameter. By minimizing the error between the true array and the block perception data, the correct robot 6D pose and / or joint angles can be estimated:
[0166]
[0167] Among them, Loss is the same indicator used in the object pose estimation part.
[0168] Similar to the above-mentioned multi-target pose estimation, robot pose estimation can optimize multiple targets or robot arms (or both at the same time), and also consists of two stages: static pose calibration and real-time pose tracking. Therefore, given an input image of a robot arm and a target object containing a known 3D model, the robot and object pose estimation method of this embodiment can first calibrate its corresponding digital twin to the correct pose, and continuously track their poses in real time when the robot arm and the target object move. The following is an example using a differential rendering library with a Universal Robots UR5 model and a teapot. First, a series of images are rendered as the poses of the target robot and the target object. Then, a robot digital twin model is constructed based on the URDF file of the UR5 robot arm, and an attempt is made to optimize its pose using differentiable rendering to match the pre-rendered robot image and the teapot.
[0169] Unified Robotics Description Format (URDF)
[0170] The Unified Robot Description Format (URDF) is an example of a data format that can be used to describe the physical geometry of a robot and its components (in the form of a point cloud mesh), as well as the degrees of freedom available on the robot to achieve its physical structural transformations (such as joint angle rotations and physical translations). Mesh components in a URDF file may include a visual mesh (a large number of triangle mesh polygons for visualization) and a collision mesh (a small number of triangle mesh polygons for efficient collision detection).
[0171] Figure 5 and Figure 6 The calibration and tracking process of the robot and teapot digital twin poses is demonstrated. Figure 5The static calibration process of the pose of the digital twin of the robot and the teapot. The leftmost figure shows the initial pose of the robot and the teapot (in color) and their digital twin (white ghost). The model of the digital twin is rendered by the known 3D model, camera parameters and the current estimated pose. From left to right, the method of this embodiment can calibrate the pose of the digital twin to match the real target. Figure 6 This is the real-time posture tracking process of the digital twin. From left to right, the physical postures of the robot and the teacup are changed at the same time. The method of this embodiment can continuously track the posture of the target and realize real-time digital twin reconstruction.
[0172] The pose estimation method of this embodiment estimates the poses of multiple objects and robots in an iterative manner. The iterative nature of such algorithms usually involves perfecting the setting of the initial pose through multiple optimization steps to achieve the purpose of convergence with the true pose. However, if the initial pose is too far away from the true pose, the optimization process may be limited by the local minimum, or it may take longer to converge, resulting in inaccurate pose estimation and high system latency. On the other hand, choosing a suitable initial pose can significantly accelerate convergence, reduce the computational burden, and improve the overall estimation accuracy. Therefore, choosing a suitable initial pose is crucial to ensuring the accuracy and efficiency of the estimation process.
[0173] In an optional implementation, the preset posture includes preset translation parameters and preset rotation parameters of the multiple targets; before step S12, the method further includes:
[0174] S31, determining preset translation parameters using the point cloud data of each target;
[0175] S32. Using preset image indicators and / or features and / or pre-trained models, calculate the similarity between the template in the template set and the reference image to determine the most matching target template, and determine the preset rotation parameters according to the rotation parameters corresponding to the target template; the template set includes the rotation parameters and the corresponding image, and the template set is generated by performing several uniform samplings on the selected rotation axis. References Nguyen, VN, Hu, Y., Xiao, Y., Salzmann, M., & Lepetit, V. Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (CVPR) (pp. 6771-6780). (2022) describes a pose estimation method based on template matching, and all the contents disclosed are incorporated into the present disclosure. Compared to generating a template by positioning a camera on a hemispherical grid and fixing the CAD model of the object at the center of the sphere, the method of some embodiments of the present disclosure can use real internal parameters and external parameters to form a fixed virtual camera and rotate the CAD model of the object in three-dimensional space using a uniformly sampled rotation matrix.
[0176] Specifically, before step S32, the method further includes:
[0177] Performing uniform sampling several times on the selected rotation axis to generate sampling data of a preset number of samplings;
[0178] Based on preset camera parameters, a three-dimensional model of the target, initial translation parameters and sampling data of a preset sampling quantity, a set of images of different poses of the three-dimensional model is rendered and generated to form a template set.
[0179] Specifically, for a given input image I (RGBD) and a CAD model M of the target, the purpose of roughly estimating the initial pose is to provide a better starting point P close to the global minimum P i , to avoid being limited by local minima and reduce the time cost of the optimization process. The rough initial object pose is composed of a 4×4 rigid transformation matrix P i =[R i ,T i ; 0, 1] means, where R i is a 3×3 rotation matrix, T i is a 3×1 translation vector. i Represented as quaternion qi , which can be easily converted into a 3×3 rotation matrix.
[0180] Initial translation parameter T i It can be obtained from the depth map D of a single camera view: First, the depth map is back-projected onto the point cloud in the camera coordinate system. The depth map is represented as D(u,v), where (u,v) are the pixel coordinates and D(u,v) are the depth values (distance from the camera). The intrinsic parameters of the camera are represented by the camera matrix K, which contains the focal length f x ,f y and the principal point (c x ,c y ). D(u,v) can be converted into 3D coordinates (x,y,z) using the following formula:
[0181]
[0182] z=D(u,v)
[0183] By applying this transformation to each pixel in the depth map, a complete point cloud can be obtained. The 3D point cloud of the corresponding object is extracted using the object segmentation mask. The rough initial translation parameter T i It can be selected from the 3D point cloud of the target, either as a specific point or as the average position of all points in the point cloud, thus providing a reasonable estimate of the initial translation parameters of the object.
[0184] In order to obtain an initial rotation parameter R close to the true value i , this embodiment adopts a similarity matching method between a pre-generated template set and a reference image. The template set consists of rendering target images with different rotation poses, where the rotation quaternion q b As the key value, the corresponding rendered image R b As a numerical value. Figure 7 It shows how to retrieve the rough initial rotation parameters. The following is a detailed description of how to generate the target template mask and match it to the target reference mask:
[0185] First, use Euler angles to uniformly sample in three-dimensional space to obtain a set of rotation quaternions q b , where the number of variations along the x-axis, y-axis, and z-axis is denoted as var i , i∈{x,y,z}. These variants are uniformly distributed along each axis in the range [0, 360) degrees, and the sampling interval Δθ is expressed as where i∈{x,y,z}. The Cartesian product of these rotation angles is created to form a comprehensive 3D mesh containing all possible rotation combinations. The Euler angles are converted to a rotation matrix according to the following formula:
[0186] R=Rz (γ)R y (β)R x (α)
[0187] Among them, R x ,R y and R z are the rotation matrices about the x-axis, y-axis, and z-axis respectively, and α, β, and γ represent the corresponding Euler angles respectively.
[0188] These matrices are then converted to quaternions q b The total number of possible rotations is J = var x ×var y ×var z By incorporating prior knowledge about the geometry of the target object, different vars can be assigned on different axes. i Specifically, for objects with poor directional symmetry, it is necessary to set a larger number of variants on the corresponding axis. According to experience, for objects that lack symmetry along a specific coordinate axis, the number of corresponding variants on the corresponding coordinate axis is usually between 6 and 12, depending on the shape and volume of the object's unique geometric features.
[0189] Figure 7 Schematic diagram of the process of retrieving rough initial rotation parameters, which includes four steps: obtaining reference mask, generating template, feature extraction and similarity matching. Based on the intrinsic parameters of the camera, the CAD model of the target object, and the initial translation parameters And the rotation list Use existing rendering tools to generate a set of images of the CAD model in different poses. b This can include masks, RGB images, or RGB-D images with corresponding masks. A template set is constructed based on the above data, where the rotation quaternion q b As the key value, the corresponding rendered image R b As a numerical value. Both the rendered template and the reference mask are cropped around the object center and resized for better matching.
[0190] The similarity between the rendered template and the reference mask can be calculated using various metrics, such as traditional image similarity metrics (such as mean square error (MSE), intersection over union (IoU), DICE coefficient), manual features (such as SIFT, HOG) or learning-based models (such as ResNet, VGG). In the method of this embodiment, a pre-trained VGG network (without a classification layer) is used to obtain the R b and R i These feature maps are represented as F b and F r, and evaluate the cosine similarity between them by comparison to determine the highest similarity at the feature level. The goal of the rough initial rotation parameter estimation is to find the best match R with the reference mask r The template image R i Then, the rough initial rotation parameters are selected as the rotation quaternion q i . Figure 8 The visualization results of different initial poses with / without initial translation parameters and initial rotation parameters at the beginning of pose estimation are shown, which shows the effectiveness of the initial 6D pose estimation strategy of this embodiment.
[0191] It should be noted that for multiple cameras, feature rendering and matching can be performed separately, but the sum of the highest similarity values is used to determine the shared coarse initial rotation parameter input. At the same time, when the camera is fixed, the camera parameters are usually unchanged, so F can be obtained in a one-time offline preprocessing process. b .
[0192] In addition, during the iterative optimization process, when the actual object moves quickly and the current object posture Far from reaching the expected global minimum P, the gradient descent method may lead to Stuck in local minima. One of the prominent problems is large angle rotations, which can lead to optimization failure for objects that lack geometric symmetry. Fig. 9 Schematic diagram of the large angle rotation problem in the optimization process. The leftmost image is the reference image for pose estimation. The right figure shows the relationship between the rotation angle and the loss value when the translation parameters are perfect. The white ghost shows the pose of the digital twin. From the initial pose (0) to the optimal pose (182), multiple local minimum traps will appear in the optimization process, and the gradient descent method is difficult to solve this problem. To solve this problem, this embodiment uses a batch optimization strategy, that is, optimizing a batch of mesh copies with different initial poses instead of a single mesh.
[0193] Therefore, before step S13, the method further includes:
[0194] S41, obtaining a plurality of replica models of the three-dimensional model of the batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters;
[0195] S42, using the high-dimensional data obtained by rendering the replica model of the batch optimized object to calculate occlusion perception data corresponding to the occlusion information of the batch optimized object;
[0196] S43, calculating the loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object to find a target replica model with the minimum loss value, wherein the target replica model is used to determine a preset pose for this round of pose estimation.
[0197] For each 3D object model in batch optimization mode, it is copied to the original mesh M B To create a batch of mesh copies Where J is the batch size of each 3D model. Each replica is assigned a corresponding initial pose compensation It comes from the collection Set P′ Δ Including translation compensation matrix and the rotation compensation matrix (Also expressed as quaternion ).parameter and It can be a fixed constant or a learnable parameter based on different configurations. Typically, the translation compensation parameter Initialize at the geometric center of the object or the origin of the object coordinate system, and rotate the compensation parameters Similar to the uniform sampling from the three-dimensional space introduced in the template generation section, it is also possible to select from the top K poses in the template set that are most similar to the current reference image. In 6D object pose estimation, since rotation estimation is more challenging than translation estimation, the batch optimization strategy is mainly to improve the accuracy of rotation estimation. The role of translation compensation is to help the rotation batch optimization move in the local space and promote optimization convergence. Based on the different learnability of translation and rotation offsets, this embodiment provides P′ Δ Three different configurations:
[0198] 1) Shared translation, fixed rotation: For all batch sizes J, the translations of all replicas have the same learnable translation compensation The rotation compensation remains unchanged. The posture compensation is expressed as:
[0199]
[0200] In this case, the translation compensation can be viewed as a learnable rotation center for rotating mesh replicas.
[0201] 2) Independent translation, fixed rotation: Each replica has its own set of independently learnable translation parameters, and the rotation compensation remains unchanged. The posture compensation is expressed as:
[0202]
[0203] In this case, each rotated replica is given an additional translational degree of freedom, enabling it to explore a larger local space.
[0204] 3) Fully learnable translation and rotation: and All can be learned independently. Posture compensation is expressed as:
[0205]
[0206] In this case, each mesh copy can be treated as an independent mesh that can be freely translated and rotated in 3D space.
[0207] By selecting different configurations according to the geometric characteristics of the target object, a balance can be achieved between computational cost and computational accuracy. In some embodiments, configuration 1) can be adapted to objects with good rotational symmetry and converge faster; configurations 2) and 3) can be adapted to objects with complex shapes and no symmetry, with a higher upper limit of accuracy, but slower convergence. Fig.10 Visualization of an example of a Teapot grid replica used in a batch optimization strategy. The batch object consists of 6 grid replicas along the Y axis.
[0208] In the context of batch optimization pose estimation, objects optimized in batch mode are called "batch optimized objects", and other objects in the scene that are not optimized in batch mode are called "non-batch optimized objects". Dynamically optimize the replica with the smallest loss value in the batch to get rid of the local minimum trap. First, use the global initial pose and compensation posture P′ Δ The initial copy model M of the batch object dup To convert, the conversion formula is:
[0209]
[0210] And convert the model of non-batch objects, the conversion formula is:
[0211]
[0212] Where N is the number of non-batch optimized objects. Then, using a differentiable renderer and the camera intrinsic parameter K, a high-dimensional rendering model of a batch of objects can be generated And another high-dimensional rendering model for non-batch objects
[0213] However, when there are multiple objects in the scene, occlusion must be considered in the loss computation. Since the pose of an object with replicas may correspond to the pose of any replica in the batch optimization and is unknown, it is very difficult to directly compute the occlusion relationship between the batch objects and other scene objects. This complexity arises from the need to enumerate all possible combinations between all replicas of all batch optimized objects and all non-batch optimized objects in the scene to accurately evaluate the mutual occlusion between them.
[0214] To solve the above problem, in this embodiment, only one target object is set to batch optimization state in one optimization iteration, and occlusion is calculated separately: first, the occlusion-aware high-dimensional data of all copies of the batch optimized object occluded by all other non-batch optimized objects is calculated. Then, the loss between the occlusion-aware high-dimensional data and the reference image of the batch optimized object is calculated, and the copy with the smallest loss value is selected as the best-fit copy. The pose of the best-fit copy can then be used to determine the pose of the batch optimized object, and all non-batch optimized objects can be calculated. The occlusion perception data. The occlusion calculation can refer to the description in the object pose estimation section.
[0215] Specifically, the loss function of a batch of objects can be expressed as:
[0216]
[0217] L dup The minimum loss in and the index of the replica corresponding to the minimum loss can be expressed as:
[0218]
[0219] in, Represents the loss value of each batch optimization (batch) j.
[0220] Then, calculate the loss for the non-batch optimized object:
[0221]
[0222] Optimize the learnable parameter P Δ ,P B ,P NB The total loss of back propagation can be expressed as
[0223] Loss total =L best +L batch +L non-batch
[0224] P Δ ,P B ,P NB=argmin Loss total
[0225] In this way, batch objects with occlusion information can be effectively included in the optimization process. The poses of the batch objects are finally determined as
[0226] In addition, single-camera systems are often subject to limitations such as occlusion and limited field of view, which can seriously affect the accuracy of posture estimation. The use of a multi-camera system can alleviate the above problem, because the multi-camera system can obtain different perspectives of the object and provide complementary information for the posture estimation of multiple objects. Therefore, step S11 may include:
[0227] S111. Use several cameras to perform surround shooting from different angles to reduce the position blur caused by mutual occlusion and self-occlusion of the target;
[0228] S112, using 3D point cloud reprojection to match the real-time target detection data of each camera.
[0229] Specifically, step S112 includes:
[0230] Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects;
[0231] Determining camera parameters, wherein the camera parameters include intrinsic parameters of the first camera and the second camera, and parameters for transforming the field of view of the first camera and the field of view of the second camera;
[0232] Based on the determined camera parameters, the image of the first camera is reprojected into the field of view of the second camera to obtain a reprojected three-dimensional point cloud;
[0233] comparing the mask information of the one or more objects in the image of the second camera with the reprojected three-dimensional point cloud to determine which object's mask information has the greatest overlap with the reprojected three-dimensional point cloud;
[0234] Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
[0235] When using a multi-camera system for object detection, the key is to solve the problem of ID mismatch. When multiple cameras detect the same object, the detection model may assign different IDs to the same object, and the problem of ID mismatch will occur. Since each camera provides an independent perspective, the model may not be able to identify whether the detection results under different perspectives correspond to the same object. For RGB images or RGBD images from a single camera view, existing object detection and segmentation methods output a dictionary containing detected object information, including object ID, category, bounding box, and mask. In order to solve the problem of ID mismatch, it is necessary to use this information to establish the correspondence between object IDs in different camera views to ensure that the same object can be consistently identified in each view. Therefore, it is usually necessary to integrate multi-view object information association techniques, such as triangulation, feature matching, etc. In this embodiment, an object ID matching algorithm based on point cloud or depth reprojection is utilized, and the algorithm includes the following steps:
[0236] 1. Camera calibration
[0237] Intrinsic calibration involves determining the internal parameters of the camera, such as focal length (f x ,f y ), main point (c x ,c y ) and distortion coefficient k 1 ,k 2 ,p 1 ,p 2 These parameters define how the camera projects real 3D image points onto the 2D image plane. This information can usually be loaded from the camera driver API or calibrated using a chessboard-like pattern.
[0238] The extrinsic matrix determines the position and orientation of the camera relative to the world coordinate system, which is usually determined by a calibration plate (ChArUco calibration plate). The extrinsic parameters include the rotation matrix R and the translation vector T, which describe the transformation from the world coordinate system to the camera coordinate system. Given a 3D point (X, Y, Z), its corresponding 2D image point (u, v) can be converted by the following formula:
[0239]
[0240] Among them, (X c ,Y c ,Z c ) is the coordinate of the point in the camera coordinate system. By solving the above formula using the corner points of the ChArUco calibration plate, the rotation matrix and translation vector can be estimated, thereby accurately determining the position and orientation of the camera in the world coordinate system.
[0241] For extrinsic calibration of multiple cameras, after intrinsic calibration, the ChArUco plate is placed in the field of view of all cameras as a common reference object. Each camera captures images of the calibration plate and from these images the pose (in terms of rotation and translation) of each camera relative to the ChArUco calibration plate is estimated. Once the pose of each camera relative to the ChArUco calibration plate is determined, a common world coordinate system can be established by selecting one camera as a reference. The extrinsic parameters of the other cameras are then transformed to the coordinate system of the reference camera, ensuring that all cameras are aligned in a unified spatial frame.
[0242] (II) Point cloud reprojection and ID matching
[0243] Taking two cameras as an example, the process of converting the depth image of camera 1 to a point cloud involves back-projecting the depth values into a 3D coordinate system using the intrinsic parameters of camera 1. The point cloud is transformed into the coordinate system of camera 2 by applying an extrinsic transformation between the two cameras. The 3D points are then reprojected into the image plane of camera 2 using the intrinsic parameters of camera 2, generating the depth information captured by camera 1 in the view of camera 2.
[0244] ID matching is performed after the points from camera 1 are reprojected onto the image plane of camera 2. Figure 1 For each object detected in , the point cloud corresponding to the surface of the target object is extracted using the mask information and the corresponding reprojected points in the camera 2 view are calculated. The reprojected points are then compared with the view Figure 2 By determining the visual Figure 2 The ID of the mask with the highest overlap with the reprojected point cloud establishes the correspondence between the object IDs in the two camera views, ensuring consistent identification of the same object in both views.
[0245] Fig.11 The multi-camera setup (left) and detection results from different viewpoints (right two images) are shown. Fig.12a -c shows the view from the camera Figure 1 To view Figure 2 Schematic diagram of the reprojected 3D point cloud. The green points in the image are the visual Figure 1 The reprojection points of the point cloud of the object surface with different IDs in the perspective of camera 2. Figure 2 The overlap ratio of the occlusion masks detected in the two camera views can determine the correspondence between the object IDs in the two camera views. Figure 1 ) and ID4 ( Figure 2 ) matches, ID2 ( Figure 1 ) and ID1 ( Figure 2 ) matches, ID3 ( Figure 1 ) and ID2 ( Figure 2 ) matches.
[0246] For each object, the ID is known in each camera view, which enables reference images of the object to be obtained from multiple angles. In each camera view, the digital twin can be rendered using shared learnable parameters (representing the object pose), and these rendered images are compared with the corresponding reference images to calculate the loss value for that view. By adding the loss values of all camera views, complementary information from multiple angles is utilized. Gradient descent is then used to optimize the shared learnable parameters, so that the pose of the object can be accurately determined by combining the data from all camera views.
[0247] Real-time pose estimation of this embodiment is critical to achieving accurate adaptive robotic manipulation, especially in dynamic environments where the spatial relationships between objects or between objects and the robot may move or change. For tasks such as pick-and-place, assembly, and bundling, the robot must accurately estimate the pose of individual objects as well as the interactions between objects and the robot. This becomes particularly important when managing multiple objects or handling complex operations such as installing or removing objects on or from the robot end-effector.
[0248] The real-time pose estimation system provides direct input to the robot control system, enabling it to perform precise movements. For example, when grasping an object, the pose data allows the control system to calculate the best approach and manipulation strategy to minimize the risk of collision or slippage. In addition, when the object is attached to the robot, the system can use more advanced manipulation strategies, such as treating the combined structure of the object and the robot as a single entity to avoid collisions and perform path planning. For example, the system can bind the relative poses of some objects and adopt a more optimal control strategy. When the robot grasps the object, it treats the combined structure of the robot and the object as a single entity for path planning and collision avoidance.
[0249] In more complex situations, such as tying multiple objects, the collective pose of the group must be managed while keeping the individual objects relatively stable. This process is further enhanced by real-time feedback from the robot, such as tactile and proprioceptive data, allowing the system to detect and compensate for unexpected events such as slippage, deformation, or external interactions. By combining visual pose tracking and feedback from the robot's sensors, the system can dynamically adjust its control strategy to ensure robustness and precision in the interaction between the objects and the robot. This integration of multimodal data is critical to improving the overall accuracy of manipulation tasks in complex and unpredictable environments.
[0250] Fig.13Schematic diagram of the interaction between the robot and the object during the grasping operation. The white ghost is a digital robot and object rendered in real time using the pose estimation results. From (a) to (b), the robot gradually approaches the handle of the teacup. In (c), the robot successfully grasps the teacup and fixes the object on the robot. From (d) to (e), the pose of the teacup is continuously updated based on its relative position to the robot. In (f), the robot releases the teacup and makes it detach from the robot.
[0251] The method for estimating multi-target poses provided in this embodiment has the following advantages:
[0252] 1. Create digital twins of objects from their high-dimensional data and simultaneously optimize their 6D poses in each iteration, taking into account physical occlusions between objects, and enhance this process by integrating RGBXY derivatives and optimal transport. RGBXY derivatives capture color and spatial variations, enabling robust gradient-based optimization for large object transformations, while optimal transport minimizes the cost function to effectively match RGBXY features, even when there are significant differences in object positioning. As a result, our pose estimation method can reflect the real-time pose of all target objects with high fidelity.
[0253] 2. The initial pose estimation method and batch optimization strategy can ensure fast convergence of the optimization while getting rid of the local minimum trap in the pose optimization process.
[0254] 3. Supports multi-camera stream input, which can alleviate the pose ambiguity problem caused by occlusion and can handle complex scene-level pose estimation.
[0255] 4. Integrating robot pose estimation and object pose estimation into a unified paradigm, which can be applied to real-time digital twin reconstruction, realizes a closed-loop and real-time feedback method for robot control.
[0256] Example 2
[0257] This embodiment provides a robot control method, the method:
[0258] Using the method for estimating multi-target poses in Example 1 to obtain pose information of the robot and pose information of the target object;
[0259] Inputting the posture information of the robot and the posture information of the target object into the control system of the robot;
[0260] Based on the input, the control system controls the interaction of the robot with the target object.
[0261] The robot control method provided in this embodiment can reflect the real-time postures of all target objects with high fidelity by utilizing the above-mentioned method for estimating multi-target postures, and can enable the control system to accurately control the interaction between the robot and the objects with the best strategy, integrating the robot posture estimation and the object posture estimation into a unified paradigm, thereby realizing a closed-loop and real-time feedback method for robot control.
[0262] Example 3
[0263] This embodiment provides a system for estimating multi-target postures, the system comprising: one or more processors, the one or more processors implementing the method for estimating multi-target postures in embodiment 1 or the robot control method in embodiment 2.
[0264] For example, the one or more processors may include one or more GPUs, one or more CPUs, or any other technically feasible combination of one or more processors such as ASIC chips. In some examples, the one or more processors such as ASIC chips may include memory.
[0265] Preferably, the system further comprises one or more cameras for acquiring video data and / or image data of the estimated target.
[0266] Preferably, the system further comprises a robot;
[0267] The system uses the method for estimating multi-target postures in Example 1 to obtain posture information of the robot and posture information of the target object;
[0268] The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
[0269] As for the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The system embodiment described above is only schematic, in which the modules described as separate components may or may not be physically separated, and the modules displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. Ordinary technicians in this field can understand and implement it without paying creative work.
[0270] Example 4
[0271] Fig.14This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the method of Embodiment 1 or Embodiment 2 is implemented when the processor executes the program. Fig.14 The electronic device 30 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0272] like Fig.14 As shown, the electronic device 30 may be in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 30 may include, but are not limited to: at least one processor 31, at least one memory 32, and a bus 33 connecting different system components (including the memory 32 and the processor 31).
[0273] The bus 33 includes a data bus, an address bus, and a control bus.
[0274] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .
[0275] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0276] The processor 31 executes various functional applications and data processing by running the computer programs stored in the memory 32, such as the method of embodiment 1 or embodiment 2 of the present disclosure.
[0277] The electronic device 30 may also communicate with one or more external devices 34 (e.g., keyboards, pointing devices, etc.). Such communication may be performed via an input / output (I / O) interface 35. Furthermore, the model-generated device 30 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. As shown, the network adapter 36 communicates with other modules of the model-generated device 30 via a bus 33. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the model-generated device 30, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0278] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.
[0279] Example 5
[0280] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method of Embodiment 1 or Embodiment 2 is implemented.
[0281] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.
[0282] In a possible implementation manner, the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the method for implementing Embodiment 1 or Embodiment 2.
[0283] Among them, the program code for executing the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed completely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or completely on the remote device.
[0284] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, but these changes and modifications all fall within the protection scope of the present disclosure.
Claims
1. A method for estimating multi-target poses, characterized in that: The method comprises: Acquire real-time video data, wherein the real-time video data includes first image frames of the multiple targets; Based on the first image frame, obtaining first high-dimensional segmented pixel data, wherein the first high-dimensional segmented pixel data includes first pixel data of each target, and the first pixel data is a three-dimensional or higher-dimensional array; Based on a first preset posture, performing differentiable rendering on the three-dimensional model of the multi-target to obtain first high-dimensional data, where the first high-dimensional data is a three-dimensional or higher-dimensional array; Based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data, the first preset posture is updated to obtain a first target posture.
2. The method for estimating multiple target poses as claimed in claim 1, characterized in that: The real-time video data also includes at least a second image frame of the multiple targets; After the step of obtaining the target posture, the method further includes: Based on the second image frame, obtaining second high-dimensional segmented pixel data; Based on the first target pose, performing differentiable rendering on the three-dimensional model of the multiple targets to obtain second high-dimensional data, where the second high-dimensional data is a three-dimensional or higher-dimensional array; Based on minimizing the error between the second high-dimensional segmented pixel data and the second high-dimensional data, updating the first target pose to obtain a second target pose; Based on other image frames of the real-time video data, the posture corresponding to the previous image frame is updated to obtain other target postures.
3. The method for estimating multiple target poses as claimed in claim 1, characterized in that: The step of obtaining first high-dimensional segmented pixel data based on the first image frame comprises: Using a segmentation algorithm to segment each object in the first image frame to obtain a mask array; The first high-dimensional segmented pixel data is obtained based on the mask array and the channel data of the first image frame.
4. The method for estimating multiple target poses as claimed in claim 1, characterized in that: The step of performing differentiable rendering on the three-dimensional model of the multi-target based on the first preset posture to obtain first high-dimensional data includes: Based on the preset poses of the multiple targets, transforming the three-dimensional models of the multiple targets from an object coordinate system to a camera coordinate system according to learnable pose transformation parameters, wherein the pose transformation parameters are used to characterize the real-time poses of the targets; Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model to obtain the first high-dimensional data.
5. The method for estimating multiple target positions as claimed in claim 4, characterized in that: After the step of performing differentiable rendering on the three-dimensional model of the multiple targets to obtain the first high-dimensional data, the method further includes: Calculating occlusion perception data corresponding to the occlusion information of the target using the depth information and contour information of the first high-dimensional data; The step of updating the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises: The gradient descent method is used to iteratively minimize the pixel difference between the first high-dimensional segmented pixel data and the occlusion-aware data to update the learnable posture transformation parameters.
6. The method for estimating multiple target poses as claimed in claim 4, characterized in that: The step of performing differentiable rendering on the converted three-dimensional model to obtain the first high-dimensional data further includes: The converted three-dimensional model is subjected to the differentiable rendering to obtain first high-dimensional data having a variable contour, wherein a range of the variable contour is determined based on a degree of difference between the first pixel data and a preset posture.
7. The method for estimating multiple target poses as claimed in claim 4, characterized in that: Performing the differentiable rendering on the converted three-dimensional model to obtain first high-dimensional data includes: Based on the first RGBXY derivative, the converted three-dimensional model is subjected to the differentiable rendering, and the first high-dimensional data and the first high-dimensional segmented pixel data are matched and mapped between pixels through optimal transmission to obtain a second RGBXY derivative of the first high-dimensional segmented pixel data, wherein the first RGBXY derivative and the second RGBXY derivative both include information of color change and spatial position change.
8. The method for estimating multiple target positions as claimed in claim 7, characterized in that: The updating of the first preset posture to obtain a first target posture based on minimizing the error between the first high-dimensional segmented pixel data and the first high-dimensional data comprises: Based on optimal transmission, the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data is minimized; based on minimizing the pixel level difference between the second RGBXY derivative of the first high-dimensional segmentation pixel data and the first RGBXY derivative of the first high-dimensional data, the first preset posture is updated to obtain the first target posture.
9. The method for estimating multiple target positions as claimed in claim 4, characterized in that: The target includes a robot; the posture conversion parameters include joint angles of the robot; The step of converting the three-dimensional model of the multiple targets to camera coordinates based on the preset poses of the multiple targets according to the learnable pose conversion parameters comprises: Based on the preset posture and preset joint angles of the robot, the three-dimensional model of each component of the robot is converted from the object coordinate system of each component to the camera coordinate system by calculating the forward kinematics matrix of the robot; The step of performing differentiable rendering on the converted three-dimensional model based on preset camera parameters to obtain the first high-dimensional data includes: Based on preset camera parameters, differentiable rendering is performed on the converted three-dimensional model of the robot to obtain the first high-dimensional data.
10. The method for estimating multiple target positions and postures according to claim 4, wherein: The preset poses include preset translation parameters and preset rotation parameters of the multiple targets; Before the step of converting the three-dimensional model of the multiple targets from the object coordinate system to the camera coordinate system according to the learnable pose conversion parameters, the method further includes: Determine preset translation parameters using point cloud data of each target; Using preset image indicators and / or features and / or pre-trained models, calculate the similarity between the templates in the template set and the reference image to determine the most matching target template, and determine the preset rotation parameters according to the rotation parameters corresponding to the target template; The template set includes rotation parameters and corresponding images, and is generated by performing a plurality of uniform samplings on a selected rotation axis.
11. The method for estimating multiple target positions and postures according to claim 10, wherein: Before calculating the similarity between the templates in the template set and the reference image to determine the most matching target template, the method further includes: Performing uniform sampling several times on the selected rotation axis to generate sampling data of a preset number of samplings; Based on preset camera parameters, a three-dimensional model of the target, initial translation parameters and a preset number of sampling data, a set of images of different positions and postures of the three-dimensional model is rendered to form the template set.
12. The method for estimating multiple target positions and postures according to claim 10, wherein: Before the step of performing differentiable rendering on the three-dimensional model of the multiple objects to obtain the first high-dimensional data, the method further includes: Acquire a plurality of replica models of the three-dimensional model of the batch optimized object, wherein the plurality of replica models have different preset rotation parameters and different translation parameters and / or rotation parameters; Calculate the occlusion perception data corresponding to the occlusion information of the batch optimized objects using the high-dimensional data obtained by rendering the replica model of the batch optimized objects; The loss value between the high-dimensional segmented pixel data and the occlusion perception data of the batch optimized object is calculated to find the target replica model with the minimum loss value. The target replica model is used to determine the preset pose for this round of pose estimation.
13. The method for estimating multiple target positions as claimed in claim 1, wherein: The step of acquiring real-time video data comprises: Use several cameras to perform surround shooting from different angles to reduce the pose ambiguity caused by mutual occlusion and self-occlusion of the target; The real-time target detection data of each camera are matched with each other by using three-dimensional point cloud reprojection to obtain the real-time video data.
14. The method for estimating multiple target positions and postures according to claim 13, wherein: The step of using the three-dimensional point cloud reprojection to match the real-time target detection data of each camera includes: Obtain a first image from a first camera and a second image from a second camera, wherein the first image includes one or more objects and the second image includes one or more objects; Determining camera parameters, the camera parameters including intrinsic parameters of the first camera and the second camera, and extrinsic parameters for transforming a field of view of the first camera and a field of view of the second camera; Based on the determined camera parameters, reprojecting the depth image or three-dimensional point cloud data of the first camera into the field of view of the second camera to obtain a corresponding reprojected two-dimensional point group; comparing the mask information of the one or more objects in the image of the second camera with the reprojected two-dimensional point group to determine which object's mask information has the greatest overlap with the reprojected two-dimensional point group; Based on the maximum overlap, it is determined that the identity information of the object in the image of the first camera matches the identity information of the object in the image of the second camera.
15. A robot control method, characterized in that: The method: Acquire the posture information of the robot and the posture information of the target object by using the method for estimating multi-target postures as described in any one of claims 1 to 14; Inputting the posture information of the robot and the posture information of the target object into the control system of the robot; Based on the input, the control system controls the interaction of the robot with the target object.
16. A computer program product, characterized in that It comprises a computer program, which, when executed by a processor, implements the method for estimating multi-target poses as described in any one of claims 1 to 14 or the method for controlling a robot as described in claim 15.
17. A system for estimating multiple target poses, characterized in that: The system comprises: One or more processors, wherein the one or more processors implement the method for estimating multi-target poses as described in any one of claims 1 to 14 or the robot control method as described in claim 15.
18. The system for estimating multiple target poses according to claim 17, wherein: The system also includes one or more cameras for acquiring video data and / or image data of the estimated target.
19. The system for estimating multiple target poses according to claim 17, wherein: The system also includes a robot; The system obtains the posture information of the robot and the posture information of the target object by using the method for estimating multi-target postures according to any one of claims 1 to 14; The system controls the interaction between the robot and the target object according to the posture information of the robot and the posture information of the target object.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the method for estimating multi-target poses as described in any one of claims 1 to 14 or the method for controlling a robot as described in claim 15.
21. A computer readable medium having computer instructions stored thereon, characterized in that: When executed by a processor, the computer instructions implement the method for estimating multi-target poses as described in any one of claims 1 to 12 or the method for controlling a robot as described in claim 15.
Citation Information
Patent Citations
Differentiable rendering pipeline for inverse graphics
US10565747B2