Robot pose estimation method and apparatus, robot, device, and medium

CN122780392APending Publication Date: 2026-09-18SHENZHEN SYBORG ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610910518.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]基于此,有必要针对上述技术问题,提供一种机器人位姿估计方法、装置、机器人、设备及介质,以解决现有技术中位姿估计易受干扰、易抖动的问题

Benefits of technology

[0010] The aforementioned robot pose estimation method, apparatus, robot, device, and medium simultaneously acquire color and depth images of the current frame; perform instance-level target detection and segmentation on the color image to generate a target list containing target masks; determine the target to be tracked in response to the user's selection operation on the target list; perform cross-frame tracking association on the target list, assign tracking identifiers to the target to be tracked, and establish a mapping relationship; based on the mapping relationship and the target mask, extract the region corresponding to the target to be tracked from the color image and the depth image respectively to generate a local region; combine the local region with a pre-registered 3D object model to obtain the initial pose of the target to be tracked; and recursively optimize the initial pose to obtain the target pose. This invention breaks down the barrier between macroscopic interactive intent and microscopic algorithm trajectory by establishing a mapping relationship between user session identifiers and underlying tracking identifiers, enabling simultaneous tracking of multiple targets. Furthermore, it utilizes recursive optimization (such as iterative error state Kalman filtering) to fuse historical temporal priors with current depth observations (such as robust features like geometric centroids) to perform iterative closed-loop correction of the error state. This not only eliminates random errors in single-frame estimation, achieving millimeter-level precision, but also endows the output pose with extremely strong temporal smoothness, completely eliminating pose jitter. It can directly meet the demands of stringent applications such as high-precision robotic arm grasping and high-fidelity AR assembly. This invention constructs a complete technical closed loop of "2D precise perception—cross-frame robust tracking—3D coarse-to-fine estimation," fundamentally solving the problems of 6DoF pose estimation being susceptible to interference, jitter, and high computational consumption in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780392A_ABST
    Figure CN122780392A_ABST
Patent Text Reader

Abstract

The present application relates to the field of robots, and particularly relates to a robot pose estimation method and device, a robot, equipment and a medium. The method comprises: acquiring a color image and a depth image of a current frame; generating a target list containing a target mask; in response to a selection operation of a user on the target list, determining a target to be tracked; performing cross-frame tracking association on the target list, assigning a tracking identifier to the target to be tracked and establishing a mapping relationship; based on the mapping relationship and the target mask, extracting a region corresponding to the target to be tracked to generate a local region; combining the local region with a pre-registered three-dimensional object model to obtain an initial pose of the target to be tracked; and recursively optimizing the initial pose to obtain a target pose. The present application can track multiple targets simultaneously, eliminates random errors of single-frame estimation, and achieves millimeter-level extreme precision. The present application fundamentally solves the problems of 6DoF pose estimation in a complex scene, such as being easily disturbed, easily dithered and large computational power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robotics, and more particularly to a robot pose estimation method, apparatus, robot, device, and medium. Background Technology

[0002] With the rapid development of industrial automation, logistics sorting, and home service robots, robots are expected to perform dexterous grasping and manipulation tasks in increasingly dynamic and open environments. The core prerequisite for achieving such tasks is that the robot can perceive the position (X, Y, Z) and attitude (yaw, pitch, roll) of the target object in three-dimensional space in real time and accurately, i.e., 6DoF pose estimation.

[0003] Traditional 6DoF pose estimation techniques for robots mainly rely on two types of methods: template matching-based methods and point cloud registration-based methods. Template matching-based methods determine the pose by comparing the 2D image features of the target object from different viewpoints with a preset template. While computationally efficient, their discriminative power and stability significantly decrease when the object is occluded, lighting changes drastically, or the background is highly cluttered, leading to estimation failure. Point cloud registration-based methods, such as the classic Iterative Closest Point algorithm, require a precise 3D computer-aided design model of the target object as a priori, and iteratively match the real-time acquired point cloud data with the model. Although these methods can achieve high accuracy under certain conditions, their iterative optimization process is computationally intensive and time-consuming, and their performance is heavily dependent on the quality of the point cloud (e.g., noise level, occlusion degree, and point cloud density). When dealing with objects without texture, with weak texture, or with high symmetry, they are prone to getting trapped in local optima due to incorrect point correspondences, leading to registration failure. This makes it difficult to guarantee real-time performance and reliability in dynamic grasping scenarios. Summary of the Invention

[0004] Therefore, it is necessary to provide a robot pose estimation method, device, robot, equipment, and medium to address the aforementioned technical problems and solve the issues of pose estimation being susceptible to interference and jitter in the prior art.

[0005] A robot pose estimation method includes: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

[0006] A robot pose estimation device, comprising: The image acquisition module is used to simultaneously acquire the color image and depth image of the current frame; The target to be tracked module is used to perform instance-level target detection and segmentation on the color image, generate a target list containing target masks, and determine the target to be tracked in response to the user's selection operation on the target list. The mapping module is used to perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; The local region module is used to extract the region corresponding to the target to be tracked from the color image and the depth image respectively, based on the mapping relationship and the target mask, to generate a local region; The initial pose module is used to combine the local region with the pre-registered 3D object model to obtain the initial pose of the target to be tracked. The target pose module is used to recursively optimize the initial pose to obtain the target pose.

[0007] A robot includes a controller for performing a pose estimation method as follows: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

[0008] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the robot pose estimation method described above when executing the computer-readable instructions.

[0009] One or more readable storage media storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the robot pose estimation method described above.

[0010] The aforementioned robot pose estimation method, apparatus, robot, device, and medium simultaneously acquire color and depth images of the current frame; perform instance-level target detection and segmentation on the color image to generate a target list containing target masks; determine the target to be tracked in response to the user's selection operation on the target list; perform cross-frame tracking association on the target list, assign tracking identifiers to the target to be tracked, and establish a mapping relationship; based on the mapping relationship and the target mask, extract the region corresponding to the target to be tracked from the color image and the depth image respectively to generate a local region; combine the local region with a pre-registered 3D object model to obtain the initial pose of the target to be tracked; and recursively optimize the initial pose to obtain the target pose. This invention breaks down the barrier between macroscopic interactive intent and microscopic algorithm trajectory by establishing a mapping relationship between user session identifiers and underlying tracking identifiers, enabling simultaneous tracking of multiple targets. Furthermore, it utilizes recursive optimization (such as iterative error state Kalman filtering) to fuse historical temporal priors with current depth observations (such as robust features like geometric centroids) to perform iterative closed-loop correction of the error state. This not only eliminates random errors in single-frame estimation, achieving millimeter-level precision, but also endows the output pose with extremely strong temporal smoothness, completely eliminating pose jitter. It can directly meet the demands of stringent applications such as high-precision robotic arm grasping and high-fidelity AR assembly. This invention constructs a complete technical closed loop of "2D precise perception—cross-frame robust tracking—3D coarse-to-fine estimation," fundamentally solving the problems of 6DoF pose estimation being susceptible to interference, jitter, and high computational consumption in complex scenarios. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an application environment for a robot pose estimation method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a robot pose estimation method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a robot pose estimation device according to an embodiment of the present invention; Figure 4This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] The robot pose estimation method provided in this embodiment can be applied to, for example, Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0015] In one embodiment, such as Figure 2 As shown, a robot pose estimation method is provided, which is then applied to... Figure 1 Taking the server-side as an example, the explanation includes the following steps: S10. Synchronously acquire the color image and depth image of the current frame.

[0016] In essence, a color image refers to a three-channel image captured by an RGB camera, containing rich information about the object's surface texture, color, and edge contours. A depth image, on the other hand, refers to a single-channel image captured by a depth sensor (such as a structured light or ToF camera), where each pixel value represents the physical distance from that point to the origin of the camera's coordinate system, containing information about the object's three-dimensional spatial geometry. Specifically, by using a calibrated camera, aligned visual appearance data and spatial distance data are acquired in real time during system operation, ensuring a strict one-to-one correspondence between color pixels and depth pixels in spatial coordinates. This avoids "artifacts" or data misalignment caused by object movement due to time differences, providing data consistency assurance for subsequent high-precision mask mapping and point cloud generation.

[0017] S20. Perform instance-level target detection and segmentation on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, determine the target to be tracked.

[0018] Understandably, by performing instance-level object detection and segmentation on a color image, not only are the categories of objects in the image detected and identified (detection), but the outline of each individual object is also delineated at the pixel level (segmentation) to distinguish different individuals of the same category (instance level). The target mask is a binary image (or matrix) with the same resolution as the original image. For example, pixels belonging to the target region are assigned a value of 1, and the background is assigned a value of 0, used to accurately represent the area occupied by the target in the image. The target list is the set of all detected and segmented instances in the current frame, each instance containing at least its category label, confidence score, and target mask. The target to be tracked refers to a specific object specified by the user or interactively selected from numerous detected targets, which the system needs to continuously monitor and locate in subsequent frames. Preferably, the target to be tracked can be one or more, depending on the user's selection. That is, at least one target to be tracked is determined. Here, by having the user select the target to be tracked, interactive intent injection is achieved, giving the system the decision-making ability to handle complex scenes (e.g., multiple similar objects crowding together), avoiding the algorithm blindly tracking incorrect targets.

[0019] Optionally, in step S20, i.e., performing instance-level target detection and segmentation on the color image to generate a target list containing target masks, the following steps are included: S201. The color image is subjected to object detection using an instance segmentation network to obtain category information for multiple detected objects. Understandably, an instance segmentation network is a multi-task deep learning model (such as YOLOv8-seg), which internally includes a feature extraction backbone, an object detection head, and a mask segmentation head, capable of parallel processing of object localization, classification, and pixel-level segmentation. Object detection refers to the technique of locating the approximate position of an object of interest in an image (usually represented by a rectangular bounding box) and identifying its category. Category information refers to the semantic classification label of the detected object (such as "cup," "wrench," "screw," etc.), usually represented by discrete category IDs or confidence vectors.

[0020] S202. Based on the category information, perform instance segmentation on multiple detected targets to obtain target masks for each detected target. Instance segmentation, as understood, refers to further outlining the precise contour of each detected target at the pixel level, based on target detection, and being able to distinguish different individuals (i.e., different instances) of the same type in the image. The target mask is a binary matrix with the same resolution as the input image or the same size as the detection box region, where pixel values ​​inside the target object's contour are set to 1 (or True), and external background pixel values ​​are set to 0 (or False), used to accurately represent the shape occupied by the target.

[0021] S203. Generate the target list based on the category and target mask of each of the detected targets. Understandably, the target list is a structured data collection (such as an array or linked list), where each element (entry) corresponds to an instance target detected in the scene, encapsulating all attribute information of that instance target.

[0022] In steps S201-S203, a leap from unstructured visual perception to structured data representation is achieved. The target list unifies and manages scattered category and mask information, providing a clear and easily indexed data interface for downstream interaction modules (such as user selection operations) and tracking modules, enabling the system to directly obtain the complete attributes of any target by traversing the list.

[0023] S30. Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked, and establish mapping relationships.

[0024] In essence, cross-frame tracking association refers to a temporal data association technique that concatenates detection boxes / masks belonging to the same physical entity in different frames by calculating the appearance similarity (such as mask IoU, feature distance) or motion consistency between targets in consecutive frames. A Track ID is a globally unique identifier assigned to a target to be tracked, used to uniquely identify a motion trajectory throughout the tracking lifecycle. It should be understood that after a new target list is input for each frame, the tracker extracts the appearance or motion features of each target and calculates similarity with the historical trajectory set maintained in the previous frame. For targets that successfully match a historical trajectory, the system allows them to inherit the original track ID; for targets appearing for the first time and not matched with a historical trajectory, the system initializes and assigns a completely new track ID.

[0025] For example, in each current frame, the detector outputs a list of targets containing information about multiple instances, each instance containing a bounding box. Category label c, confidence score and pixel-level segmentation mask Information such as confidence level is collected. Detected instances are categorized into high-scoring and low-scoring instances. Then, trajectory state prediction is performed for each instance: before a new frame arrives, each trajectory's predicted bounding box in the current frame is predicted using a uniform motion model. ,in The bounding box of the previous frame. To estimate velocity, Δt is the frame interval. This prediction is used for subsequent data association. Data association includes primary and secondary association. Primary association involves matching high-resolution instances with trajectories, first associating all high-resolution instances with the predicted trajectory bounding boxes. IoU distance is used to measure positional consistency; successfully matched trajectory-detection pairs update their states (position, features, liveness count, etc.), while unmatched trajectories enter a "mismatch count" state. Secondary association involves performing a second match between the unmatched trajectories from the first level (i.e., targets that may not be covered by high-resolution detections due to occlusion) and low-resolution instances. After matching, a motion trajectory management strategy is executed. This strategy includes new trajectory creation (unmatched high-resolution detections are treated as new targets, and new trajectories are created), trajectory destruction (trajectories with more than a threshold (30 frames) of consecutive mismatch frames are destroyed), and trajectory restoration (if a low-resolution detection successfully matches a trajectory that is about to be destroyed, its active state is restored).

[0026] The mapping relationship refers to the two-way binding between the session identifier at the user interaction level and the Track ID of the underlying tracking algorithm. It should be understood that for a user-selected target to be tracked, if it is consistently matched successfully, the system assigns or inherits a fixed tracking identifier for it and maintains the "session identifier" in memory. A mapping table for "tracking identifiers" is used. In this way, even if the target to be tracked is briefly occluded, deformed, or similar objects overlap, the tracking association can maintain its identity ID, solving the ID jump problem caused by independent calculations for each frame in simple detection.

[0027] Optionally, in step S30, namely, performing cross-frame tracking association on the target list, assigning tracking identifiers to the targets to be tracked and establishing mapping relationships, the following steps are included: S301. Obtain the session identifier of the target to be tracked; understandably, the session identifier is a unique logical number dynamically assigned by the interactive system when the user performs a selection operation. It belongs to the system application layer concept, represents the starting point of a "user-specified tracking task", and is bound to the user's operation intention. When the system responds to the user's selection operation on the target list (such as clicking on a target mask on the screen), the interactive module captures the event, generates or assigns a unique session identifier for the selected target, and stores it in the system's state machine for subsequent use by the mapping module.

[0028] S302. Perform cross-frame tracking association on the target list and assign a tracking identifier to the target to be tracked. Cross-frame tracking association is a temporal data association technique that calculates the similarity (e.g., appearance feature distance, bounding box intersection-union ratio, IoU) between the detected target in the current frame (an instance in the target list) and existing motion trajectories in historical frames, matching and concatenating observations of the same physical entity in different frames. A tracking identifier is a globally unique number assigned by the underlying tracking algorithm (e.g., ByteTrack). It is an algorithm-level concept that persists throughout the target's entire lifecycle before it is lost or occluded, uniquely identifying a continuous motion trajectory within the tracker.

[0029] It should be understood that after a new target list is input in each frame, the tracker extracts the appearance or motion features of each target and calculates the similarity with the historical trajectory set maintained in the previous frame. For targets that successfully match a historical trajectory, the system allows them to inherit the tracking identifier of the original trajectory; for targets that appear for the first time and do not match a historical trajectory, the system initializes and assigns a completely new tracking identifier.

[0030] S303. Establish a mapping relationship between the target to be tracked based on the session identifier and the tracking identifier. The mapping relationship refers to the bidirectional binding relationship between the application layer's session identifier and the algorithm layer's tracking identifier (such as a key-value pair mapping in a hash table). It establishes an index channel between macroscopic interaction intentions and microscopic algorithmic trajectories.

[0031] In steps S301-S303, the data barrier between the human-computer interaction layer and the underlying algorithm layer is broken down. The mapping relationship allows the user's initial choice (session identifier) ​​to seamlessly penetrate the complex underlying tracking logic (tracking identifier iteration, prediction, and recovery), achieving continuous transmission of intent across frames. This enables the system to instantly retrieve the latest and most accurate tracking identifier in the current frame using only the unchanging session identifier during subsequent region extraction, thereby accurately locking onto the target.

[0032] S40. Based on the mapping relationship and the target mask, extract the region corresponding to the target to be tracked from the color image and the depth image respectively to generate a local region.

[0033] Understandably, a local region refers to a local image patch cropped from a full-size image with the target mask as the smallest bounding rectangle, including local color patches and local depth patches.

[0034] Specifically, the tracking identifier corresponding to the target in the current frame is found through a mapping relationship, thereby retrieving the latest mask of the target in the current frame (the target mask of the current frame). Then, the bounding box of this mask is calculated. Finally, this bounding box is used to perform cropping operations on both the color image and the depth image, and the mask itself is applied to the cropped patches to filter out a small amount of background within the bounding box, generating a clean local region. This drastically reduces the search space of subsequent pose estimation algorithms from the entire image (e.g., 1920×1080) to a local region (e.g., 200×200), greatly reducing computational load and ensuring real-time performance. Simultaneously, it eliminates a large number of irrelevant background and interfering objects around the target, allowing subsequent pose calculations to focus on the target's own features, significantly improving robustness.

[0035] Optionally, in step S40, that is, based on the mapping relationship and the target mask, extracting the region corresponding to the target to be tracked from the color image and the depth image respectively to generate a local region, includes: S401. Based on the mapping relationship, the target mask corresponding to the target to be tracked is selected from subsequent frames and used as the target mask of the current frame. Understandably, subsequent frames refer to each new image in the video stream continuously acquired by the system after the user completes the target selection operation. The target mask of the current frame refers to the mask of the "target to be tracked" selected by the user through the mapping relationship in the latest frame.

[0036] It should be understood that the target to be tracked may move, deform, or change viewpoint in the video stream, and its mask will be different in each frame. By filtering through mapping relationships, it is ensured that the system always extracts the outline of the target in its "latest state," rather than the outdated outline of the initial frame. This avoids misalignment between the mask and the actual image caused by target movement, and provides a dynamic tracking template for subsequent accurate cropping.

[0037] S402. Based on the target mask of the current frame, crop local color patches and local depth patches corresponding to the target to be tracked from the color image and the depth image, respectively. In essence, calculate the minimum bounding rectangle of the target mask of the current frame. Then, using the coordinate parameters of this bounding rectangle, perform matrix slicing operations on the full-resolution color image and depth image of the current frame, respectively, to extract image data within the rectangular region. Further, the target mask of the current frame can be applied to the cropped patches, setting the pixel values ​​outside the mask to zero to completely mask the background, ultimately generating clean local color patches and local depth patches.

[0038] S403. Generate the local region based on the local color patch and the local depth patch. This understandably represents a leap from separate perception to joint representation. Pose estimation (especially 6DoF pose) requires both texture features to identify pose and depth features to anchor space. The generation of the local region tightly integrates the 2D appearance with the 3D geometric content, eliminating the overhead of asynchronous multimodal data calls and providing a one-stop, highly cohesive input interface for subsequent pose estimation algorithms.

[0039] S50. Combine the local region with the pre-registered 3D object model to obtain the initial pose of the target to be tracked.

[0040] Understandably, a pre-registered 3D object model refers to a complete 3D point cloud or mesh model of the target object obtained in advance through 3D scanning or CAD modeling, with its coordinate system being the object coordinate system. Initial 6DoF Pose refers to the six degrees of freedom parameters describing the object's spatial state in the camera coordinate system, including a three-dimensional rotation vector (pose) and a three-dimensional translation vector (position). Initial pose refers to a coarse pose that has not been refined by optimization algorithms.

[0041] Specifically, local color tiles (providing texture cues), local depth tiles (providing geometric contours), and a pre-stored 3D object model are input into a pose estimator (such as the FoundationPose algorithm). The pose estimator then solves for the rotation and translation matrices that best align the current observation with the 3D model, i.e., the initial pose. This achieves a transition from 2D image pixels to 3D spatial coordinates, giving the target absolute coordinates in the physical world. Furthermore, deep learning-based pose estimation algorithms can quickly provide a globally approximate correct pose assumption, avoiding the risk of traditional optimization algorithms (such as ICP) getting trapped in local optima or failing to converge.

[0042] S60. Recursively optimize the initial pose to obtain the target pose.

[0043] In essence, recursive optimization refers to an iterative approximation algorithm based on a temporal state-space model. It utilizes the observation data of the current frame, combined with the system state prediction from the previous frame, to gradually approximate the true value through repeated closed-loop error correction. The target pose refers to the final 6DoF pose result with high precision and high smoothness output after the recursive optimization algorithm filters out noise and corrects residuals.

[0044] Optionally, in step S60, i.e., recursively optimizing the initial pose to obtain the target pose, the following steps are included: S601. Back-project the depth information in the local area to generate a three-dimensional point cloud, and use the geometric centroid of the three-dimensional point cloud as the pose observation value. S602. Based on the iterative error state Kalman filter and the pose observations, the initial pose is recursively optimized to obtain the target pose.

[0045] Specifically, the initial pose is used as the prior state. A 3D point cloud is generated by backprojection of local depth tiles. The geometric centroid of the 3D point cloud is calculated and used as the pose observation value. In the recursive optimization stage, based on the prior state, the pre-registered 3D object model is projected onto the camera coordinate system to generate a virtual model point cloud. By matching the 3D point cloud with the virtual model point cloud, the difference between the observed geometric centroid (pose observation value) and the model geometric centroid is calculated to construct the position observation residual. Simultaneously, after centroidalization of the matched point pairs, the pure angular deviation is extracted to construct the rotation residual of the point cloud registration. The position observation residual and the rotation residual are combined to construct a complete nonlinear observation equation. The nonlinear observation equation is then subjected to a first-order Taylor expansion and iterative solution using the iterative error state Kalman filter algorithm to obtain the optimal estimate of the error state. This error state is then used to compensate for the initial pose, ultimately outputting a high-precision target pose.

[0046] In steps S10-S60, the color image and depth image of the current frame are acquired synchronously; instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; cross-frame tracking association is performed on the target list, a tracking identifier is assigned to the target to be tracked, and a mapping relationship is established; based on the mapping relationship and the target mask, the region corresponding to the target to be tracked is extracted from the color image and the depth image respectively to generate a local region; the initial pose of the target to be tracked is obtained by combining the local region with a pre-registered 3D object model; the initial pose is recursively optimized to obtain the target pose. This invention breaks down the barrier between macroscopic interactive intent and microscopic algorithm trajectory by establishing a mapping relationship between user session identifiers and underlying tracking identifiers, enabling simultaneous tracking of multiple targets. Furthermore, it utilizes recursive optimization (such as iterative error state Kalman filtering) to fuse historical temporal priors with current depth observations (such as robust features like geometric centroids) to perform iterative closed-loop correction of the error state. This not only eliminates random errors in single-frame estimation, achieving millimeter-level precision, but also endows the output pose with extremely strong temporal smoothness, completely eliminating pose jitter. It can directly meet the demands of stringent applications such as high-precision robotic arm grasping and high-fidelity AR assembly. This invention constructs a complete technical closed loop of "2D precise perception—cross-frame robust tracking—3D coarse-to-fine estimation," fundamentally solving the problems of 6DoF pose estimation being susceptible to interference, jitter, and high computational consumption in complex scenarios.

[0047] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0048] In one embodiment, a robot pose estimation device is provided, which corresponds one-to-one with the robot pose estimation methods described in the above embodiments. For example... Figure 3 As shown, the robot pose estimation device includes an image acquisition module 10, a target tracking module 20, a mapping relationship module 30, a local region module 40, an initial pose module 50, and a target pose module 60. Detailed descriptions of each functional module are as follows: Image acquisition module 10 is used to simultaneously acquire the color image and depth image of the current frame; The target tracking module 20 is used to perform instance-level target detection and segmentation on the color image, generate a target list containing target masks, and determine the target to be tracked in response to the user's selection operation on the target list. The mapping module 30 is used to perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked, and establish mapping relationships. Local region module 40 is used to extract the region corresponding to the target to be tracked from the color image and the depth image respectively, based on the mapping relationship and the target mask, to generate a local region; The initial pose module 50 is used to combine the local region with the pre-registered 3D object model to obtain the initial pose of the target to be tracked. The target pose module 60 is used to recursively optimize the initial pose to obtain the target pose.

[0049] Optionally, the target module to be tracked includes: The category unit is used to perform target detection on the color image using an instance segmentation network to obtain category information of multiple detected targets; A masking unit is used to perform instance segmentation on multiple detected targets based on the category information to obtain a target mask for each detected target; The list unit is used to generate the target list based on the category of each of the detected targets and the target mask.

[0050] Specific limitations regarding the robot pose estimation device can be found in the limitations of the robot pose estimation method described above, and will not be repeated here. Each module in the aforementioned robot pose estimation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0051] In one embodiment, a robot is provided, including a controller for performing a pose estimation method as follows: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

[0052] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computational and control capabilities. The memory includes a readable storage medium and internal memory. The non-volatile storage medium stores the operating system and computer-readable instructions. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The network interface is used to communicate with an external server via a network connection. When the computer-readable instructions are executed by the processor, they implement a robot pose estimation method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0053] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor performs the following steps when executing the computer-readable instructions: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

[0054] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which, when executed by one or more processors, perform the following steps: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; in response to the user's selection operation on the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

[0055] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0056] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0057] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A robot pose estimation method, characterized in that, include: Simultaneously acquire the color image and depth image of the current frame; Instance-level target detection and segmentation are performed on the color image to generate a target list containing target masks; In response to the user's selection of the target list, the target to be tracked is determined; Perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions; By combining the local region with the pre-registered 3D object model, the initial pose of the target to be tracked is obtained; The initial pose is recursively optimized to obtain the target pose.

2. The robot pose estimation method as described in claim 1, characterized in that, The step of performing instance-level object detection and segmentation on the color image to generate a target list containing target masks includes: The color image is used to perform target detection using an instance segmentation network to obtain category information for multiple detected targets; Based on the category information, multiple detection targets are segmented into instances to obtain the target mask for each detection target. The target list is generated based on the category and target mask of each of the detected targets.

3. The robot pose estimation method as described in claim 1, characterized in that, The step of performing cross-frame tracking association on the target list, assigning tracking identifiers to the targets to be tracked, and establishing mapping relationships includes: Obtain the session identifier of the target to be tracked; Perform cross-frame tracking association on the target list and assign tracking identifiers to the targets to be tracked; A mapping relationship is established for the target to be tracked based on the session identifier and the tracking identifier.

4. The robot pose estimation method as described in claim 1, characterized in that, Based on the mapping relationship and the target mask, the regions corresponding to the target to be tracked are extracted from the color image and the depth image respectively to generate local regions, including: Based on the mapping relationship, the target mask corresponding to the target to be tracked is selected in subsequent frames and used as the target mask of the current frame; Based on the target mask of the current frame, local color patches and local depth patches corresponding to the target to be tracked are cropped from the color image and the depth image, respectively; The local region is generated based on the local color patch and the local depth patch.

5. The robot pose estimation method as described in claim 1, characterized in that, The recursive optimization of the initial pose to obtain the target pose includes: The depth information in the local area is back-projected to generate a three-dimensional point cloud, and the geometric centroid of the three-dimensional point cloud is used as the pose observation value. Based on the iterative error state Kalman filter and the pose observations, the initial pose is recursively optimized to obtain the target pose.

6. A robot pose estimation device, characterized in that, include: The image acquisition module is used to simultaneously acquire the color image and depth image of the current frame; The target to be tracked module is used to perform instance-level target detection and segmentation on the color image, generate a target list containing target masks, and determine the target to be tracked in response to the user's selection operation on the target list. The mapping module is used to perform cross-frame tracking association on the target list, assign tracking identifiers to the targets to be tracked and establish mapping relationships; The local region module is used to extract the region corresponding to the target to be tracked from the color image and the depth image respectively, based on the mapping relationship and the target mask, to generate a local region; The initial pose module is used to combine the local region with the pre-registered 3D object model to obtain the initial pose of the target to be tracked. The target pose module is used to recursively optimize the initial pose to obtain the target pose.

7. The robot pose estimation device as described in claim 6, characterized in that, The target module to be tracked includes: The category unit is used to perform target detection on the color image using an instance segmentation network to obtain category information of multiple detected targets; A masking unit is used to perform instance segmentation on multiple detected targets based on the category information to obtain a target mask for each detected target; The list unit is used to generate the target list based on the category of each of the detected targets and the target mask.

8. A robot, comprising a controller, characterized in that, The controller is used to execute the robot pose estimation method as described in any one of claims 1 to 5.

9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the robot pose estimation method as described in any one of claims 1 to 5.

10. One or more readable storage media storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by one or more processors, the one or more processors cause the robot pose estimation method as described in any one of claims 1 to 5 to be performed.