Robot end effector pose generation method and system, storage medium

CN122442677BActive Publication Date: 2026-09-15SHANGHAI SEER INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610913790.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-15
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

[0009]为此,本发明的主要目的在于提供一种机器人末端执行器姿态生成方法及系统、存储介质,以解决现有抓取姿态估计方法动作类型单一、与夹爪外形强绑定、扩散模型推理速度慢以及遮挡场景下点云缺失导致姿态生成不稳定的问题

Benefits of technology

[0041] Parallel estimation of multiple behavior types: Simultaneously output independent feasibility scores for multiple behavior types such as gripping, pushing, and pulling in the same visual input. Through a multi-label non-exclusive classification mechanism, multiple operation feasibility can be achieved at the same pixel position at the same time, breaking through the limitations of traditional single-label classification and supporting the robot to flexibly select the optimal interaction strategy according to the scene context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122442677B_ABST
    Figure CN122442677B_ABST
Patent Text Reader

Abstract

The application discloses a robot end effector posture generation method and system, and a storage medium, wherein the method steps comprise: establishing a gripper coordinate system rule decoupled from the gripper shape; acquiring an RGB image, a depth map and camera parameters of a scene and generating a scene point cloud; generating a behavior candidate heat map based on the RGB image; screening out behavior candidate points and rough behavior directions according to each behavior type heat map and converting them into an initial end posture, adding noise to the initial end posture at a truncated time step to obtain an initial value of a diffusion model, and inputting the initial value and a condition vector into a conditional diffusion posture generation model, performing few-step posture denoising based on truncated diffusion, generating an end behavior posture, performing executability evaluation and sorting, and outputting an executable end behavior posture of the robot. Thus, the problems of single action type, strong binding with the gripper shape, slow diffusion model reasoning speed and unstable posture generation caused by point cloud loss in the occluded scene of the existing grasping posture estimation method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to robot end effector posture planning technology, and more particularly to a method, system, and storage medium for generating general behavior postures of multi-type robot end effectors based on behavior estimation. Background Technology

[0002] With the widespread application of robotics technology in intelligent manufacturing, smart logistics, and home services, the demand for robots to perform object manipulation tasks in unstructured environments is increasing. In complex scenarios such as logistics sorting, warehouse handling, industrial loading and unloading, and mobile robots opening and retrieving items, robots need to autonomously determine the behavior type, operating position, and spatial posture of their end effectors based on environmental information collected by visual sensors in order to achieve effective interaction with the environment. Specifically, robots not only need to determine whether a certain spatial position is suitable for grasping an object, but also whether the task should be completed through different interaction methods such as pushing or pulling, and further generate a six-degree-of-freedom end effector posture that matches the gripper structure and robot motion constraints.

[0003] Currently, existing robot grasping and detection methods can be mainly divided into two categories: methods based on RGB images and methods based on depth maps or point clouds. RGB image-based methods typically utilize convolutional neural networks to extract semantic information and scene context from images, outputting two-dimensional grasping boxes or grasping points. However, due to the lack of three-dimensional spatial structural information, it is difficult to directly generate executable six-degree-of-freedom end-effector poses.

[0004] Methods based on depth maps or point clouds directly utilize the geometric information of the target surface for six-DOF grasping posture regression. These methods can leverage the target's 3D shape, surface normals, and other geometric features to output more accurate spatial posture. However, existing point cloud-based methods typically tightly bind the model training process, evaluation metrics, and end effector structure. When the gripper shape, finger spacing, approach direction, or action type changes, it is often necessary to redefine labels, re-collect data, or retrain the model, resulting in low versatility and transfer efficiency, making it difficult to adapt to the application requirements of multi-end effector systems.

[0005] Furthermore, current robot manipulation technologies for multiple behaviors still face numerous challenges. In real-world scenarios, different behaviors such as pushing, pulling, and gripping may share the same area of ​​action in an image. For example, the same edge region of an object can be lifted by gripping, moved by pushing, and dragged by pulling. Traditional single-label classification methods treat different behavior types as mutually exclusive categories, failing to represent situations where the same pixel location can be suitable for multiple interaction types simultaneously. This limits the expressive power of behavior candidate regions, making it difficult to cover the complex and diverse operational needs in real-world scenarios.

[0006] In recent years, six-DOF grasping posture generation methods based on diffusion models have attracted attention due to their excellent multimodal generation capabilities. Diffusion models can sample from random distributions and iteratively denoise to generate diverse postures that meet physical constraints. However, standard diffusion models typically start with pure random Gaussian noise and require a complete multi-step denoising process, making the inference speed difficult to meet the requirements of online real-time response control for robots.

[0007] In scenarios with occlusion or limited sensor field of view, existing methods also have significant shortcomings. RGB images, with their rich semantic information, can usually identify valid candidate regions for actions; however, depth or point cloud data may be missing at the corresponding locations due to occlusion, reflection, or sensor blind spots. This can easily lead to unstable or erroneous pose estimations, and even cause operation failures.

[0008] Therefore, there is an urgent need in this field for a robot end effector behavior and posture generation scheme that can comprehensively utilize RGB semantic information and point cloud geometric information, support multiple behavior type labels, decouple from gripper shape, have real-time reasoning capabilities, and be stable and reliable. Summary of the Invention

[0009] Therefore, the main objective of this invention is to provide a robot end effector posture generation method and system, and storage medium, to solve the problems of existing grasping posture estimation methods, such as single action type, strong binding to gripper shape, slow inference speed of diffusion model, and unstable posture generation caused by missing point clouds in occluded scenes.

[0010] To achieve the above objectives, according to one aspect of the present invention, a method for generating pose of a robot end effector is provided, comprising the steps of:

[0011] Establish a gripper coordinate system rule decoupled from the gripper shape, uniformly describe the different gripper ends as a gripper coordinate system including the approach direction, gripping direction and lateral direction, and describe the key geometric position of the gripper through control points;

[0012] Acquire the RGB image, depth map, and camera parameters of the scene, and generate the scene point cloud based on the depth map;

[0013] The RGB image is input into the behavior estimation network, which independently outputs a heatmap of feasibility scores for multiple behavior types and the proximity direction corresponding to that location at each pixel location. Each behavior type is output in a multi-label manner, and the same pixel location is allowed to have multiple non-zero behavior scores at the same time.

[0014] Based on the heatmap of each behavior type, candidate pixels are selected for behavior. The candidate pixels are back-projected into three-dimensional candidate action points. Local point clouds are extracted with the candidate action points as the center. Condition vectors are constructed using the candidate action points, interaction type labels, approach directions and gripper control point geometric descriptions.

[0015] The candidate action point and approach direction are converted into the initial end pose, and noise is added to the initial end pose at the truncation time step to obtain the initial values ​​for diffusion model inference.

[0016] The initial values ​​and conditional vectors of the diffusion model inference are input into the conditional diffusion attitude generation model. A few steps of denoising are performed between the truncation time step and zero to generate a six-degree-of-freedom end-effector attitude that satisfies the gripper coordinate system rules.

[0017] The generated candidate poses are evaluated and ranked for feasibility, and the robot's executable end-effector poses are output.

[0018] In a possible preferred embodiment, the method steps further include: if the candidate pixel location or any depth in its neighborhood is missing, performing at least one of the following compensation methods: planning supplementary observation actions based on the candidate region and the robot's current pose, controlling the camera or robot to move to a new observation pose to acquire point clouds of the target region; or inputting behavior type, target category information and visible point clouds into the point cloud completion model to complete the missing target geometry under the current viewpoint; or estimating an approximate action point in the neighborhood of the candidate region based on edge, normal and depth continuity.

[0019] In a possible preferred embodiment, the behavior estimation network at each pixel location The output is:

[0020]

[0021] in In order to obtain a feasibility heatmap, To promote feasibility heatmaps, To generate a feasibility heatmap, Encode the proximity direction or direction vector; , , Through respectively The function outputs independently, without using Mutually exclusive categories.

[0022] In a possible preferred embodiment, the training loss of the behavior estimation network is:

[0023]

[0024] in For multi-label heatmap loss, To approximate directional loss, The direction loss weight is used.

[0025] In a possible preferred embodiment, the step of filtering candidate pixels based on the heatmap of each behavior type includes:

[0026] For each behavior type, a threshold is applied to the heatmap to retain pixels with scores greater than the threshold. Then, local non-maximum suppression is applied to the retained pixels to obtain the top K candidate points with the highest scores for each behavior type. Candidate points of different types are retained independently.

[0027] In a possible preferred embodiment, the initial values ​​for diffusion model inference are obtained using the following formula:

[0028]

[0029] in To truncate the time step The corresponding noise scheduling parameters, It is Gaussian noise. Less than the standard diffusion total time step T.

[0030] In a possible preferred embodiment, the few-step denoising is performed at the truncation time step. Perform N denoising steps between 0 and 0, where N is less than the full time step of the standard diffusion model; the denoising process uses either DDIM or DDPM sampling method.

[0031] In a possible preferred embodiment, the method steps further include: sampling multiple noises for the same candidate action point to obtain multiple candidate poses, and uniformly screening them in the executability evaluation.

[0032] To achieve the above objectives, corresponding to the above method example, according to another aspect of the present invention, a robot end effector posture generation system is also provided, comprising:

[0033] The storage module is used to store the program of the pose generation method steps as described in any of the above examples, so that the visual perception module, behavior estimation module, candidate point generation module, pose generation module, and evaluation execution module can call it up and execute it as needed.

[0034] The visual perception module is used to acquire the RGB image, depth map and camera parameters of the scene, and generate scene point cloud based on the depth map;

[0035] The behavior estimation module is used to input RGB images into the behavior estimation network and output a heatmap of feasibility scores for multiple behavior types and the proximity direction corresponding to each pixel location. Each behavior type adopts a multi-label output method, and the same pixel location is allowed to have multiple non-zero behavior scores at the same time.

[0036] The candidate point generation module is used to filter candidate pixels based on the heatmap of each behavior type, back-project the candidate pixels into three-dimensional candidate action points, and extract local point clouds with the candidate action points as the center.

[0037] The attitude generation module is used to construct a condition vector with candidate action points, interaction type labels, approach directions and gripper control point geometric descriptions, convert candidate action points and approach directions into initial end poses, add noise at the truncation time step to obtain the initial value of diffusion model inference, input the initial value of diffusion model inference and condition vector into the conditional diffusion attitude generation model, perform few-step denoising, and generate a six-degree-of-freedom end behavior attitude that satisfies the gripper coordinate system rules.

[0038] The evaluation and execution module is used to evaluate and rank the executability of the generated candidate poses, and output the robot's executable end-effector poses.

[0039] To achieve the above objectives, in accordance with the above method examples, according to another aspect of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the attitude generation method as described in any of the above examples.

[0040] The beneficial effects of the robot end effector posture generation method, system, and storage medium provided by this invention, corresponding to various examples, are as follows:

[0041] Parallel estimation of multiple behavior types: Simultaneously output independent feasibility scores for multiple behavior types such as gripping, pushing, and pulling in the same visual input. Through a multi-label non-exclusive classification mechanism, multiple operation feasibility can be achieved at the same pixel position at the same time, breaking through the limitations of traditional single-label classification and supporting the robot to flexibly select the optimal interaction strategy according to the scene context.

[0042] Gripper shape decoupling: A unified gripper coordinate system and control point description method are designed to normalize different grippers into three directions: clamping, lateral, and approach. By replacing the control point set, multiple gripper shapes can be adapted without retraining the model, which significantly improves versatility and transfer efficiency.

[0043] Real-time and efficient attitude generation: The diffusion attitude generation is guided by candidate points of RGB behavior heatmap. The candidate behavior position and approach direction are used as the initial conditions for diffusion generation. Combined with truncated time steps and low-step denoising, the number of inference steps is greatly reduced. It can meet the real-time control requirements while maintaining the multimodal generation capability.

[0044] Robustness in occluded scenarios: By actively supplementing observations, completing category-conditional point clouds, or estimating neighborhood depth, the problem of unstable pose generation caused by missing depth is solved, and the success rate of operation in occluded and view-limited scenarios is improved.

[0045] Semantic-geometric fusion: Through a hierarchical architecture of "RGB heatmap localization and behavior judgment → point cloud providing spatial constraints → diffusion model generating precise pose", semantic perception and geometric reasoning are organically integrated to improve the quality of end-effector behavior pose generation in unstructured complex scenarios. Attached Figure Description

[0046] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0047] Figure 1 This is a schematic diagram of the gripper coordinate system design in the posture generation method of the present invention, wherein the X-axis represents the gripping force or finger closing direction, the Y-axis represents the lateral direction pointing into the robot, and the Z-axis represents the approach direction;

[0048] Figure 2 This is a schematic diagram illustrating the output of a heatmap of multiple types of behavior in the posture generation method of the present invention.

[0049] Figure 3 This is a schematic diagram of truncated diffusion attitude generation based on RGB candidate points in the attitude generation method of the present invention.

[0050] Figure 4 This is a schematic diagram of the posture generation method steps of the present invention;

[0051] Figure 5 This is a schematic diagram of the attitude generation system of the present invention. Detailed Implementation

[0052] To enable those skilled in the art to better understand the technical solutions of the present invention, the specific technical solutions of the present invention will be clearly and completely described below in conjunction with embodiments, so as to help those skilled in the art further understand the present invention. Obviously, the embodiments described in this application are merely some embodiments of the present invention, and not all embodiments. It should be noted that, for those skilled in the art, the embodiments and features in the embodiments of this application can be combined with each other without departing from the concept of the present invention and without conflict. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the disclosure and protection scope of the present invention.

[0053] Furthermore, the terms "first," "second," "S1," "S2," etc., used in the specification, claims, and drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such features can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those described herein. At the same time, the stages described in each step are not necessarily to be implemented in the same step; it should be understood that the implementation order of the contents of each step stage can be adjusted and interchanged without violating the inventive concept, so that embodiments of the invention described herein can be implemented in orders other than those described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. Unless otherwise expressly specified and limited, the terms "set," "arrange," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this case based on the specific circumstances and in conjunction with existing technology.

[0054] To address the issues of existing grasping pose estimation methods, such as limited action types, strong dependence on gripper shape, slow inference speed of diffusion models, and unstable pose generation due to missing point clouds in occluded scenes, such as... Figures 1 to 4 As shown, the present invention provides a method for generating the attitude of a robot end effector, the example steps of which include:

[0055] Step S1: Establish the gripper coordinate system rules

[0056] like Figure 1 As shown, this example uniformly represents any gripper end in a gripper coordinate system G, where the X-axis represents the direction of the gripping force or the direction of finger closure, the Y-axis represents the lateral direction of the gripper with its positive direction pointing inwards towards the robot, and the Z-axis represents the end-effector approach direction. For different gripper shapes, several control points are used to describe the key geometric positions of the gripper, and the coordinates of the control points are normalized to the gripper coordinate system G.

[0057] For example, the gripper coordinate system G is defined as:

[0058]

[0059] in, The center of action of the gripper is usually located at the midpoint between the two fingers of the gripper or at the center of the suction cup; This refers to the direction of the clamping force or the direction in which the fingers close. The gripper is facing laterally, with its positive direction pointing inwards from the robot. To indicate the approach direction, point towards the target object.

[0060] Examples of coordinate system definitions for different types of end effectors are as follows:

[0061] Two parallel gripper fingers: Defined as the direction of the line connecting two fingertip control points; Defined as the direction from the midpoint of the line connecting the fingertips to the target; Depend on and The cross product is determined and adjusted so that its positive direction points inward into the robot.

[0062] Suction cup grippers: Defined as the adsorption normal direction (usually perpendicular to the suction cup surface and pointing outwards). Defined according to the installation direction or the desired direction of action (e.g., along a fixed direction of the robot wrist). Depend on and Cross product is determined.

[0063] Hook pull end: Defined as the approach direction (the direction in which the hook points towards the target); Defined as the direction of the pulling force (the direction in which the hook bears the pulling force); Depend on and Cross product is determined.

[0064] Among them, the gripper control point description Examples include a set of key point coordinates for the grippers in open, closed, or contact states:

[0065]

[0066] in The number of control points. For a two-finger gripper, control points may include: left fingertip point, right fingertip point, left finger root point, right finger root point, and palm point; for a suction cup gripper, control points may include: suction cup center point and four evenly distributed boundary points on the edge of the suction cup; for a hook-pulling end, control points may include: hook tip contact point, hook back support point, and hook root connection point.

[0067] Each control point is represented in the gripper coordinate system G, allowing the same pose generation model to learn the geometric constraints of different grippers. For example, the control point coordinates of a two-finger gripper in the open state can be represented as:

[0068] - Left fingertip: (-w / 2, 0, l)

[0069] - Right fingertip: (w / 2, 0, l)

[0070] - Palm point: (0, 0, 0)

[0071] Where w is the fingertip distance and l is the distance from the fingertip to the palm along the Zg direction. By unifying and normalizing the geometric parameters of different grippers to the gripper coordinate system G, the behavior pose generation model does not need to be retrained for each gripper; it only needs to replace the corresponding control point set Gi during inference.

[0072] Step S2: Obtain scene visual information

[0073] The example acquires a scene color image M, a depth map D, camera intrinsic parameters K, and camera-to-robot extrinsic parameters Tcb using an RGBD camera (e.g., Intel RealSense D435 or Azure Kinect). A scene point cloud P is generated based on the depth map D and camera intrinsic parameters K, and the color image M is normalized in size (e.g., scaled to 640×480 or 1280×720), color normalized (pixel values ​​are normalized to the [0,1] range), and distortion corrected.

[0074] Step S3: Generate candidate heatmaps for behavior based on RGB images

[0075] Existing deep learning-based affordance estimation networks are widely used in robotic manipulation tasks. They predict pixel-level feasibility heatmaps from RGB images using convolutional neural networks or Transformers to determine the position of the end effector. However, current affordance estimation methods typically treat different interaction types as mutually exclusive categories for single-label classification or only output heatmaps for a single behavior type. Therefore, they cannot represent situations where the same pixel location can be suitable for multiple interaction methods simultaneously. Furthermore, existing methods often directly use affordance heatmaps for manipulator selection, lacking effective coupling with subsequent six-DOF pose generation modules, making it difficult to leverage heatmap information to guide high-precision pose generation. In addition, existing affordance networks typically do not output end-effector approach direction information, forcing subsequent pose generation to blindly search a larger space, reducing generation efficiency and stability.

[0076] Therefore, such as Figure 2 As shown, an example of the present invention uses a color image. The input behavior estimation network (AffordanceEstimation Network) allows the network to estimate the behavior at each pixel location. Output independent feasibility scores for multiple behavior types (such as gripping, pushing, and pulling). ,in Belongs to the set of behavior types , It includes at least three behaviors: grab, push, and pull. Each behavior type uses a multi-label output method, and different types are not mutually exclusive normalized, so that the same pixel location can have multiple non-zero behavior scores at the same time.

[0077] Step S4: Generate candidate behavior points and rough behavior directions

[0078] Example heatmap based on behavior type Non-Maximum Suppression (NMS) and thresholding are performed to obtain a set of candidate pixels. ,in Indicates the candidate pixel coordinates, Indicates the type of behavior. This represents the behavior score of the candidate point. The behavior estimation network also outputs the proximity direction for each candidate pixel. Alternatively, it can output a direction encoding vector used to recover the approximate direction.

[0079] Specifically, the following examples illustrate the network structure, output format, and training method of the behavior estimation network in steps S2 to S4.

[0080] 1) Network Structure

[0081] In this example, the behavior estimation network can adopt an encoder-decoder architecture. Preferably, the encoder uses a ResNet-50 or ResNet-101 backbone network to extract multi-scale features of the RGB image; the decoder uses a Feature Pyramid Network (FPN) or U-Net structure to fuse deep semantic features with shallow spatial features and output a behavior heatmap with the same resolution as the input image.

[0082] In one specific implementation, the behavior estimation network includes:

[0083] - Shared encoder: Extracts deep features from RGB images;

[0084] - Multi-task decoding head: contains three parallel heatmap prediction heads (corresponding to gripping, pushing, and pulling respectively) and one direction prediction head;

[0085] - Each heatmap prediction head outputs a single-channel heatmap, and the pixel values ​​are constrained to the [0,1] interval by the Sigmoid activation function;

[0086] - The direction prediction head outputs a direction encoding vector or directly uses a classification method to output a discretized proximity direction.

[0087] 2) Output format

[0088] This behavior estimates the network at each pixel location. The output is:

[0089] ,

[0090] in In order to obtain a feasibility heatmap, To promote feasibility heatmaps, To generate a feasibility heatmap, Encode the proximity direction or direction vector.

[0091] , , Each pixel is output independently using the Sigmoid function, without employing Softmax mutual exclusion classification. This means that for any pixel... , Multiple values ​​can be non-zero simultaneously; for example, a pixel may output multiple values ​​simultaneously. (High clamping feasibility) and (High feasibility of implementation), thus truly reflecting the physical scenario where the same location is suitable for multiple interaction methods.

[0092] 3) Training loss

[0093] The training loss of the behavior estimation network is:

[0094]

[0095] in For multi-label heatmap loss, binary cross-entropy loss (BCE Loss), Focal Loss, or mean squared error loss (MSE Loss) can be used.

[0096] Preferably, Focal Loss is used to address the imbalance between positive and negative samples:

[0097]

[0098] in To predict probabilities for the model, For category weights, For focusing parameters.

[0099] Lang represents the proximity direction loss. When the direction prediction head outputs continuous direction vectors, cosine distance loss is used.

[0100]

[0101] in To predict the direction vector, This represents the true direction vector. When the direction prediction head outputs discrete direction categories, the direction classification cross-entropy loss is used.

[0102] The directional loss weight is preferably set to a value between 0.1 and 1.0.

[0103] 4) Candidate point screening

[0104] The candidate pixel set Q is obtained by the following method: a heatmap for each behavior type c. Perform threshold filtering and retain Pixels, of which For behavior type related thresholds (e.g.) Then, local non-maximum suppression (NMS) is applied to the retained pixels to obtain the top K candidate points with the highest scores for each behavior type. Different types of candidate points are retained independently, that is, pinch candidate points, push candidate points, and pull candidate points are screened separately without interference.

[0105] Step S5 converts candidate pixels into three-dimensional candidate action points.

[0106] For candidate pixels Read the corresponding depth value from depth map D. And using the camera intrinsic parameter K, it is back-projected into a three-dimensional candidate action point in the camera coordinate system. If the corresponding depth is missing, a valid depth is searched within the neighborhood of the candidate pixel, or the point cloud compensation process in step S10 is triggered.

[0107] Step S6: Extract local point cloud of candidate region

[0108] Three-dimensional candidate action points Centered on the point cloud P, extract a local point cloud Pi within a bounding box of radius r (e.g., r = 0.05m to 0.15m) or fixed size (e.g., 0.1m × 0.1m × 0.1m), and then select the appropriate local point cloud Pi based on the candidate behavior type. Approach direction Description of gripper control points Construct condition vector Condition vector It should include at least the coordinates of the candidate action point, behavior type encoding, approach direction, local point cloud features, and geometric description of the gripper control point.

[0109] Step S7: Construct the initial state of the behavior pose diffusion model

[0110] Candidate action points and approach direction Convert to initial end pose The initial end pose includes three-dimensional translation. and rotation And satisfy the condition that the Z-axis in the gripper coordinate system is consistent with or approximately consistent with the approach direction. Encoded as attitude vector and at the cutoff time step τ to By adding a small amount of noise, we obtain the initial values ​​for the diffusion model inference. .

[0111] Step S8: Perform few-step attitude denoising based on truncated diffusion

[0112] Will Local point cloud Pi, behavior type Description of gripper control points The input conditional diffusion pose generation model performs N denoising steps between the truncated time step τ and 0, where N is less than the number of full time steps in the standard diffusion model, preferably 2 to 10 steps. The denoising process uses either DDIM (Denoising Diffusion Implicit Models) or DDPM (Denoising Diffusion Probabilistic Models) sampling methods, outputting one or more candidate six-DOF end-effector poses. .

[0113] The following examples illustrate the six-degree-of-freedom end-effector behavior pose generation method based on the truncated diffusion model in steps S6 to S8.

[0114] 1) Construction of condition vectors

[0115] Example using three-dimensional candidate action points Centered on the point cloud P, extract the radius... The local point cloud Pi is defined within a spherical neighborhood. The local point cloud Pi is downsampled (e.g., downsampled to 1024 or 2048 points using voxels), and its features are extracted using a point cloud encoder such as PointNet++ or DGCNN. .

[0116] Condition vector The construction method is as follows:

[0117]

[0118] in:

[0119] - The coordinates of the candidate point of action;

[0120] - One-hot encoding for behavior types, Number of behavior types;

[0121] - It is a unit vector in the approximate direction;

[0122] - Here, d represents the feature vector of a local point cloud, and d is the feature dimension.

[0123] - is the flattened vector of the control point coordinates of the gripper, and m is the number of control points.

[0124] 2) Initial attitude and initial values ​​of truncation diffusion

[0125] Candidate action points As a translation component Approaching direction The initial rotation matrix is ​​constructed as follows, with the Z-axis as the reference direction. :

[0126] Let the approach direction be... The Z-axis direction is defined. A reference vector in the world coordinate system (e.g., [0,0,1] or the robot's wrist orientation vector) is selected, and the Z-axis is constructed using Gram-Schmidt orthogonalization. and ,make sure It is a valid rotation matrix and satisfies the gripper coordinate system rules.

[0127] Initial end attitude .Will Encoded as attitude vector In this embodiment, a six-dimensional continuous representation using translation and rotation is employed:

[0128]

[0129] in Rotation matrix The rotation can be represented in six dimensions (e.g., by flattening the first two columns of the rotation matrix). In alternative implementations, rotations can also be represented using quaternions, axis angles, or Lie algebras.

[0130] The initial value for cutoff diffusion is obtained using the following formula:

[0131]

[0132] in To truncate the time step The corresponding noise scheduling parameters, It is Gaussian noise. The total time step is less than T in the standard diffusion process. Through this truncation and noise-adding process, the starting point of the diffusion model's inference changes from pure random noise to a perturbation point near the candidate pose, which preserves the generation diversity of the diffusion model and avoids starting the search from a completely random state.

[0133] 3) Low-step noise reduction generation

[0134] Will Conditional vector A time-step encoded τ-input conditional diffusion pose generation model is proposed. This model employs a denoising network. Its input includes the current noise pose. Time step τ and condition vector Output the predicted noise or the denoised pose:

[0135]

[0136] in This includes local point cloud features, behavior type encoding, gripper control point descriptions, candidate action points, and proximity directions. The model is updated iteratively from... Obtain y0_hat and decode it into a candidate end-behind behavior pose.

[0137] In one specific implementation A Transformer architecture or a diffusion model based on graph neural networks can be used, where time step τ is embedded through sinusoidal position encoding and conditional vectors. After encoding using a multilayer perceptron (MLP), the noise is fused with pose features. The denoising process is performed in N steps between the truncated time step τ and 0, where N is much smaller than the full time step T of the standard diffusion model. Preferably, N ranges from 2 to 10 steps.

[0138] After each denoising step, the output pose is decoded: the translation component is extracted from y0_hat. and rotational components The generated posture is mapped back to the gripper coordinate system G to ensure that the generated posture meets geometric constraints such as the Z-axis being consistent with the approach direction and the control point Gi being located in a reasonable spatial position.

[0139] 4) Multimodal attitude sampling

[0140] To generate multimodal poses (i.e., multiple feasible gripper orientations corresponding to the same candidate action point), multiple different noise samples can be taken from the same candidate action point. M initial diffusion values ​​were obtained respectively. After denoising, M candidate poses are obtained. These candidate poses are uniformly screened and evaluated in step S9.

[0141] Step S9: Perform attitude executability assessment and ranking.

[0142] For each candidate pose The system performs gripper control point collision detection, robot kinematic reachability detection, behavior type consistency detection, and local geometric matching scoring. Candidate poses are ranked based on a combination of heatmap score, diffusion model confidence, and executability score, and the optimal end-effector pose is output.

[0143] The following example illustrates the method for comprehensively evaluating candidate poses in step S9 of this embodiment.

[0144] 1) Collision Detection

[0145] For example, for each candidate pose The gripper control point Gi is transformed to the scene coordinate system using this pose transformation, resulting in the transformed set of control points:

[0146]

[0147] The system detects whether the transformed control points penetrate or are too close to obstacles in the scene point cloud P. Preferably, a nearest neighbor search based on a KD-Tree or octree is used to calculate the shortest distance from each control point to the scene point cloud. If the shortest distance is less than a safety threshold δ (e.g., δ = 0.002m), a collision is determined, and the collision penalty term Scoll for the candidate pose is assigned a large positive value; otherwise, Scoll = 0 or a smoothing penalty value is assigned based on the distance.

[0148] 2) Robot kinematic reachability detection

[0149] Candidate poses Transform from the camera coordinate system to the robot base coordinate system (see step S11), and call the robot inverse kinematics (IK) solver to determine whether there is a joint angle solution that satisfies the joint limits and singular point constraints. If a valid IK solution exists, the reachability score Sreach = 1; if no solution exists or an approximate solution is required, Sreach is assigned a value between 0 and 1 according to the degree of proximity.

[0150] 3) Behavioral type consistency detection

[0151] Set geometric constraints according to different behavior types:

[0152] Grasp behavior: Detects whether there are grippable surfaces on both sides of the gripper along the X-axis. Specifically, it emits rays or searches the local point cloud along the positive and negative Xg directions to determine whether there are opposing surfaces on both sides (the angle between the surface normal and Xg is less than a threshold, such as 30°). Simultaneously, it detects whether the gripper's closing path (the trajectory of the control point during the opening and closing process) collides with obstacles.

[0153] Push behavior: Detects whether the angle between the end effector direction (usually the Z-axis approach direction or the X-axis action direction) and the target surface normal or desired motion direction meets a preset range (e.g., the angle is between 60° and 120° to ensure effective contact and pushing). Simultaneously, it detects whether the pushing contact area has sufficient support area.

[0154] Pull behavior: Detects whether the end effector can establish stable contact with the target edge, hole, handle, or hookable area. Specifically, it detects whether the local geometry near the candidate point of action has depressions, holes, or elongated structures that would allow the hook-like end or gripper fingertip to establish a mechanically stable hook.

[0155] 4) Overall score and ranking

[0156] The comprehensive scoring function for candidate poses is:

[0157]

[0158] in:

[0159] - Scoring the RGB heatmap reflects the semantic feasibility of the behavior.

[0160] - The attitude confidence of the diffusion model can be calculated from the predicted noise amplitude or the model output probability during the denoising process.

[0161] - The local geometric matching score is calculated by combining various indicators of behavior type consistency detection.

[0162] - Assign a score to the robot's reachability;

[0163] - This is a collision penalty item;

[0164] - w1 to w5 are weighting coefficients, for example... It can be adjusted according to the specific scenario.

[0165] All candidate poses are sorted in descending order based on the overall score, and the pose with the highest score is selected as the optimal end-effector pose for output. If the highest score is lower than the global threshold, the point cloud compensation process in step S10 is triggered or an operation failure is reported to the user.

[0166] Step S10: Handle missing point clouds in candidate regions

[0167] If the candidate pixel location or its neighborhood lacks depth, at least one of the following compensation methods can be performed: First, plan supplementary observation actions based on the candidate region and the robot's current pose, controlling the camera or robot to move to a new observation pose to acquire the target region point cloud; Second, input the behavior type, target category information, and visible point cloud into the point cloud completion model to complete the missing target geometry from the current viewpoint; Third, estimate the approximate action point in the candidate region's neighborhood based on edge, normal, and depth continuity. After compensation, return to steps S5 to S9 to generate the behavior pose.

[0168] Specifically, the compensation method for missing point clouds in the candidate region in step S10 is illustrated below:

[0169] 1) Triggering conditions

[0170] The point cloud compensation process is triggered under at least one of the following conditions:

[0171] - The candidate pixel (ui, vi) has a null depth value (NaN or 0) in the depth map D.

[0172] - The number of valid points in the local point cloud Pi centered at Xi is less than the threshold Nmin (e.g., Nmin=50).

[0173] - The confidence level of the normal estimate of the local point cloud Pi is lower than the threshold (e.g., the eigenvalue ratio λ3 / λ1 obtained by PCA analysis < 0.05).

[0174] - The candidate pose is unstable in collision detection (e.g., a small perturbation causes a drastic change in the collision determination).

[0175] - The combined score difference of multiple candidate poses is below the threshold (making it difficult to distinguish the optimal pose).

[0176] 2) Compensation Method 1: Active Supplementary Observation

[0177] Based on the position Xi of the candidate region in the camera coordinate system and the current robot pose, supplementary observation actions are planned. Specifically, a new pose is calculated to align the camera optical axis with the candidate region and shorten the observation distance. The robot's kinematics planner generates a kinematics sequence from the current pose to... A collision-free path is established, and the robot arm or mobile chassis moves the camera to a new observation pose. RGBD data is reacquired in the new observation pose, the depth map D and point cloud P are updated, and then the process returns to step S5 to re-perform candidate point backprojection and local point cloud extraction.

[0178] 3) Compensation Method Two: Category Condition Point Cloud Completion

[0179] The behavior type ci, target category information (e.g., object category labels obtained from semantic segmentation of RGB images, such as "drawer," "box," and "cup"), and visible point cloud Pi are input into a pre-trained point cloud completion model (e.g., a completion network based on PCN, PoinTr, or implicit neural representation). The point cloud completion model infers the complete shape of the missing region based on category priors and visible geometry, and outputs the completed local point cloud Pi'. Pi' replaces the original Pi, and the process returns to step S6 to continue execution.

[0180] 4) Compensation Method 3: Neighborhood Depth Estimation

[0181] Within the neighborhood of the candidate pixel (ui, vi) (e.g., a 15×15 pixel window), an approximate application point is estimated based on edge, normal, and depth continuity. Specifically, a planar fit or quadratic surface fit is performed on the effective depth points within the neighborhood, and the candidate pixel position is substituted into the fitted surface to obtain the estimated depth. Then, the back projection is used to obtain an approximate three-dimensional candidate action point. Simultaneously, the local normal of the candidate region is estimated based on the normal distribution of effective points in the neighborhood, serving as a substitute for the approach direction.

[0182] By employing methods such as robot-assisted observation supplementation, category-based conditional point cloud completion, or neighborhood depth estimation, local geometric information of candidate regions is obtained. This addresses the issue of unstable pose generation caused by the lack of depth information at corresponding locations despite the identifiable candidate regions in RGB images, thus improving the robustness of pose generation in scenarios with occlusion or limited sensor field of view.

[0183] Step S11: Transform the end effector's pose to the robot's base coordinate system

[0184] Based on the extrinsic parameter Tcb from the camera to the robot base, the end effector pose in the camera coordinate system is determined. Conversion to pose in robot base coordinate system And send it to the robot motion planning module for execution.

[0185] The following examples, using steps S1 and S11, illustrate the adaptation methods for different gripper shapes and the coordinate transformation process.

[0186] 1) Multi-claw adapter

[0187] During the system deployment phase, corresponding control point sets Gi are pre-generated for different end effectors and stored in the configuration library. When the robot changes its end effector, the corresponding Gi is loaded from the configuration library, eliminating the need to retrain the behavior estimation network or spread the pose generation model.

[0188] For example, when changing from a two-finger gripper to a suction cup gripper:

[0189] - Redefine Zg of the gripper coordinate system G as the suction cup adsorption normal;

[0190] - Load the set of control points corresponding to the suction cup (including the four boundary points of the suction cup center and edge);

[0191] - In the collision detection in step S9, the collision detection of the suction cup boundary point is replaced with the fingertip point detection;

[0192] - In the behavior type consistency detection, the clamping constraint is replaced with the adsorption plane constraint (requiring that the angle between the surface normal of the adsorption region and Zg is less than 10°, and the surface roughness / curvature meets the adsorption requirements).

[0193] 2) Coordinate transformation

[0194] In step S11, the transformation relationship of the end effector pose from the camera coordinate system to the robot base coordinate system is as follows:

[0195]

[0196] in The end pose in the camera coordinate system is represented by a 4×4 homogeneous transformation matrix. Tcb is the end-effector pose in the robot base coordinate system, and Tcb is the extrinsic transformation matrix from the camera coordinate system to the robot base coordinate system (obtained in advance through hand-eye calibration).

[0197] Specifically, if In the camera coordinate system, represented by the rotation matrix Rc and translation tc, then:

[0198]

[0199]

[0200] in The transformed posture The data is sent to the robot's motion planning module (such as MoveIt or OMPL), where the motion planner generates and executes the joint space trajectory.

[0201] Furthermore, in other alternative implementations, the behavior estimation network in the above examples can use VisionTransformer (ViT) or Swing Transformer as the backbone network to capture global contextual relationships through an attention mechanism; it can also adopt a hybrid architecture of convolution and Transformer (such as CoAtNet) to balance local details and global semantics. In addition, the behavior estimation network can also be extended to a multimodal visual language model, guiding the generation of heatmaps for specific behavior types through natural language instructions (such as "Please open the box").

[0202] Furthermore, in other alternative implementations, the pose generation model in the example steps S7 and S8 above can also be implemented using a conditional variational autoencoder (CVAE), a normalizing flow, an energy model (EBM), a Transformer autoregressive model, or a sampling-based optimization model (such as CEM, MPPI). These models can also generate six-DOF poses using candidate action points, local point clouds, behavior types, and gripper geometry as conditions, but the preferred option is the truncated diffusion model, as it achieves the best balance between multimodal generation quality and geometric constraint satisfaction.

[0203] Furthermore, in other alternative implementations, the initial pose of the truncation diffusion in the above example... In addition to being determined by RGB heatmap candidate points and proximity direction, it can also be determined or replaced by the following methods:

[0204] - Attitude anchor points obtained by clustering: Cluster the historical successful operation attitudes and select the cluster center closest to the candidate point as the initial value;

[0205] - Historical successful operation postures: directly reuse historical postures in similar scenarios;

[0206] - Robot teaching trajectory: Obtain a rough posture through manual teaching;

[0207] - Heuristic geometry rules: Generate initial poses heuristically based on object category and local geometry (such as box edges, cylinder axes).

[0208] Furthermore, in other alternative implementations, the gripper geometry described in the above examples can be represented, in addition to the set of control points, using a mesh model, a symbolic distance field (SDF), a bounding box, key segments, contact patches, or a parametric gripper model (such as geometric parameters in a URDF). These representations can all be preprocessed to a unified description in the gripper coordinate system G, which can then be used as conditional input for the diffusion model.

[0209] On the other hand, corresponding to the above method examples, such as Figure 5 As shown, the present invention also provides a robot end effector posture generation system, which includes:

[0210] The storage module is used to store the program of the pose generation method steps as described in any of the above examples, so that the visual perception module, behavior estimation module, candidate point generation module, pose generation module, and evaluation execution module can call it up and execute it as needed.

[0211] The visual perception module is used to acquire the RGB image, depth map and camera parameters of the scene, and generate scene point cloud based on the depth map;

[0212] The behavior estimation module is used to input RGB images into the behavior estimation network and output a heatmap of feasibility scores for multiple behavior types and the proximity direction corresponding to each pixel location. Each behavior type adopts a multi-label output method, and the same pixel location is allowed to have multiple non-zero behavior scores at the same time.

[0213] The candidate point generation module is used to filter candidate pixels based on the heatmap of each behavior type, back-project the candidate pixels into three-dimensional candidate action points, and extract local point clouds with the candidate action points as the center.

[0214] The attitude generation module is used to construct a condition vector with candidate action points, interaction type labels, approach directions and gripper control point geometric descriptions, convert candidate action points and approach directions into initial end poses, add noise at the truncation time step to obtain the initial value of diffusion model inference, input the initial value of diffusion model inference and condition vector into the conditional diffusion attitude generation model, perform few-step denoising, and generate a six-degree-of-freedom end behavior attitude that satisfies the gripper coordinate system rules.

[0215] The evaluation and execution module is used to evaluate and rank the executability of the generated candidate poses, and output the robot's executable end-effector poses.

[0216] On the other hand, corresponding to the above method examples, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the attitude generation method as described in any of the above examples.

[0217] In summary, the robot end effector posture generation method, system, and storage medium provided by this invention support parallel estimation of multiple behavior types and can simultaneously output independent feasibility scores for behavior types such as gripping, pushing, and pulling from the same visual input. Through a multi-label non-exclusive classification mechanism, it can express situations where the same pixel location can simultaneously possess multiple operation modes, breaking through the limitation of traditional single-label classification which can only output a single behavior type, enabling the robot to select the optimal interaction strategy based on the scene context.

[0218] Furthermore, by designing a unified gripper coordinate system description method that is decoupled from the gripper shape, this invention enables the behavior posture generation model to adapt to various gripper shapes such as two-finger grippers, suction cup grippers, and hook-shaped ends, without the need to retrain the model for each gripper, thus improving the versatility and transfer efficiency of the method.

[0219] To improve the speed and stability of pose generation, this invention also designs a method for guiding diffusion pose generation using candidate points from an RGB behavior heatmap. The candidate behavior position and approach direction are used as initial conditions for diffusion generation, avoiding the generation of poses entirely from random Gaussian noise. By truncating time steps and using fewer DDIM / DDPM denoising steps, the number of inference steps is significantly reduced while maintaining the multimodal generation capability of the diffusion model, thus significantly improving the speed and stability of pose generation and meeting the requirements of online real-time robot control.

[0220] Furthermore, this invention presents a six-DOF end-effector pose generation process based on innovative truncated diffusion and minimal denoising. By adding a small amount of noise near the candidate pose as the diffusion starting point, this invention retains the ability of the standard diffusion model to generate diverse candidate poses while significantly reducing the number of denoising iterations, achieving an effective balance between generation quality and inference efficiency.

[0221] Furthermore, this invention fully leverages the advantages of RGB images for semantic region judgment and behavior type estimation, while utilizing point clouds or depth maps to provide spatial geometric constraints. Through a hierarchical architecture of "RGB heatmap to determine operation location and type → point cloud to provide spatial constraints → diffusion model to generate precise pose," it achieves the organic integration of semantic perception and geometric reasoning, improving the success rate of operations in complex unstructured scenarios.

[0222] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The present invention is limited only by the claims and their full scope and equivalents. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.

[0223] Those skilled in the art will understand that, besides implementing the system, apparatus, unit, and its modules provided by this invention in purely computer-readable program code, the same program can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system, apparatus, and its modules provided by this invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; alternatively, modules for implementing various functions can be considered both software programs implementing the method and structures within the hardware component.

[0224] Furthermore, all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a microcontroller, chip, or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0225] Furthermore, various different implementations of the present invention can be combined arbitrarily, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed in the present invention.

Claims

1. A method for generating the attitude of a robot end effector, comprising the following steps: Establish a gripper coordinate system rule decoupled from the gripper shape, uniformly describe the different gripper ends as a gripper coordinate system including the approach direction, gripping direction and lateral direction, and describe the key geometric position of the gripper through control points; Acquire the RGB image, depth map, and camera parameters of the scene, and generate the scene point cloud based on the depth map; The RGB image is input into the behavior estimation network, which independently outputs a heatmap of feasibility scores for multiple behavior types and the proximity direction corresponding to that location at each pixel location. Each behavior type is output in a multi-label manner, and the same pixel location is allowed to have multiple non-zero behavior scores at the same time. Based on the heatmap of each behavior type, candidate pixels are selected for behavior. The candidate pixels are back-projected into three-dimensional candidate action points. Local point clouds are extracted with the candidate action points as the center. Condition vectors are constructed using the candidate action points, interaction type labels, approach directions and gripper control point geometric descriptions. The candidate action point and approach direction are converted into the initial end pose, and noise is added to the initial end pose at the truncation time step to obtain the initial values ​​for diffusion model inference. The initial values ​​and conditional vectors of the diffusion model inference are input into the conditional diffusion attitude generation model. A few steps of denoising are performed between the truncation time step and zero to generate a six-degree-of-freedom end-effector attitude that satisfies the gripper coordinate system rules. The generated candidate poses are evaluated and ranked for feasibility, and the robot's executable end-effector poses are output. If a candidate pixel location or any depth within its neighborhood is missing, perform at least one of the following compensation methods: plan supplementary observation actions based on the candidate region and the robot's current pose, and control the camera or robot to move to a new observation pose to acquire point clouds of the target region; or input the behavior type, target category information, and visible point clouds into the point cloud completion model to complete the missing target geometry under the current viewpoint; or estimate an approximate action point in the neighborhood of the candidate region based on edge, normal, and depth continuity.

2. The method of claim 1, wherein the behavior estimation network at each pixel location The output is: , in In order to obtain a feasibility heatmap, To promote feasibility heatmaps, To generate a feasibility heatmap, Encode the proximity direction or direction vector; , , Through respectively The function outputs independently, without using Mutually exclusive categories.

3. The method according to claim 1, wherein the training loss of the behavior estimation network is: , in For multi-label heatmap loss, To approximate directional loss, The direction loss weight is used.

4. The method according to claim 1, wherein the step of filtering candidate pixels for each behavior type based on the heatmap includes: For each behavior type, a threshold is applied to the heatmap, and pixels with scores greater than the threshold are retained. Then, local non-maximum suppression is applied to the retained pixels to obtain the top K candidate points with the highest scores under each behavior type. Candidate points of different types are retained independently.

5. The method according to claim 1, wherein the initial values ​​for diffusion model inference are obtained by the following formula: , in To truncate the time step The corresponding noise scheduling parameters, It is Gaussian noise. Less than the standard diffusion total time step T.

6. The method of claim 1, wherein the few-step denoising is performed at the truncation time step. Perform N denoising steps between 0 and 0, where N is less than the full time step of the standard diffusion model; the denoising process uses either DDIM or DDPM sampling method.

7. The method of claim 1, wherein the step further comprises: Multiple noise samples are taken from the same candidate action point to obtain multiple candidate poses, which are then uniformly screened during the executability assessment.

8. A robot end effector posture generation system, comprising: The storage module is used to store the program of the pose generation method steps as described in any one of claims 1 to 7, so that the visual perception module, behavior estimation module, candidate point generation module, pose generation module, and evaluation execution module can retrieve and execute it as needed. The visual perception module is used to acquire the RGB image, depth map and camera parameters of the scene, and generate scene point cloud based on the depth map; The behavior estimation module is used to input RGB images into the behavior estimation network and output a heatmap of feasibility scores for multiple behavior types and the proximity direction corresponding to each pixel location. Each behavior type adopts a multi-label output method, and the same pixel location is allowed to have multiple non-zero behavior scores at the same time. The candidate point generation module is used to filter candidate pixels based on the heatmap of each behavior type, back-project the candidate pixels into three-dimensional candidate action points, and extract local point clouds with the candidate action points as the center. The attitude generation module is used to construct a condition vector with candidate action points, interaction type labels, approach directions and gripper control point geometric descriptions, convert candidate action points and approach directions into initial end poses, add noise at the truncation time step to obtain the initial value of diffusion model inference, input the initial value of diffusion model inference and condition vector into the conditional diffusion attitude generation model, perform a few-step denoising, and generate a six-degree-of-freedom end behavior attitude that satisfies the gripper coordinate system rules. The evaluation and execution module is used to evaluate and rank the executability of the generated candidate poses, and output the robot's executable end-effector poses.

9. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the attitude generation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hybrid control method and device for mobile operation robot, robot and medium

    CN121937828A

  • Robot control method and device, electronic equipment, storage medium and product

    CN121989233A