Robot grabbing control method and device, robot and storage medium

By generating the optimal grasping posture and adjusting the torque value using a conditional diffusion model and tactile latent space, the problem of the disconnect between posture selection and force control in robot grasping is solved, achieving a balance between stability and efficiency and improving the reliability of grasping diverse objects.

CN121572301APending Publication Date: 2026-02-27深圳市真保科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511778232.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing robot grasping technologies, posture selection and force control are disconnected, and there is a lack of explicit consideration of contact mechanics characteristics, making it difficult to balance grasping stability and efficiency. Fixed force control strategies lack adaptability, and the force adjustment process is lagging and passive.

Method used

By acquiring point cloud data of the target object, multiple candidate grasping postures are generated. The optimal posture is selected using the physical consistency index. The desired tactile image is predicted by combining the conditional diffusion model. The torque value of the gripper is iteratively calculated in the tactile latent space to achieve closed-loop precise adjustment.

Benefits of technology

It achieves a bridge between attitude selection and force control, ensuring optimized force distribution during grasping, strong adaptability, overcoming the lag and high energy consumption problems of traditional methods, and improving grasping stability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121572301A_ABST
    Figure CN121572301A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, and discloses a robot grabbing control method and device, a robot and a storage medium, and the method comprises the steps: firstly obtaining point cloud data of a target object, and generating a plurality of candidate grabbing postures; then calculating a physical consistency index based on the point cloud data of the contact area, and reordering the candidate postures to screen out an optimal grabbing posture; then, a contact area depth image corresponding to the optimal grabbing posture, physical condition parameters of an object and a current touch observation image serve as input, and an expected touch image under the force optimal stable grabbing condition is predicted through a conditional diffusion model; and finally, according to the real-time tactile image, a clamping jaw torque value is iteratively calculated in the tactile submerged space by using a linear quadratic regulator, and the clamping jaw is controlled to realize stable grabbing with the minimum necessary force. The problem that in a traditional method, due to disjunction of attitude planning and force control, grabbing stability and efficiency are difficult to consider at the same time is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot control, and in particular to a robot grasping control method and device, a robot and a storage medium. BACKGROUND

[0002] As the core link of intelligent manufacturing, logistics sorting and service robots, the performance of robot grasping technology directly affects the intelligent level and operation efficiency of the entire system. In recent years, significant progress has been made in grasping planning methods based on visual perception and data-driven.

[0003] However, such methods still have obvious limitations in practical application. First, existing methods mainly rely on geometric force closure criteria or network confidence scores for pose selection, lacking explicit consideration of contact mechanics characteristics. For example, a grasping pose that is considered to be well closed in geometry may have problems such as surface curvature discontinuity or inconsistent normal direction in the contact area, resulting in uneven force distribution in the actual grasping process.

[0004] Secondly, in terms of grasping force control, existing methods mostly use fixed or predefined force values. This open-loop force control method is difficult to adapt to the physical characteristics of different objects. For example, when grasping a fragile egg, a fixed force value that is set too large may cause the eggshell to break, while a fixed force value that is set too small may cause the object to slip due to insufficient grasping force. Similarly, when grasping a smooth metal part, a fixed clamping force may not be able to adaptively adjust according to the actual contact state, thereby affecting the stability and reliability of grasping.

[0005] In addition, traditional force control methods mostly rely on passive response during grasping. For example, increasing the clamping force when detecting object slip, this lag compensation method not only has slow response speed, but also easily causes oscillation or overshoot in the force regulation process, which is particularly disadvantageous for soft, deformable or fragile objects.

[0006] In summary, there are the following technical problems in existing robot grasping technology that need to be solved: the grasping pose selection method based on pure geometry or data-driven is mismatched with the actual contact mechanics characteristics; the fixed or predefined force control strategy lacks the ability to adapt to the physical characteristics of objects; and the lag and passivity of the force regulation process. These problems restrict the ability of robots to perform agile, delicate and reliable grasping of diverse objects in complex scenarios. SUMMARY

[0007] Therefore, it is necessary to propose a robot grasping control method, device, robot and storage medium to solve the technical problem that the grasping stability and efficiency are difficult to balance due to the disconnection between pose selection and force control in the existing robot grasping process.

[0008] In a first aspect, a robot grasping control method is provided, and the method comprises: obtaining point cloud data of a target object, and generating a plurality of candidate grasping poses based on the point cloud data; reordering the plurality of candidate grasping poses based on a physical consistency index to filter out an optimal grasping pose; the physical consistency index is calculated based on contact area point cloud data of a contact area between a gripper corresponding to each candidate pose and the target object; inputting a depth image of the contact area corresponding to the optimal grasping pose, physical condition parameters of the target object, and a current haptic observation image of the target object into a conditional diffusion model to predict an expected haptic image under a force-optimal stable grasping condition; According to the real-time obtained haptic image, the torque value of the gripper is iteratively calculated in the haptic latent space by using a linear quadratic regulator, and the gripper is controlled to grasp the target object by using the torque value, so that the current haptic image of the gripper approximates to the expected haptic image.

[0009] In a second aspect, a robot grasping control device is provided, and the device comprises: an acquisition module configured to obtain point cloud data of a target object, and generate a plurality of candidate grasping poses based on the point cloud data; reordering the plurality of candidate grasping poses based on a physical consistency index to filter out an optimal grasping pose; the physical consistency index is calculated based on contact area point cloud data of a contact area between a gripper corresponding to each candidate pose and the target object; inputting a depth image of the contact area corresponding to the optimal grasping pose, physical condition parameters of the target object, and a current haptic observation image of the target object into a conditional diffusion model to predict an expected haptic image under a force-optimal stable grasping condition; According to the real-time obtained haptic image, the torque value of the gripper is iteratively calculated in the haptic latent space by using a linear quadratic regulator, and the gripper is controlled to grasp the target object by using the torque value, so that the current haptic image of the gripper approximates to the expected haptic image.

[0010] In a third aspect, a robot is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the robot grasping control method described above when executing the computer program.

[0011] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the robot grasping control method described above when executed by a processor.

[0012] Advantages of the present application: By introducing a physical consistency index based on contact area point cloud data to reorder the candidate grasping poses, a bridge between pose selection and force control is built. This design breaks through the limitations of traditional pure reliance on geometric closure criteria, making the selection of the optimal grasping pose not only consider geometric feasibility, but also fully consider the contact mechanics characteristics, ensuring that the grasping pose is conducive to achieving an optimized force distribution from the source.

[0013] An expected haptic image is generated using a conditional diffusion model, establishing a precise mapping from object physical properties to haptic goals. By using object physical condition parameters as constraint conditions for the generation model, it is ensured that the generated haptic goals not only meet the requirements of optimal stable grasping, but also meet the physical realizability, effectively solving the problem of lack of adaptability of fixed force strategy.

[0014] A linear quadratic regulator control is implemented in the haptic latent space to achieve closed-loop precise adjustment of the haptic state. This design directly operates in the latent space, reducing computational complexity, and through optimization control theory, it ensures rapid convergence of the haptic state with minimal control action, fundamentally overcoming the problems of traditional force regulation lag and high energy consumption.

[0015] The present application constitutes a complete technical chain from pose selection, target generation to force control, achieving the unity of stability and efficiency in the robot grasping process, and providing reliable technical support for diversified object grasping in complex scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0017] Among them: Figure 1 It is a structural schematic diagram of the robot provided in an embodiment; Figure 2 It is a flowchart of the robot grasping control method in an embodiment; Figure 3 It is a structural block diagram of the robot grasping control device in an embodiment; Figure 4 It is a structural block diagram of the robot in an embodiment. DETAILED DESCRIPTION

[0018] With reference to the accompanying drawings: the technical solutions in the embodiments of the present application will be apparently and completely described, obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by one of ordinary skill in the art without creative labor fall within the scope of the present application.

[0019] The structural schematic diagram of the robot provided by the embodiments of the present application comprises: The gripper 10, the mechanical arm body 11, the driving system 12, the sensing system and the control system Figure 1 are not drawn in the figure).

[0020] The gripper 10 is a grasping structure installed at the end of the mechanical arm, and is a part directly interacting with the grasped object. The gripper 10 can be a parallel two-finger gripper, which realizes grasping through the relative parallel movement of two fingers, and is suitable for regular objects such as boxes and blocks. The gripper 10 can also be a multi-fingered dexterous hand, which is designed by imitating human hands, has multiple fingers and joints, and can perform more complex grasping and operation, such as holding tools and grasping irregular objects.

[0021] The mechanical arm body 11 is usually a multi-degree-of-freedom serial or parallel mechanism composed of multiple links and joints, and is responsible for realizing the movement of the gripper in three-dimensional space. Its structure determines the workspace, flexibility and carrying capacity of the robot.

[0022] The driving system 12 provides power for the joints of the mechanical arm and the gripper. Its driving mode can be motor drive, pneumatic drive or hydraulic drive, etc. Motor drive includes servo motor and stepper motor, which is used to provide accurate position and torque control. Pneumatic drive is driven by air cylinder, which is usually used in occasions requiring fast and large power but not high precision. The hydraulic drive has larger output and is suitable for heavy robots.

[0023] The sensing system is used for the robot to perceive its own state and external environment, including: internal sensors and external sensors. The internal sensors are, for example, encoders (measuring joint angles) and torque sensors (measuring joint or end force). The external sensors are, for example, vision sensors (2D camera, 3D depth camera, used to obtain object position and shape), and tactile sensors (installed on the fingertips or surface of the gripper, used to measure the pressure distribution and force of contact with the object).

[0024] The control system is the "brain" of the robot, which is usually composed of a processor and a memory. The memory stores computer programs, such as control software. The control device receives the information of the sensors, calculates the action instructions of the mechanical arm and the gripper according to the preset algorithm program, and sends them to the driving system for execution.

[0025] The application will be described in detail below through specific embodiments.

[0026] Please refer to Figure 2 as shown, Figure 2 A flowchart of a robot grasping control method provided by an embodiment of the application includes the following steps: S1, obtaining point cloud data of a target object, and generating a plurality of candidate grasping poses based on the point cloud data.

[0027] Among them, the processor obtains the depth information and color image of the target working scene through the depth camera. The depth camera actively projects an optical pattern to the scene and receives the reflected signal, and obtains the distance value of each pixel point by calculating the deformation of the light spot or the flight time of the light. The processor registers and fuses the depth information and the color image to construct point cloud data describing the three-dimensional space coordinates of the object surface. The point cloud data is composed of a large number of space points, each point contains three-dimensional coordinate information, and collectively forms a digital expression of the geometric shape of the object surface.

[0028] After completing the construction of the point cloud data, the processor calls the pre-set grasping pose generation model to process the point cloud data. The grasping pose generation model is based on a deep learning network architecture, which analyzes the overall geometric profile and local shape features of the target object in the point cloud data, and combines the grasping mechanics principle to calculate a plurality of candidate grasping poses that are feasible in the geometric level. Each candidate grasping pose defines the spatial pose of the gripper in the three-dimensional space, including the position coordinates and the orientation angle, forming a set of initial grasping schemes for subsequent screening.

[0029] For example, when the robot needs to grasp an apple placed in a fruit tray, the depth camera first obtains the scene information containing the apple and its surrounding environment. After the processor fuses the depth information and the color image, a three-dimensional point cloud model of the apple is generated, which accurately reflects the spherical main structure of the apple, the stem feature of the top depression, and the continuous and smooth curvature change of the surface.

[0030] Subsequently, the grasping pose generation model analyzes the apple point cloud and may generate several different candidate grasping poses. The candidate grasping poses include a scheme of clamping from the side of the apple along the maximum circumference, a scheme of vertically grasping the apple from the top to avoid the stem, and a specific angle grasping scheme designed for the irregular shape of the apple. Each scheme provides complete gripper spatial pose parameters, laying a foundation for subsequent optimization and screening In one possible embodiment of the present application, S1, the plurality of candidate grasping poses are generated based on the point cloud data, comprising: S11, outputting a plurality of candidate grasping poses and the confidence of each candidate grasping pose through the GraspNet model after processing the point cloud data.

[0031] Specifically, the point cloud data obtained in the foregoing step is input into a pre-trained GraspNet model. The model is a grasp pose prediction network based on a deep learning architecture, and its internal structure includes a variational grasp sampler, a grasp pose evaluator, and an iterative grasp pose optimizer. The model first performs multi-level feature extraction on the input point cloud data to understand the overall geometric profile and local surface characteristics of the object. Then, it generates multiple possible gripper spatial pose schemes through its internal mechanism, and evaluates the stability and feasibility of each scheme.

[0032] After the GraspNet model is run, a set of structured grasp pose data is output. The data includes multiple candidate grasp pose parameters, each of which defines the position coordinates and rotation angles of the gripper in three-dimensional space. At the same time, the GraspNet model provides each candidate pose with a corresponding confidence score, which reflects the model's probability estimate of successfully achieving stable grasping based on learned grasp priors.

[0033] The training process of the GraspNet model is described below: The training of the GraspNet model is based on a large-scale grasp dataset, which contains a large amount of point cloud data of different categories of objects and their corresponding real grasp pose labels. The training data is constructed in the following way: use a depth camera to collect three-dimensional point cloud information of tens of thousands of different objects, each point cloud sample contains data of the object in multiple poses. At the same time, the data acquisition system records the actual spatial pose of the gripper when the robot successfully grasps, including the position coordinates, rotation angles, and finger opening width of the gripper base, which constitutes the grasp pose true value required for training.

[0034] In the data preprocessing stage, the original point cloud data is downsampled and normalized to ensure the regularity of the input data. At the same time, the grasp pose true value is unified in the coordinate system and parameterized to adapt to the output format requirements of the model. The dataset is divided into training set, validation set and test set according to the preset proportion, among which the training set is used for model parameter learning, the validation set is used for hyperparameter tuning, and the test set is used for final performance evaluation.

[0035] The training of the GraspNet model adopts an end-to-end learning method, which gradually improves the grasp pose generation ability of the model through a multi-stage optimization strategy. The training process first initializes the network parameters of the model, including the encoder-decoder weights in the variational grasp sampler, the classification network parameters of the grasp pose evaluator, and the gradient update module of the iterative optimizer.

[0036] Model training uses a combined loss function, which comprises three main components: a variational sampling loss function that measures the difference in distribution between the generated and actual grasping poses; a pose evaluation loss function that assesses the accuracy of confidence predictions; and an iterative optimization loss function that ensures the convergence of the pose fine-tuning process. During training, the stochastic gradient descent algorithm and its variant optimizer are employed. The gradient of the loss function with respect to the model's parameters at each layer is calculated using backpropagation, and the network weights are iteratively updated.

[0037] An early stopping mechanism is implemented during training. Training is terminated when the performance metrics on the validation set no longer improve after several consecutive training epochs to prevent overfitting. A dynamic learning rate adjustment strategy is also employed, using a larger learning rate in the early stages of training to accelerate convergence and gradually decreasing it in later stages to improve optimization accuracy. Data augmentation techniques are also used during training, randomly rotating, translating, and adding noise to the point cloud data to enhance the model's generalization ability and robustness.

[0038] After training, the model is comprehensively evaluated on an independent test set. Test metrics include the success rate of grasping pose generation, stability score, and spatial error compared to the actual grasping pose. A fully trained GraspNet model can generate candidate grasping poses with high geometric and mechanical feasibility based on input point cloud data, and assign an appropriate confidence score to each pose, providing a reliable basis for subsequent grasping decisions.

[0039] S2. Based on the physical consistency index, the multiple candidate grasping postures are reordered to select the optimal grasping posture.

[0040] In this process, the processor performs refined evaluation and screening of multiple candidate grasping postures obtained in step S1. For each candidate posture, the processor first extracts the contact area point cloud data, specifically calculating the local 3D point cloud set where the gripper fingertip surface and the target object surface are expected to make contact under that posture. Then, through spatial coordinate transformation, the extracted contact area point cloud data is uniformly converted to the local coordinate system of the gripper fingertip, and points located inside the gripper's mechanical structure or obviously unreasonable spatial points are filtered out.

[0041] Based on the standardized contact area point cloud, the processor calculates a set of physical consistency indices. These indices are a parameter system used to quantitatively evaluate the mechanical rationality of the grasping posture. This system predicts the stability and efficiency of the grasping action by mathematically modeling and statistically analyzing the geometric characteristics of the contact area between the gripper and the object. These indices are derived from the principles of contact mechanics and can identify posture defects that may lead to grasping failure or inefficiency in advance.

[0042] The surface roughness index reflects the flatness of the contact area surface, and is obtained by calculating the variance of the normal vectors of adjacent points in the contact point cloud. The higher the index value, the more uneven the contact surface, and the more likely stress concentration occurs during grasping, affecting the stability of grasping. The normal consistency index measures the degree of coordination of the normal direction of the contact surface, which is quantified by the distribution of the main normal direction of the contact point. The lower the index value, the more consistent the direction of the applied grasping force, and the higher the force transmission efficiency. The curvature uniformity index represents the uniformity of the curvature distribution of the contact area, which is evaluated by analyzing the standard deviation of the principal curvature of the contact surface. The higher the index value, the more uniform the pressure distribution during grasping, and the less likely local slip occurs.

[0043] The processor weights and fuses the above-mentioned physical consistency index and the original confidence output by the GraspNet model, wherein the weight coefficients of each index are obtained through pre-experiment calibration. After fusion, a comprehensive evaluation score of each candidate posture is generated, and the processor ranks all candidate postures in descending order according to the score, and selects the highest ranked posture as the final optimal grasping posture.

[0044] It should be noted that when the target object is of rigid material, there may be a microscopic gap between the gripper fingertips and the object surface without physical contact during the grasping posture planning stage. In such cases, the contact area point cloud data described in the present application refers to the point cloud data set corresponding to the projection area of the gripper fingertip surface on the target object surface. The projection area is determined by a geometric projection algorithm, which specifically projects the contact surface of the gripper fingertip onto the target object surface along its normal direction to form a theoretical contact area.

[0045] The processor constructs the projected contact area by calculating the spatial position relationship between the gripper fingertip model and the object point cloud. The projected contact area accurately reflects the most likely contact range between the gripper and the object in the ideal closed state. The point cloud data extracted based on this projected contact area retains the geometric features and curvature distribution of the contact area, which can fully support the calculation requirements of the subsequent physical consistency index.

[0046] S3, taking the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, predicting the expected tactile image under the condition of force-optimal stable grasping through a conditional diffusion model.

[0047] Wherein, the processor transmits a plurality of input parameters to the pre-trained conditional diffusion model when performing this step. These input parameters include the optimal grasping posture data optimized and screened in step S2, the depth image of the contact area corresponding to the posture, the physical conditions of the target object obtained by measuring or querying the database, and the current tactile observation image collected by the tactile sensor in real time.

[0048] wherein the physical condition parameters are a set of quantified features describing the inherent physical properties of the target object, which directly affect the mechanical interaction characteristics in the grasping process. The physical condition parameters involved in this scheme mainly include two core elements: material type and mass value.

[0049] The material type parameter characterizes the mechanical properties of the object surface, representing the physical properties of different materials through classification coding. This parameter covers mechanical characteristics such as rigidity and flexibility, smoothness and roughness, elasticity and plasticity, directly affecting the distribution of contact pressure and the required grasping force. The mass value parameter characterizes the inertia characteristics of the object, quantified by standard mass units. This parameter is directly related to the support force required in the grasping process and is an important basis for determining the basic grasping force range.

[0050] The acquisition of physical condition parameters is achieved through various sensing methods. The material type can be obtained through visual texture analysis, spectral detection, or a pre-set object database. The mass value can be directly measured by a weighing sensor or indirectly obtained based on the volume and density estimation of the object. The processor standardizes the raw parameters obtained to ensure that the parameter values have a uniform dimension and value range, facilitating subsequent feature fusion and processing. The physical condition parameters provide necessary physical constraints for haptic target generation. The material type determines the mechanical behavior of the contact interface and directly affects the pressure distribution pattern in the desired haptic image. The mass value determines the minimum grasping force range required to maintain stable grasping, ensuring that the generated haptic target meets the stability requirements and avoids excessive force.

[0051] The condition diffusion model is a generative neural network based on physical law constraints, whose core architecture includes an encoder module, a latent space diffusion module, and a decoder module. The encoder module first maps the input haptic observation image and contact area depth image to feature vectors in the latent space. The physical condition parameters are converted to feature vectors through an embedding layer and injected into the model's inference process through cross-attention mechanisms.

[0052] In the latent space diffusion module, the model performs an iterative denoising inference process. Starting with randomly initialized Gaussian noise, under the guidance of the physical condition vector, the process gradually corrects the numerical distribution of the latent space vector through multiple rounds of feature extraction and fusion, making it converge to a feature representation that meets the optimal stability grasping condition. Finally, the decoder module reconstructs the optimized latent space vector into the desired haptic image.

[0053] The training of the conditional diffusion model is based on a specially constructed physical condition haptic dataset. The dataset contains multiple sets of interrelated data samples, each sample consisting of four elements: the physical condition parameters of the target object, the contact area depth image corresponding to the grasping pose, the haptic observation image collected by the tactile sensor, and the real haptic image in the force-optimal stable grasping state. The physical condition parameters are obtained by precise instruments, including object material type and mass value. The contact area depth image is derived from the depth information conversion of point cloud data, and the haptic image is collected by a high-resolution tactile sensor in a controlled experimental environment.

[0054] The dataset covers 34 different categories of objects, covering various physical properties such as rigid bodies, flexible bodies, and viscoelastic bodies, with a total of 8600 samples. In the data preprocessing stage, all haptic images are standardized, including size normalization, grayscale value calibration, and noise filtering. The contact area depth image is unified in coordinates and adjusted in resolution to ensure spatial alignment with the haptic image. The physical condition parameters are normalized and encoded into machine-readable feature vectors.

[0055] The model training is divided into two main stages. The first stage trains the variational autoencoder, which is responsible for compressing high-dimensional haptic images into low-dimensional latent space representations. During training, the mean square error loss function and the KL divergence loss function are used to ensure that the latent space features not only retain the key information of the original image but also conform to the Gaussian distribution characteristics.

[0056] The second stage trains the core denoising network of the latent space diffusion model. This network uses the U-Net architecture and embeds cross-attention modules inside to receive physical condition parameters. During training, Gaussian noise is gradually added to the latent space feature vectors, and the denoising network learns to predict the added noise at each step based on the physical condition parameters and the contact area depth image. The training loss function is defined as the mean square error between the predicted noise and the real noise, and the network parameters are iteratively optimized using the gradient descent algorithm.

[0057] The training process uses a phased strategy, first training the network's basic feature extraction capability at a fixed noise level, and then gradually expanding to the full noise scheduling range. The optimizer uses an adaptive learning rate algorithm to dynamically adjust the learning rate size based on the training progress. The model uses an early stopping mechanism during training, which automatically terminates training when the validation set loss does not decrease for multiple consecutive training periods, preventing overfitting.

[0058] The trained model was fully validated on an independent test set. Test metrics included structural similarity index of haptic image generation, peak signal-to-noise ratio, and perceptual loss function value. The model also underwent physical plausibility testing to verify that the generated haptic images conformed to the physical constraints of the object. For haptic images generated from objects with special materials, manual evaluation by mechanical experts was required to ensure that the pressure distribution pattern met theoretical expectations. For example, when a processor needs to grasp a ripe fruit, physical condition parameters indicate that the object has a soft texture and a relatively light mass. After receiving these parameters, the conditional diffusion model combines them with a current tactile observation image captured from the fruit's surface to generate a corresponding desired tactile image. This image exhibits a uniformly distributed low-pressure region characteristic, meeting the requirements for a gentle grasp.

[0059] Conversely, when handling a metal tool, the physical parameters indicate that the object has a hard texture and a large mass. The desired tactile image generated by the model based on these conditions shows several concentrated high-pressure areas, which correspond to the key contact points required to ensure gripping stability. In this way, the model can adaptively adjust the tactile expectations according to different object characteristics, providing a precise reference standard for subsequent gripping control.

[0060] S4. Based on the real-time acquired tactile image, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator, and the gripper is controlled to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image.

[0061] Specifically, the processor establishes a closed-loop servo control mechanism within the tactile latent space. This mechanism uses the desired tactile image generated by the conditional diffusion model as the control target, and converts it into a target latent space vector through a pre-trained encoder. During the grasping process, the tactile sensor continuously acquires real-time tactile images, and the processor synchronously encodes them into the current latent space vector.

[0062] The linear quadratic regulator constructs a state-space model within the latent space, defining the latent space vector error as the system state variable and the gripper torque adjustment as the control variable. The regulator estimates the system dynamic parameters online using recursive least squares, including the state transition matrix and the control action matrix. Based on the estimated system model, the regulator solves the algebraic Riccati equations to obtain the optimal control gain matrix.

[0063] In each control cycle, the processor calculates the error between the current latent space vector and the target vector, and then calculates the torque adjustment command based on the optimal control gain. This calculation process simultaneously considers the balance between error convergence speed and control energy consumption, achieving multi-objective optimization through a carefully designed weight matrix. The generated torque command drives the gripper to perform the corresponding action through the robot's drive system.

[0064] In one possible embodiment of this application, S2, the reordering of the plurality of candidate grasping postures based on the physical consistency index to select the optimal grasping posture, includes: S21. For each candidate grasping posture, extract the contact area point cloud data of the contact area between the gripper and the target object.

[0065] Specifically, the processor performs point cloud data extraction of the contact area for each candidate grasping posture. This operation is based on the spatial positional relationship between the geometric model of the gripper fingertips and the point cloud data of the target object. The processor first determines the specific pose of the gripper in three-dimensional space based on the candidate posture parameters, and then identifies the object surface areas where contact may occur by calculating the minimum spatial distance between the surface of the gripper fingertips and the object point cloud.

[0066] After identifying the potential contact area, the processor uses a point cloud spatial query algorithm to extract the set of points corresponding to that area from the complete point cloud data of the target object. These spatial points together constitute the contact area point cloud data, and their distribution accurately reflects the expected contact pattern between the gripper and the object surface under this grasping posture. The extraction process ensures the integrity and geometric accuracy of the point cloud data, providing a reliable data foundation for subsequent physical property analysis.

[0067] S22. Align the point cloud data of the contact area to the coordinate system of the gripper fingertip through rigid transformation.

[0068] Specifically, the processor performs a spatial coordinate system transformation operation on the contact area point cloud data. This operation transforms the extracted contact area point cloud data from its original coordinate system to the local coordinate system of the gripper fingertip by applying a rigid transformation. The rigid transformation is defined by a rotation matrix and a translation vector, where the rotation matrix is ​​determined by the orientation angle of the gripper in the candidate grasping posture, and the translation vector is determined by the spatial position coordinates of the gripper fingertip.

[0069] During the coordinate system transformation, the processor first constructs a transformation matrix from the original coordinate system of the point cloud to the coordinate system of the gripper fingertip. This transformation matrix is ​​then applied to each data point in the contact area point cloud, achieving batch coordinate transformation of all point cloud data. This process maintains the relative spatial relationships between the data points in the point cloud, only changing the reference datum for their coordinate representation. After the transformation, the point cloud data uses the geometric center of the gripper fingertip as the coordinate origin and the main structural direction of the fingertip as the coordinate axis direction, forming a standardized spatial data representation.

[0070] S23. Calculate one or more physical consistency indices for the contact area point cloud data in the gripper index coordinate system; the physical consistency indices include at least one of surface roughness, normal consistency, and curvature uniformity.

[0071] Specifically, the surface roughness index is used to quantify the unevenness of the contact area surface. This index reflects the microscopic undulations in the local geometric features of the contact area, and its value is negatively correlated with surface smoothness. Excessive surface roughness can lead to stress concentration at the contact point, affecting the stability and reliability of the gripping action.

[0072] During the calculation, the processor first performs local neighborhood analysis on the point cloud of the contact area, calculating the surface normal vector for each data point. Then, it statistically analyzes the deviation of the normal vector direction from the average normal vector of the local area, quantifying the surface roughness by calculating the variance of the angle between the normal vectors. Specifically, the processor selects a set of neighboring points within a predetermined radius around each point, obtains the normal vector direction of that point through principal component analysis, and finally calculates the statistical average of the variances of the normal vector directions of all points to obtain the comprehensive surface roughness evaluation value.

[0073] The normal consistency index is used to evaluate the degree of coordination of the normal directions of the contact area surfaces. This index reflects the consistency of force transmission direction at each contact point during the gripping process, and its value directly affects the efficiency of gripping force transmission. The higher the normal consistency, the more uniform the direction of the applied gripping force, and the better the force transmission efficiency.

[0074] The calculation method is based on the distribution characteristics of the surface normal direction at all points in the contact area. The processor first calculates the average normal direction of the entire contact area, and then statistically analyzes the deviation angle of the normal direction at each point from the average direction. By analyzing the dispersion of these deviation angles, especially calculating their standard deviation and extreme value range, a quantitative index of normal consistency is obtained. This index comprehensively considers the central tendency and dispersion of the normal direction, and can accurately reflect the cooperative efficiency of the force transmission direction.

[0075] Curvature uniformity is used to characterize the evenness of the surface curvature distribution in the contact area. This index predicts the pressure distribution during the gripping process by analyzing the principal curvature distribution characteristics of the contact area. Higher curvature uniformity indicates a more uniform pressure distribution and better gripping stability.

[0076] During the calculation, the processor performs differential geometric analysis on the point cloud of the contact area, calculating the principal curvature value of each point using a local surface fitting method. Subsequently, the distribution characteristics of the principal curvature values ​​at all points are statistically analyzed, including the range, variance, and distribution entropy of the curvature values. The curvature uniformity index comprehensively considers these statistical characteristics, paying particular attention to the dispersion and variation patterns of the curvature distribution. Through entropy analysis of the principal curvature distribution, the uniformity level of the curvature distribution in the contact area can be accurately assessed, providing a reliable basis for predicting pressure distribution.

[0077] S24. The physical consistency index is weighted and fused with the confidence level of the candidate grasping posture to obtain a comprehensive score.

[0078] Specifically, the processor performs a weighted fusion calculation of the physical consistency metric and the confidence level of the candidate grasping posture. This calculation process is based on a pre-defined weight allocation scheme, integrating evaluation metrics from different dimensions into a unified comprehensive score. The weight coefficients were optimized and determined through extensive experimental data, reflecting the importance of each metric in the grasping stability assessment.

[0079] In the specific calculation process, the processor first standardizes each physical consistency index to eliminate dimensional differences and ensure data comparability. Then, the standardized physical consistency indexes are linearly combined with the normalized grasping posture confidence score according to predetermined weights. The sum of the overall weight of the physical consistency indexes and the weight of the grasping posture confidence score is a complete unit value, ensuring a reasonable range for the overall score.

[0080] The weighted fusion formula is a weighted sum of the physical consistency index function and the grasping posture confidence. The physical consistency index function itself is a weighted combination of three sub-indicators: surface roughness, normal consistency, and curvature uniformity. This hierarchical weighting design preserves the independent influence of each sub-indicator while effectively integrating multi-dimensional evaluations.

[0081] The formula for calculating the overall score W_p is: W_p = δ·(1-S) + (1-δ)·F(P_t); Where F(P_t) =α·S_rough +β·(1-C_N) +γ·U_C; The parameters in the formula are defined as follows: S represents the crawl confidence of the GraspNet model output, S_rough represents the surface roughness index, C_N represents the normal consistency index, U_C represents the curvature uniformity index, and α, β, γ, and δ represent the weight coefficients, for example: α=0.2, β=0.6, γ=0.2, δ=0.5.

[0082] The weighting coefficients α, β, γ, and δ were determined through a systematic experimental optimization process. This process, based on a large amount of crawling experimental data, employs a multi-objective optimization method, using the crawling success rate and the force-optimal stable crawling achievement rate as the core evaluation indicators.

[0083] The experiment first constructs a representative set of test objects, containing typical objects with different geometries, material properties, and mass distributions. The test set covers objects with different physical properties, including rigid, flexible, and viscoelastic bodies, ensuring that the weighting coefficients have good generalization ability. In each experimental cycle, the system executes a standardized grasping test sequence under a specific combination of weighting coefficients. By changing the combination of weighting coefficient values, the system records the grasping performance under different weighting configurations.

[0084] The weight optimization process employs a phased optimization strategy. First, the relative weights within the physical consistency index are determined. By analyzing extensive crawling experiment data, a correlation model between each physical index and crawling stability is established. Based on this, the balanced weights between the comprehensive physical consistency index and network confidence are further optimized.

[0085] The optimization process employs the response surface methodology, gradually approximating the optimal weight combination through multiple rounds of experiments. In each round, the search direction and step size of the weight coefficients are adjusted based on the results of previous experiments, ultimately obtaining the range of weight coefficient values ​​that maximizes crawling performance.

[0086] The final weighting coefficients underwent rigorous cross-validation and statistical significance testing. Validation results on the independent test set show that, by using the optimized combination of weighting coefficients, the system can achieve the best balance between data-driven prediction and physical constraints, significantly improving the reliability and adaptability of the crawling system.

[0087] The specific values ​​of the weighting coefficients are adjusted within a certain range according to the actual application scenario and the requirements of the crawling task, but their relative magnitudes remain stable. Among them, the normal consistency index is usually given a higher weight, while the surface roughness and curvature uniformity indexes have relatively lower weights. A proper balance is maintained between the network confidence and physical consistency indexes.

[0088] S25. Based on the comprehensive score, sort the multiple candidate grasping postures in ascending or descending order, and select the one with the best score as the optimal grasping posture.

[0089] Specifically, after obtaining the overall score W_p for all candidate grasping poses, the processor sorts them in ascending order according to the score value. This sorting method ensures that candidate poses with lower overall scores are placed at the beginning of the sequence, while candidate poses with higher overall scores are placed at the end of the sequence. Under this sorting mechanism, the candidate pose with the lowest score has the optimal force distribution characteristics.

[0090] When selecting the optimal grasping posture, the processor chooses the candidate posture that is first in the sorted sequence, i.e., the posture with the smallest overall score W_p value, as the final optimal grasping posture. This selection mechanism ensures that the selected posture has the most favorable force distribution in terms of contact geometry.

[0091] Taking an apple grasping task as an example, suppose the W_p values ​​of the five candidate grasping postures, after comprehensive scoring, are 0.15, 0.22, 0.31, 0.45, and 0.58, respectively. The processor sorts these postures according to their W_p values ​​from smallest to largest, forming an ordered sequence. The side gripping posture with a score of 0.15 is at the top of the sequence, followed by the horizontal gripping posture with a score of 0.22, and the remaining postures are arranged in order. The processor selects the side gripping posture with a W_p value of 0.15 as the optimal grasping posture, which has been proven to achieve the best force distribution effect in subsequent actual grasping processes.

[0092] This technical solution effectively compensates for the shortcomings of relying solely on data-driven models by introducing a physical consistency index to reorder candidate grasping postures. This optimization and selection mechanism based on physical laws significantly improves the scientific rigor and reliability of grasping posture selection, ensuring that the final selected optimal grasping posture not only meets geometric feasibility but also conforms to mechanical rationality.

[0093] In one possible embodiment of this application, S3, the step of predicting the desired tactile image under force-optimal stable grasping conditions using a conditional diffusion model, with the depth image of the contact area, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, includes: S31. Encode the current tactile observation image into a first latent space vector using a VAE encoder; S32. Encode the depth image into a second latent space vector using the VAE encoder; S33. The physical condition parameters of the target object are fused into a condition vector; S34. Determine the weight coefficients of the first latent space vector, the second latent space vector, and the conditional vector based on the attention mechanism; S35. Based on the weight coefficients, the first latent space vector, the second latent space vector, and the condition vector are fused to obtain a fused feature vector; S36. Using the fused feature vector as input, generate a tactile target vector using a pre-trained conditional diffusion model; S37. Encode the tactile target vector into a desired tactile image.

[0094] In step S31, the processor uses the encoder component in a pre-trained variational autoencoder to extract and compress features from the current tactile observation image. This encoder extracts deep features from the tactile image through a multi-layer convolutional neural network, mapping them to a first latent space vector in a low-dimensional space. This process preserves key information related to pressure distribution in the tactile image while eliminating unnecessary details and noise.

[0095] In step S32, the processor uses the same variational autoencoder architecture to encode the depth image of the contact area corresponding to the optimal grasping posture. The depth image records the three-dimensional geometric shape information of the contact area, and the encoder converts it into a second latent space vector through feature extraction. This vector represents the geometric characteristics of the contact area, such as curvature distribution and surface orientation, in a compact form.

[0096] In step S33, the processor performs numerical processing and feature fusion on the physical condition parameters of the target object. The physical condition parameters include material type and mass value. First, the material type is converted into a numerical vector through an embedding layer, and the mass value is normalized before being input into a fully connected layer. Then, these two feature vectors are concatenated and linearly transformed to generate a unified condition vector.

[0097] For example, for a metal object with a mass of 0.2 kg, the processor maps the metal material to a specific embedding vector, normalizes the mass value to a standard value, and after fusion, generates a conditional vector that comprehensively represents the physical properties of the object. This vector is significantly different from the conditional vector generated for a sponge object with a mass of 0.1 kg.

[0098] In step S34, the processor analyzes the intrinsic relationships between the first latent space vector, the second latent space vector, and the conditional vector based on an attention mechanism, and adaptively calculates the weight coefficients of each vector. This mechanism obtains the attention weights by calculating the similarity score between the query vector and the key vector, and then normalizing it using a softmax function. These weight coefficients reflect the importance of different input features to the final tactile target generation.

[0099] The processor employs a cross-attention mechanism to calculate the weight coefficients of the first latent space vector, the second latent space vector, and the conditional vector. This calculation process is based on a similarity metric in the feature space and achieves weight allocation by constructing a mapping relationship between the query vector, key vector, and value vector.

[0100] The specific calculation process first uses the condition vector as the query vector, and the first and second latent space vectors together as the key and value vectors. The processor maps the query vector to the query space and the key vectors to the key space through a linear transformation layer. Then, the dot product similarity between the query vector and each key vector is calculated to obtain the initial attention score.

[0101] These attention scores, after being adjusted by a scaling factor, are normalized using a softmax function to generate the final attention weight coefficients. The magnitude of the weight coefficients reflects the correlation between each input feature and the current physical conditions, and the sum of the coefficients is a fixed value of 1. The entire calculation process ensures a reasonable distribution and interpretability of the weight coefficients.

[0102] For example, taking the apple-grabbing task as an example, the processor uses the apple's physical condition vector as the query vector and the first latent space vector of the current tactile observation and the second latent space vector of the contact area depth image as the key vectors. Calculations show that the second latent space vector has a high similarity to the query vector, thus receiving a weight coefficient of 0.6, while the first latent space vector receives a weight coefficient of 0.4. This indicates that in the current grabbing scenario, the geometric features of the contact area have a more significant impact on tactile target generation than real-time tactile observation.

[0103] Taking sponge block grasping as another example, due to the deformable nature of the sponge, the first latent space vector of the current tactile observation has a higher correlation with the physical condition vector, thus obtaining a weight coefficient of 0.7, while the second latent space vector only obtains a weight coefficient of 0.3. This adaptive weight allocation ensures that the system can intelligently adjust the importance of each input feature according to different object characteristics and grasping states.

[0104] In step S35, the processor performs a weighted fusion of the three input vectors based on the calculated attention weights. The fusion process employs an element-wise weighted summation method to ensure that the feature information of each vector is reasonably integrated. The final generated fused feature vector simultaneously contains information from tactile observation, geometric features, and physical conditions, providing comprehensive input for subsequent tactile target generation.

[0105] In step S36, the processor inputs the fused feature vector into the pre-trained conditional diffusion model and generates a tactile target vector through an iterative denoising process. The conditional diffusion model performs multiple rounds of diffusion operations in the latent space. In each round, it gradually optimizes the vector representation based on the conditional information provided by the fused feature vector, and finally outputs a tactile target vector that meets the requirements of optimal force and stable grasping.

[0106] In step S37, the processor uses the decoder component of a variational autoencoder to reconstruct the tactile target vector into the desired tactile image. The decoder progressively upsamples the low-dimensional latent space vector into a complete tactile image through a deconvolutional neural network layer, which clearly shows the pressure distribution pattern expected under force-optimal stable gripping conditions.

[0107] This technical solution achieves accurate mapping from the physical characteristics of an object to the desired tactile image through multi-source information fusion and conditional generation technology. This solution effectively solves the problem of reliance on experience in setting tactile targets in traditional grasping control, generating tactile expectations that conform to physical laws through a data-driven approach, thus providing a reliable technical foundation for achieving optimal force-based stable grasping.

[0108] In one possible embodiment of this application, S4, the step of iteratively calculating the torque value of the gripper in the tactile latent space using a linear quadratic modulator based on the real-time acquired tactile image, includes: S41. Encode the real-time acquired tactile image into a current latent space vector.

[0109] Specifically, the processor performs feature extraction and compression on real-time acquired tactile images using a pre-trained variational autoencoder. This encoder, composed of multiple convolutional neural networks, extracts spatial features from the tactile image layer by layer and maps high-dimensional tactile data to a low-dimensional latent space. During encoding, the processor first performs standardization preprocessing on the input tactile image, adjusting the image size and numerical range. Then, through forward propagation calculation by the encoder, it finally outputs a current latent space vector representing the core features of the tactile image. This vector preserves key information about the pressure distribution in the original tactile image while eliminating sensor noise and redundant details.

[0110] S42. Calculate the error between the current latent space vector and the target latent space vector; the target latent space vector is obtained by encoding the desired tactile image.

[0111] Specifically, the processor performs vector error calculation in the tactile latent space. This operation uses the current latent space vector obtained from real-time tactile image encoding and the target latent space vector obtained from the desired tactile image encoding as input data. The processor employs a vector distance metric algorithm to calculate the differences between the two latent space vectors across various feature dimensions, obtaining a comprehensive error quantification result. This error calculation process comprehensively considers the Euclidean distance and cosine similarity between the vectors, ensuring a comprehensive reflection of the degree of difference between the actual tactile state and the desired tactile state.

[0112] S43. Based on the error, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator.

[0113] Specifically, the processor constructs a linear quadratic regulator in the tactile latent space. First, state representation is performed using a pre-trained variational autoencoder to encode the current tactile image x_c into a current latent space vector z_c, and the desired tactile image into a target latent space vector z_g. The processor calculates the latent error e_c = diag(1 / S_scale)·(z_c - z_g), where S_scale is a normalized scaling vector calculated from the training data distribution to ensure consistent dimensions across error dimensions.

[0114] In the dynamic modeling phase, a linear discrete-time state-space model is established: e_{c+1} = A e_c + B Δu_c + d + w_c. Here, e_{c+1} represents the error state at the next time step, Δu_c is the gripper displacement increment used as the control input, d is the constant bias, and w_c is the process noise. The system dynamic parameter matrices A and B are estimated online using a recursive least squares algorithm. This algorithm continuously updates the model parameters based on real-time acquired system state data, with initial parameter values ​​learned from historical operating data.

[0115] In the control law design process, a quadratic cost function is defined as: J = Σ_{t=0}^∞ (e_c^TQ e_c + Δu_c^TR Δu_c), where Q ≥ 0 and R>0 are the weight matrices. The Q matrix emphasizes error minimization, and the R matrix emphasizes the economy of control actions. The optimal control gain matrix K is obtained by solving the algebraic Riccati equation, leading to the optimal control law Δu_c = -K e_c. This control output drives the gripper to adjust its grasping force, causing the tactile state to converge towards the target state.

[0116] S44. When the preset iteration exit condition is met, the clamping is set to hold mode; wherein, the cost function of the linear quadratic regulator is configured to simultaneously minimize the error and the movement amplitude of the gripper.

[0117] Specifically, the processor continuously monitors the error state in the tactile latent space. When the preset iteration exit condition is met, it automatically switches the gripper control mode from active adjustment to holding mode. The iteration exit condition is based on the norm of the error vector. When the error norm remains below a predetermined threshold for a specific time window, the system determines that the optimal force stable gripping state has been reached. At this point, the processor stops the iterative calculation of the torque value, maintains the current gripper configuration parameters, and enters a stable gripping and holding phase.

[0118] The cost function of the linear quadratic regulator is specially configured to simultaneously optimize two key indicators during the control process. On the one hand, it aims to minimize the tactile latent space error, ensuring that the actual tactile state converges quickly to the target state. On the other hand, it strictly controls the gripper's movement amplitude to avoid energy loss and mechanical wear caused by over-adjustment. This dual-objective optimization design enables the system to achieve both grasping accuracy and economic efficiency and reliability in the control process.

[0119] Furthermore, the iteration exit condition is set based on the stability theory of control systems and reliability considerations in engineering practice. This condition includes two key parameters: the error norm threshold and the stability duration, which together constitute the necessary and sufficient condition for determining that the system has reached a stable state.

[0120] The error norm reflects the combined distance between the current tactile state and the target state in a multi-dimensional feature space. When the error norm is below a preset threshold, it indicates that the actual output of the system has entered the neighborhood of the target state, and the control benefit of further adjustments will be significantly reduced. However, threshold attainment at a single moment may be due to random factors or measurement noise, so a duration requirement needs to be introduced for verification.

[0121] The duration requirement, based on statistical process control principles, ensures that the system reaches a sustained steady state rather than experiencing momentary fluctuations. This design demands that the system remain stable over multiple consecutive control cycles, effectively filtering out random disturbances and avoiding misjudgments caused by momentary fluctuations. The duration is set based on the system's time constant and control cycle to ensure the capture of the system's true dynamic characteristics.

[0122] During the control process, the processor calculates the L2 norm of the current latent space error vector in each sampling period. This value comprehensively represents the overall level of error across all dimensions. The processor maintains a continuous counter, which increments when the error norm is below a threshold and resets to zero when it exceeds the threshold. The iteration exit condition is triggered only when the accumulated counter value reaches the number of periods corresponding to a preset duration.

[0123] This dual verification mechanism ensures the reliability of the state determination. The threshold parameter is determined comprehensively based on the accuracy of the tactile sensor, the system noise level, and the control accuracy requirements, and is usually taken as the boundary value after the system enters the stable region. The duration parameter takes into account the system's response speed and anti-interference requirements, and is generally taken as several times the time constant of the system's main dynamic process.

[0124] Taking the grasping of a sponge block as an example, the error norm threshold is set to 0.15, and the duration is 5 control cycles. When the actual error norm gradually decreases from the initial value of 0.8 to 0.12 and remains stable for 5 consecutive cycles, the system confirms that a stable grasping state has been reached. During this process, even if the error briefly rises to 0.18 in a single cycle, the system continues to perform adjustment control because the duration requirement is not met, ensuring that a true stable state is eventually reached.

[0125] The design principle of this iterative exit condition ensures the accuracy and reliability of control mode switching, avoiding both insufficient control precision due to premature exit and energy waste and equipment wear caused by over-adjustment. Through rigorous mathematical judgment and engineering verification, it provides a stable and reliable termination judgment basis for the grasping system. This technical solution achieves precise optimization and adjustment of gripper torque through a closed-loop control mechanism based on tactile latent space. This solution effectively solves the problems of force adjustment lag and excessive squeezing in traditional gripping control, ensuring rapid convergence of the tactile state with minimal control actions through iterative optimization algorithms. This servo control method based on latent space error significantly improves the force control accuracy and energy utilization efficiency of the gripping process, providing a reliable technical guarantee for achieving optimal force-based stable gripping.

[0126] Please see Figure 3 As shown, in one embodiment, a robot grasping control device is provided, the device comprising: The acquisition module 301 is used to acquire point cloud data of the target object and generate multiple candidate grasping postures based on the point cloud data. The filtering module 302 is used to reorder the multiple candidate grasping postures based on the physical consistency index in order to filter out the optimal grasping posture; the physical consistency index is calculated based on the contact area point cloud data of the contact area between the gripper and the target object corresponding to each candidate posture. The prediction module 303 is used to predict the desired tactile image under the optimal force-stable grasping condition by taking the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as inputs, and through a conditional diffusion model. Control module 304 is used to iteratively calculate the torque value of the gripper in the tactile latent space using a linear quadratic regulator based on the real-time acquired tactile image, and to control the gripper to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image. In one possible embodiment, generating multiple candidate grasping poses based on the point cloud data includes: After processing the point cloud data using the GraspNet model, multiple candidate grasping poses and the confidence scores of each candidate grasping pose are output.

[0127] In one possible embodiment, the step of reordering the plurality of candidate grasping postures based on a physical consistency metric to select the optimal grasping posture includes: For each candidate grasping posture, extract the contact area point cloud data of the contact area between the gripper and the target object; The point cloud data of the contact area is aligned to the coordinate system of the gripper fingertip through a rigid transformation; In the gripper index coordinate system, one or more physical consistency indices of the contact area point cloud data are calculated; the physical consistency indices include at least one of surface roughness, normal consistency and curvature uniformity. The physical consistency index is weighted and fused with the confidence level of the candidate grasping posture to obtain a comprehensive score; The candidate grasping postures are sorted in ascending or descending order based on the comprehensive score, and the one with the best score is selected as the optimal grasping posture.

[0128] In one possible embodiment, the surface roughness is quantified by calculating the normal variance or local curvature variation of the contact region point cloud data; the normal uniformity is quantified by calculating the variance of the surface normal angle of the contact region point cloud data; and the curvature uniformity is evaluated by calculating the entropy value of the curvature distribution of the contact region point cloud data.

[0129] In one possible embodiment, the step of predicting the desired tactile image under force-optimal stable grasping conditions using a conditional diffusion model, taking the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, includes: The current tactile observation image is encoded into a first latent space vector using a VAE encoder; The VAE encoder is used to encode the depth image of the contact area corresponding to the optimal grasping posture into a second latent space vector. The physical condition parameters of the target object are fused into a condition vector; The weight coefficients of the first latent space vector, the second latent space vector, and the conditional vector are determined based on the attention mechanism. Based on the weight coefficients, the first latent space vector, the second latent space vector, and the conditional vector are fused to obtain a fused feature vector; The fused feature vector is used as input to generate a tactile target vector using a pre-trained conditional diffusion model; The tactile target vector is encoded into a desired tactile image.

[0130] In one possible embodiment, the step of iteratively calculating the torque value of the gripper in the tactile latent space using a linear quadratic modulator based on the real-time acquired tactile image includes: The real-time acquired tactile image is encoded into a current latent space vector; Calculate the error between the current latent space vector and the target latent space vector; the target latent space vector is obtained by encoding the desired tactile image; Based on the error, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator; When the preset iteration exit condition is met, the clamping is set to hold mode; wherein, the cost function of the linear quadratic regulator is configured to simultaneously minimize the error and the movement amplitude of the gripper.

[0131] In one possible embodiment, the iteration exit condition is: Monitor the norm of the error in real time; When the norm of the error is lower than a preset threshold and continues for a preset duration, it is determined that the iteration exit condition is met.

[0132] In one embodiment, a structural schematic diagram of a robot is provided, the internal structure of which can be as follows: Figure 4 As shown, the robot includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the control device is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a robot grasping control method.

[0133] In one embodiment, a robot is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps: Acquire point cloud data of the target object, and generate multiple candidate grasping postures based on the point cloud data; Based on the physical consistency index, the multiple candidate grasping postures are reordered to select the optimal grasping posture; the physical consistency index is calculated based on the point cloud data of the contact area between the gripper and the target object corresponding to each candidate posture. Using the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, the desired tactile image under the optimal force-stable grasping condition is predicted through a conditional diffusion model. Based on the real-time acquired tactile images, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator, and the gripper is controlled to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image.

[0134] This technical solution optimizes and filters candidate grasping postures by introducing a physical consistency index, combines it with a tactile target prediction method based on physical constraints, and employs a latent space optimal control strategy to achieve an organic unity between posture selection and force control during the grasping process. This solution effectively improves the grasping system's adaptability to the physical characteristics of objects, significantly reducing grasping force requirements while ensuring grasping stability. It avoids object damage or grasping failure caused by force control lag or inappropriate force values ​​in traditional methods, providing a reliable technical guarantee for precise robotic grasping operations.

[0135] In one embodiment, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, performs the following steps: Acquire point cloud data of the target object, and generate multiple candidate grasping postures based on the point cloud data; Based on the physical consistency index, the multiple candidate grasping postures are reordered to select the optimal grasping posture; the physical consistency index is calculated based on the point cloud data of the contact area between the gripper and the target object corresponding to each candidate posture. Using the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, the desired tactile image under the optimal force-stable grasping condition is predicted through a conditional diffusion model. Based on the real-time acquired tactile images, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator, and the gripper is controlled to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image.

[0136] This technical solution optimizes and filters candidate grasping postures by introducing a physical consistency index, combines it with a tactile target prediction method based on physical constraints, and employs a latent space optimal control strategy to achieve an organic unity between posture selection and force control during the grasping process. This solution effectively improves the grasping system's adaptability to the physical characteristics of objects, significantly reducing grasping force requirements while ensuring grasping stability. It avoids object damage or grasping failure caused by force control lag or inappropriate force values ​​in traditional methods, providing a reliable technical guarantee for precise robotic grasping operations.

[0137] It should be noted that the functions or steps that the computer-readable storage medium or robot can achieve are described in the relevant descriptions of the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0140] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A robot grasping control method, characterized in that, The method includes: Acquire point cloud data of the target object, and generate multiple candidate grasping postures based on the point cloud data; Based on the physical consistency index, the multiple candidate grasping postures are reordered to select the optimal grasping posture; the physical consistency index is calculated based on the point cloud data of the contact area between the gripper and the target object corresponding to each candidate posture. Using the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, the desired tactile image under the optimal force-stable grasping condition is predicted through a conditional diffusion model. Based on the real-time acquired tactile images, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator, and the gripper is controlled to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image.

2. The robot grasping control method according to claim 1, characterized in that, The generation of multiple candidate grasping postures based on the point cloud data includes: After processing the point cloud data using the GraspNet model, multiple candidate grasping poses and the confidence scores of each candidate grasping pose are output.

3. The robot grasping control method according to claim 1, characterized in that, The step of reordering the multiple candidate grasping postures based on physical consistency metrics to select the optimal grasping posture includes: For each candidate grasping posture, extract the contact area point cloud data of the contact area between the gripper and the target object; The point cloud data of the contact area is aligned to the coordinate system of the gripper fingertip through a rigid transformation; In the gripper index coordinate system, one or more physical consistency indices of the contact area point cloud data are calculated; the physical consistency indices include at least one of surface roughness, normal consistency and curvature uniformity. The physical consistency index is weighted and fused with the confidence level of the candidate grasping posture to obtain a comprehensive score; The candidate grasping postures are sorted in ascending or descending order based on the comprehensive score, and the one with the best score is selected as the optimal grasping posture.

4. The robot grasping control method according to claim 3, characterized in that, The surface roughness is quantified by calculating the normal variance or local curvature variation of the contact area point cloud data; the normal uniformity is quantified by calculating the variance of the surface normal angle of the contact area point cloud data; and the curvature uniformity is evaluated by calculating the entropy value of the curvature distribution of the contact area point cloud data.

5. The robot grasping control method according to claim 3 or 4, characterized in that, The process involves using the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as input, and predicting the desired tactile image under the optimal force-stable grasping condition through a conditional diffusion model. This includes: The current tactile observation image is encoded into a first latent space vector using a VAE encoder; The VAE encoder is used to encode the depth image of the contact area corresponding to the optimal grasping posture into a second latent space vector. The physical condition parameters of the target object are fused into a condition vector; The weight coefficients of the first latent space vector, the second latent space vector, and the conditional vector are determined based on the attention mechanism. Based on the weight coefficients, the first latent space vector, the second latent space vector, and the conditional vector are fused to obtain a fused feature vector; The fused feature vector is used as input to generate a tactile target vector using a pre-trained conditional diffusion model; The tactile target vector is encoded into a desired tactile image.

6. The robot grasping control method according to claim 1, characterized in that, The step of iteratively calculating the torque value of the gripper in the tactile latent space using a linear quadratic modulator based on the real-time acquired tactile image includes: The real-time acquired tactile image is encoded into a current latent space vector; Calculate the error between the current latent space vector and the target latent space vector; the target latent space vector is obtained by encoding the desired tactile image; Based on the error, the torque value of the gripper is iteratively calculated in the tactile latent space using a linear quadratic regulator; When the preset iteration exit condition is met, the clamping is set to hold mode; wherein, the cost function of the linear quadratic regulator is configured to simultaneously minimize the error and the movement amplitude of the gripper.

7. The robot grasping control method according to claim 6, characterized in that, The iteration exit condition is: Monitor the norm of the error in real time; When the norm of the error is lower than a preset threshold and continues for a preset duration, it is determined that the iteration exit condition is met.

8. A robot grasping control device, characterized in that, The device includes: The acquisition module is used to acquire point cloud data of the target object and generate multiple candidate grasping postures based on the point cloud data. The filtering module is used to reorder the multiple candidate grasping postures based on a physical consistency index in order to select the optimal grasping posture; the physical consistency index is calculated based on the point cloud data of the contact area between the gripper and the target object corresponding to each candidate posture. The prediction module is used to predict the desired tactile image under the optimal force-stable grasping condition by taking the depth image of the contact area corresponding to the optimal grasping posture, the physical condition parameters of the target object, and the current tactile observation image of the target object as inputs, and by using a conditional diffusion model. The control module is used to iteratively calculate the torque value of the gripper in the tactile latent space using a linear quadratic regulator based on the real-time acquired tactile image, and to control the gripper to grasp the target object using the torque value, so that the current tactile image of the gripper approximates the desired tactile image.

9. A robot, characterized in that, The robot includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the robot grasping control method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the robot grasping control method as described in any one of claims 1 to 7.