A mechanical arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation
By constructing a dynamic semantic geometric feature field and performing real-time dynamic compensation, the optimal grasping point is generated and corrected online, which solves the problem of insufficient grasping of irregular objects in the existing technology and achieves adaptive grasping effect in complex scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU SUPERVISION & INSPECTION INST FOR PROD QUALITY
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-31
AI Technical Summary
Existing dynamic grasping systems struggle to understand the local graspable areas, avoidance areas, and deformation risks of irregular objects, and lack semantic-geometric fusion and real-time dynamic compensation capabilities, resulting in poor grasping performance in complex scenes.
By synchronously collecting multimodal data, a dynamic semantic geometric feature field is constructed, the optimal gripping point is generated and real-time dynamic compensation is performed. The gripping force control parameters are generated by combining semantic category, material estimation and local contact geometry, and online correction is performed by integrating target motion prediction and robotic arm status to achieve force control adaptation and stable gripping.
It significantly improves the success rate of grasping complex-shaped objects, avoids excessive compression of flexible objects and structural damage to fragile objects, reduces positioning deviation and cumulative posture error during dynamic grasping, and ensures the safety and stability of the grasping process.
Smart Images

Figure CN122480966A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot perception and control technology, and in particular to an adaptive grasping system and method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation. Background Technology
[0002] With the development of automated logistics, flexible sorting, and intelligent manufacturing, the dynamic grasping task of robotic arms has expanded from grasping structured workpieces to comprehensive operations in complex environments. Traditional dynamic grasping systems typically rely on 2D image recognition, PnP positioning, and preset trajectories, which are suitable for regular boxes or workpieces with limited posture changes. However, in real industrial scenarios, targets often have characteristics such as irregular shapes, partial occlusion, large material differences, and limited graspable areas. They may also contain flexible packaging or easily deformable items. In such cases, relying solely on RGB vision and geometric matching is insufficient to meet grasping requirements. Semantic perception, geometric understanding, and execution control must be integrated. Furthermore, the actual grasping effect of the robotic arm platform depends not only on the rationality of the grasping point but also on the control system's ability to compensate for the target's motion state and execution errors. High-precision robotic arms, such as the Franka FR3, face problems such as end-effector lag, posture error accumulation, and contact impact if static grasping points or low-frequency correction strategies are used in high-speed dynamic target scenarios. Therefore, dynamic grasping systems are gradually involving the joint optimization of visual perception, target positioning, grasping point generation, motion planning, gripper control, and feedback execution.
[0003] Existing dynamic grasping systems mainly rely on RGB vision and geometric positioning, which makes it difficult to understand the local graspable area, avoidance area, and deformation risk of an object. This results in insufficient adaptability to irregular objects, occluded scenes, and flexible objects. At the same time, fixed gripping force or simple threshold strategies cannot dynamically adjust the gripping force according to the semantic category, material properties, and local contact geometry of the object, which can easily cause slippage or compression deformation. In addition, existing compensation schemes are mostly performed at the vision or trajectory layer, without combining the joint state of the robotic arm, end effector error, and contact feedback for real-time correction. This makes it difficult to guarantee the positioning accuracy and stability during dynamic grasping. Therefore, how to solve the problems of missing semantic-geometric fusion, lack of adaptive gripping force control, and insufficient coupling between dynamic compensation and the platform, and achieve adaptive grasping of the robotic arm for complex dynamic scenes, is the problem that this invention aims to solve. Summary of the Invention
[0004] To overcome the shortcomings of the prior art, the present invention provides an adaptive grasping system and method for robotic arms based on semantic-geometric fusion perception and real-time dynamic compensation, which can effectively solve the problems involved in the prior art.
[0005] The objective of this invention can be achieved through the following technical solution: Firstly, this invention provides an adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, comprising the following steps: Step 1: Simultaneously collect multimodal data including RGB images, depth maps, point clouds, conveyor belt speed, robotic arm joint states, and end-effector poses. Through hand-eye calibration and spatiotemporal alignment, unify the multimodal data to the robotic arm base coordinate system, reconstruct the dynamic target point cloud under the unified coordinate system, ensure spatiotemporal consistency of multi-source data, and provide an accurate target pose reference for dynamic grasping. Step 2: Construct a dynamic semantic geometric feature field, assigning category semantics, local geometry, accessibility, and occlusion / visibility features to each point in the target point cloud set, forming a dense joint representation for grasping decisions, enabling the robotic arm to understand the local graspable area and avoidance area of the object. Step 3: In the dynamic semantic-geometric feature field, establish semantic-geometric graspability scores for different candidate regions of the same target, and generate several candidate grasping points and corresponding end poses on each target. Filter the optimal grasping points and end poses to improve the rationality of grasping point selection and enhance the grasping success rate of complex-shaped objects. Step 4: Generate clamping force control parameters based on the semantic category, material estimation, weight estimation, local contact geometry features, and task type of the target object to match the rigid-flexible heterogeneous grasping requirements, avoid rigid objects from slipping and flexible objects from being squeezed and deformed, and achieve force control adaptation. Step 5: Integrate target motion prediction, FR3 joint status, end-effector pose error and visual feedback to construct a position-attitude-temporal joint compensation term, correct the grasping point and execution trajectory online, reduce the position deviation and attitude accumulation error in dynamic grasping, and improve tracking accuracy. Step 6: When the end approaches the target and a contact event is detected, the gripper displacement change, contact force change and gripping success criteria are integrated to dynamically adjust the gripping force increment and closing speed to prevent slippage or excessive squeezing, achieve a smooth force control transition after contact, and ensure the safety and stability of the gripping process. Step 7: After the grasping is completed, the force control parameters and scoring weights of various objects are updated based on the error data during the grasping process. A transfer learning mechanism is introduced to generalize the correction strategy to similar unlabeled objects, forming a semantic-force control adaptive closed-loop database that can be iteratively optimized across tasks. This enables the system to have cross-task self-optimization capabilities and accelerates the grasping and deployment of unknown objects.
[0006] Preferably, step 1 specifically includes: Synchronously trigger the RGB-D camera, conveyor belt encoder, and Franka FR3 robotic arm's joint status feedback device to collect multimodal data including RGB images, depth maps, point clouds, conveyor belt linear velocity, joint angles, and end-effector six-dimensional pose, ensuring strict consistency of multi-source data on the time base. Based on the pre-calibrated hand-eye transformation matrix and depth camera intrinsic parameters, the RGB image and depth map are back-projected pixel by pixel to the robotic arm base coordinate system. The spatiotemporal alignment of the conveyor belt motion state and the robotic arm state is achieved through timestamp interpolation and Kalman filtering, eliminating coordinate and time deviations and improving dynamic tracking accuracy. Adaptive voxel mesh downsampling, statistical outlier removal based on local density distribution, and region growing segmentation based on normal consistency constraints are sequentially performed on the fused point cloud under a unified coordinate system to reconstruct a dynamic target point cloud set that retains geometric edge features and instantaneous motion vectors, thereby reducing data redundancy and filtering out noise to obtain complete and clear target instances.
[0007] Preferably, step 2 specifically includes: The dynamic target point cloud set is input into a pre-trained semantic segmentation network and a local geometric feature extractor to generate a category semantic vector and a local geometric descriptor containing normal, curvature and edge distance for each point, thereby improving the accuracy of object part recognition and local geometric perception. Based on the gripper contact surface parameters and the local flatness of the point cloud, the accessibility score of each point is calculated. At the same time, combined with the ray visibility analysis under the current view, occlusion / visibility confidence features are generated to enhance the assessment ability of gripping contact feasibility and field of view occlusion. By channel concatenating and spatially normalizing semantic vectors, geometric descriptors, accessibility scores, and occlusion confidence, a dense semantic-geometric joint feature field corresponding to each spatial location is constructed, forming a unified representation for grasping decisions and improving feature robustness.
[0008] Preferably, step 3 specifically includes: Within the same target point cloud set in a unified coordinate system, a neighborhood sphere query is performed along the local normal direction. For each candidate grabbing region, the matching probability of semantic category and gripper contact surface material, the flatness represented by the ratio of eigenvalues of the local point cloud covariance matrix, and the occlusion rate of the ray from the current camera viewpoint to the region being intercepted by other objects are calculated to generate three basic scoring components, ensuring that the grabbing point has semantic matching, surface flatness, and clear field of view. The variance of the angle between the point cloud normal and the approach direction of the gripper within the candidate region is evaluated to quantify the low damage risk. The compressive deformation sensitivity is calculated by local thickness estimation and curvature change rate. At the same time, the dynamic gripping timing adaptation component is generated by combining the time alignment deviation between the conveyor belt speed and the target motion trajectory to reduce the gripping damage risk and improve the dynamic gripping timing matching degree. The five scoring components are non-linearly weighted and fused using an adaptive weight adjustment mechanism based on the historical success rate of crawling. The weight coefficients are updated online based on the backpropagation error value in the previous round of crawling feedback to obtain the comprehensive semantic-geometric crawlable score for each candidate crawling region. The accuracy of the comprehensive score is continuously optimized by utilizing historical feedback.
[0009] Preferably, step 3 further includes: Within the candidate region where the comprehensive score exceeds the dynamic threshold, based on the umbrella-shaped normal projection distribution of the local point cloud and the geometric constraints of the parallel opening of the gripper, a multi-resolution voxel mesh sampling strategy is used to discretize and generate multiple candidate gripping points and their corresponding quaternion end poses and estimated contact widths constrained by the alignment of the proximity axis and normal, ensuring that the gripping points satisfy geometric alignment and pose unambiguity at different resolutions. For flexible bagged products, semantic edge detection is used to extract the point cloud skeleton of the sealing area and prioritize nodes with uniform stress on the skeleton line. For fragile packaging boxes, the side wall label output by semantic segmentation and flatness filtering results are used to screen areas without creases. For irregularly shaped parts, normal change detection and curvature peak positioning are used to actively avoid protruding tips, through holes and semantically labeled functional assembly surfaces, so as to realize semantically guided collision avoidance gripping point screening, reduce the risk of gripping damage and improve the gripping safety of complex shaped targets. A path cost function is introduced, which combines the shortest collision-free trajectory length from the candidate gripping point to the current end effector of the robotic arm, the displacement offset of the dynamic target within the gripping time window, and the contact angle change rate during the gripper closure process into the cost evaluation. The optimal gripping point, end effector posture, and approach path with the highest comprehensive score and the lowest path cost are output, thereby improving the gripping execution efficiency and gripper closure stability in dynamic scenarios.
[0010] Preferably, step 4 specifically includes: Based on the semantic category index of the target object, a pre-constructed semantic-material association mapping table is used to back-calculate the minimum anti-slip clamping force by combining the local normal and friction cone constraints. The weight estimation is then integrated to generate the basic clamping force and initial closing velocity values, ensuring that the clamping force is sufficient to prevent slipping without damaging the object surface. A dynamic stiffness estimation model based on local contact geometry is introduced. Based on the curvature change of the gripping area and the compliance coefficient of the contact surface, the deformation risk adjustment factor is calculated. The clamping force rise slope and peak amplitude are adaptively corrected to avoid local indentation or structural damage to the target due to force control overshoot. A semantic category-driven gripping strategy decision tree is constructed, employing position-force hybrid control for rigid objects, force closed-loop progressive pressing for flexible objects, and touch-sensing segmented approximation gripping for fragile objects, to achieve non-destructive and stable gripping of rigid-flexible heterogeneous objects.
[0011] Preferably, step 5 specifically includes: Based on the position sequence of the dynamic target at continuous time, a Kalman predictor is used to estimate the target point cloud pose at the future grasping time. At the same time, the current joint angle, end-effector actual pose and trajectory tracking error of the FrankaFR3 robotic arm are collected to effectively suppress visual measurement noise and improve the target pose prediction accuracy. A position compensation term is constructed by weighting the visual observation error, target velocity prediction error and robotic arm tracking error, and an attitude compensation quaternion correction term is generated based on the deviation between the end-effector approach direction and the local surface normal, thereby reducing the end-effector positioning deviation and enhancing the reliability of grasping alignment. Based on the estimated timing deviation between the target's movement speed and the gripper's closing delay, timing compensation terms are generated for the gripping start time and the gripper closing time. This forms a joint online correction command for position, attitude, and timing, compensating for dynamic lag and ensuring that the gripping action is synchronized with the target's movement.
[0012] Preferably, step 6 specifically includes: When the end effector approaches the target and the gripper’s built-in six-dimensional force / torque sensor detects that the contact force exceeds the preset dynamic contact threshold, a contact event response is triggered. The gripper’s displacement change rate, contact force rise gradient, and end attitude angle deviation are simultaneously latched to ensure that the state at the moment of contact is traceable and to provide a precise reference for subsequent adaptive control. Based on the contact pressure distribution and displacement-force ratio fed back by the distributed tactile array of the two-sided finger sleeves of the gripper, the target's local compression, creep slippage or rigid contact state is determined. If the displacement continues to increase but the contact force does not increase synchronously, the clamping force increment is increased stepwise. If the contact force exceeds the limit, the closing speed is decelerated exponentially and the end pose is locally fine-tuned, realizing intelligent identification and differentiated response of contact state, effectively preventing slippage or overload damage. A multimodal grasping success criterion matrix is adopted, which integrates the gripper opening and closing stroke, contact force holding time, optical flow displacement detection results of the target relative to the conveyor belt, and visual servo locking signal. The gripping force is finely adjusted in real time by a secondary closed-loop PID until the grasping is stable. The comprehensive multi-source information closed-loop control ensures the reliability of stable grasping and locking.
[0013] Preferably, step 7 specifically includes: After the grasping is completed, the visual observation residual sequence, joint torque error trajectory, contact force-time curve feature points and grasping success / failure binary labels of this grasping process are recorded in a structured manner. The data is stored in the force control parameter library according to the three-level index of semantic category-material type-deformation level, realizing fine-grained structured storage of grasping experience, which is convenient for quick retrieval and reuse according to object attributes. Based on the historical success rate of crawling, a weighted sliding window is used to perform Bayesian inference and maximum a posteriori estimation iterative updates on the semantic-geometric scoring weight coefficients and basic clamping force parameters. This allows the parameters to converge to the optimal value of the task one by one and automatically removes abnormal crawling samples, thereby improving the convergence speed and estimation robustness of the parameters and effectively avoiding the contamination of the model by extreme samples. By introducing a feature space similarity transfer learning mechanism based on Siamese networks, the semantic-force control joint parameter vector of the optimized object is adaptively mapped to an unlabeled object with similar semantic or geometric features. This enables zero-shot cross-task rapid generalization and accelerated parameter library cold start, allowing for zero-shot rapid deployment of unknown objects and significantly reducing the parameter tuning cost in new scenarios.
[0014] Secondly, the present invention also provides an adaptive grasping system for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, for implementing the aforementioned adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, comprising the following modules: The multimodal perception and acquisition module is used to simultaneously acquire RGB images, depth maps, point cloud data of the target scene, as well as conveyor belt speed, target motion information, and robotic arm status information, ensuring the spatiotemporal consistency of multi-source data and providing complete input for dynamic grasping; The coordinate calibration and scene reconstruction module is used to complete the hand-eye calibration between the camera and the FR3 robotic arm, coordinate system transformation and 3D reconstruction of the target point cloud, establish a unified spatial reference, and achieve accurate positioning and reconstruction of the target point cloud; The Dynamic Semantic Geometric Feature Field (DSGF) construction module is used to map RGB-D and point cloud data into a dense 3D feature field containing category semantics, local geometry and operable attributes, and to integrate semantic and geometric information to improve the intelligence level of grasping decision-making. The candidate grab point generation and filtering module is used to generate candidate grab points based on information such as local normal, edge distribution, contact accessibility, occlusion degree and flexibility risk in the semantic feature field, and output the optimal grab point and end pose. It intelligently filters the optimal grab point to improve the grab success rate and security. The semantic-driven clamping force control module is used to adaptively calculate the clamping force, closing speed and holding strategy of the gripper based on the semantic category, material estimation, rigidity and flexibility characteristics, local surface contact features and task requirements of the target object. It also adaptively adjusts the force control parameters according to the object properties to avoid slippage or damage. The FR3 real-time dynamic compensation module is used to correct the gripping point position, gripping posture and execution time online based on dynamic target motion prediction, FR3 current joint status, end effector error, visual feedback error and contact feedback error, thereby correcting gripping errors online and ensuring dynamic gripping accuracy and stability. The execution and closed-loop feedback module is used to control the FR3 robotic arm and gripper to perform grasping actions, and continuously receive visual, force / torque or gripper status feedback to make closed-loop adjustments to the grasping process, realize closed-loop control of the entire grasping process, and ensure reliable execution and self-adaptation.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. This is a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation. By constructing a dense feature field that integrates semantic category, local geometry, accessibility and occlusion visibility, the robotic arm can understand the graspable area and avoidance area of the object, effectively cope with complex working conditions such as irregular shape, partial occlusion and flexible deformation, and significantly improve the success rate of dynamic grasping.
[0016] 2. This is a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation. It generates clamping force parameters based on semantic category and local contact geometry, and dynamically limits the slope and peak value of the clamping force by combining deformation risk adjustment mechanism. At the same time, it adopts differentiated clamping strategies for different object types to effectively avoid excessive compression of flexible objects and structural damage to vulnerable objects.
[0017] 3. This is a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation. It integrates target motion prediction, robotic arm joint state and multi-source error feedback, constructs position-attitude-temporal joint compensation terms and corrects the grasping trajectory online, enabling the robotic arm to track the moving target in real time and significantly reducing the positioning deviation and attitude accumulation error in the dynamic grasping process.
[0018] 4. This is a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation. Through contact event-triggered state discrimination and adaptive adjustment mechanism, combined with a multimodal grasping success criterion matrix, the clamping force is finely adjusted in a second closed loop to ensure a smooth transition from free space motion to contact interaction control, effectively preventing slippage and contact instability, and improving the safety of the grasping process. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the workflow of an adaptive grasping system and method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to the present invention. Figure 2 This is a flowchart illustrating an adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to the present invention. Figure 3 This is a schematic diagram of the semantic-driven grasping point scoring of a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation according to the present invention. Figure 4 This is a schematic diagram of the real-time dynamic compensation for FR3 in a robotic arm adaptive grasping system and method based on semantic-geometric fusion perception and real-time dynamic compensation according to the present invention. Figure 5 This is a schematic diagram illustrating the application of a complex target grasping system and method for robotic arms based on semantic-geometric fusion perception and real-time dynamic compensation, according to the present invention. Detailed Implementation
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0021] Example 1, please refer to Figures 1 to 5 This invention provides a technical solution: an adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, comprising the following steps: Step 1: Simultaneously acquire multimodal data including RGB images, depth maps, point clouds, conveyor belt speed, robotic arm joint states, and end-effector pose. Through hand-eye calibration and spatiotemporal alignment, unify the multimodal data to the robotic arm's base coordinate system, reconstructing a dynamic target point cloud in the unified coordinate system. This ensures spatiotemporal consistency of the multi-source data, providing a precise target pose reference for dynamic grasping. Simultaneously trigger the RGB-D camera, conveyor belt encoder, and Franka... The FR3 robotic arm's joint state feedback system collects multimodal data including RGB images, depth maps, point clouds, conveyor belt linear velocity, joint angles, and end-effector six-dimensional pose. This ensures strict consistency of the multi-source data in terms of time reference. Based on the pre-calibrated hand-eye transformation matrix and depth camera intrinsic parameters, the RGB images and depth maps are back-projected pixel by pixel to the robotic arm's base coordinate system. Time stamp interpolation and Kalman filtering are used to achieve spatiotemporal alignment between the conveyor belt motion state and the robotic arm state, eliminating coordinate and time deviations and improving dynamic tracking accuracy. Adaptive voxel mesh downsampling, statistical outlier removal based on local density distribution, and region growing segmentation based on normal consistency constraints are sequentially performed on the fused point cloud under the unified coordinate system to reconstruct a dynamic target point cloud set that retains geometric edge features and instantaneous motion vectors. This reduces data redundancy and filters out noise, resulting in a complete and clear target instance. It should be noted that the RGB-D depth camera, conveyor encoder, and Franka FR3 robotic arm's joint status feedback are synchronously enabled via hardware trigger signal lines, with a uniform sampling frequency of 30 Hz set to ensure the consistency of the time base of multimodal data. The RGB-D camera acquires RGB images and depth maps of the target scene at a resolution of 640×480. Simultaneously, the encoder obtains the real-time linear velocity of the conveyor belt, and the joint feedback reads the angular displacement of the seven joints of the robotic arm and the six-dimensional pose vector of the end effector in the base coordinate system. After acquisition, using a pre-calibrated hand-eye transformation matrix and camera intrinsic parameter matrix, the three-dimensional space point corresponding to each pixel is transformed from the camera coordinate system to the robotic arm's base coordinate system. To address the inconsistency in timestamps caused by transmission delays or jitter, cubic spline interpolation is used to align the conveyor belt speed sequence and joint state sequence in time. The interpolated state variables are then input into a Kalman filter framework to filter out high-frequency measurement noise and predict the optimal motion state estimate for the current moment. After coordinate transformation and spatiotemporal alignment, a fused point cloud is obtained with the robotic arm's base coordinate system as the global reference. This point cloud simultaneously contains spatial information of both the static background and the dynamic target. First, adaptive voxel mesh downsampling is performed on the fused point cloud, setting the voxel edge length to be inversely proportional to the local density of the point cloud; higher density results in smaller voxel edge lengths. To preserve the geometric details of the target surface while reducing data redundancy, a statistical outlier removal method based on local density distribution is then employed. For each point, the number of neighboring points within a 0.02-meter radius is searched, and the mean and standard deviation of the local point spacing are calculated. Points deviating from the mean by more than three times the standard deviation are identified as outliers and removed. Based on this, region growth segmentation is implemented using normal consistency constraints. First, principal component analysis is used to estimate the normal vector of each point. A normal angle threshold of 15 degrees is set as the growth criterion. Growth proceeds from the seed point along the surface, segmenting points with continuous, smooth, and spatially connected normal vectors. Cloud clustering is used to identify independent target instances, ultimately outputting a dynamic target point cloud set that preserves the original geometric edge features and instantaneous motion vectors. In the segmented dynamic target point cloud set, each target instance is accompanied by its three-dimensional spatial coordinates in the base coordinate system, local surface normal distribution, and instantaneous velocity vector calculated through inter-frame difference. For each target point cloud, its bounding box size is further calculated to fit the maximum opening width of the gripper. Each point in the point cloud retains RGB color information and depth confidence, where the depth confidence is calculated based on the distance between the pixel and the optical center of the depth camera and the angle of incidence. For every 0.The confidence level decreases linearly over 1 meter, and is reset to zero when the incident angle exceeds 60 degrees. Furthermore, by matching the nearest neighbor points between two consecutive point clouds, the instantaneous displacement vector of each spatial point is estimated, thereby obtaining the overall velocity and direction of the target. This motion information is cross-validated with conveyor belt encoder readings to correct velocity deviations caused by conveyor belt slippage or target sliding. After processing, a structured dynamic target point cloud set containing geometric coordinates, normals, color, confidence level, and motion vectors is output. Step 2: Construct a dynamic semantic-geometric feature field. Assign category semantics, local geometry, accessibility, and occlusion / visibility features to each point in the target point cloud set, forming a dense joint representation for grasping decisions. This enables the robotic arm to understand the local graspable and avoidable areas of an object. Input the dynamic target point cloud set into a pre-trained semantic segmentation network and a local geometric feature extractor to generate a category semantic vector and a local geometric descriptor containing normal, curvature, and edge distance for each point, improving the accuracy of object component recognition and local geometric perception. Calculate the accessibility score for each point based on the gripper contact surface parameters and the local flatness of the point cloud. Simultaneously, combine the ray visibility analysis under the current viewpoint to generate occlusion / visibility confidence features, enhancing the assessment ability of grasping contact feasibility and field of view occlusion. Concatenate and spatially normalize the semantic vector, geometric descriptor, accessibility score, and occlusion confidence to construct a dense semantic-geometric joint feature field corresponding to each spatial location, forming a unified grasping decision representation and improving feature robustness. It should be noted that the reconstructed dynamic target point cloud set is input as a whole into a pre-trained semantic segmentation network. This network adopts the U-Net architecture and performs parameter optimization based on a large-scale crawling scene dataset. It outputs a C-dimensional category semantic vector for each point, representing the probability distribution of its belonging to different object parts or crawling-related categories. At the same time, the same point cloud set is input into a local geometric feature extractor. This extractor fits a local plane in a neighborhood with a radius of 0.015 meters using principal component analysis, calculates the normal vector of each point, and uses the eigenvalue combination of the covariance matrix of the neighborhood point cloud. Based on the relationship that the eigenvalues λ1≥λ2≥λ3, the curvature is defined as λ3 / (λ1... +λ2+λ3); The edge distance is estimated by the standard deviation of the projection distribution of the point to the neighboring point cloud in the normal direction. Through parallel processing, each spatial point obtains a set of original feature vectors containing category semantics, normal, curvature, and edge distance. Based on the semantic and geometric feature extraction, the fitness of each point as a gripper contact point is calculated. According to the gripper contact surface parameters, including gripper finger width, contact surface curvature, and surface friction coefficient, combined with the local flatness of the point cloud, i.e., the root mean square error of the neighboring point cloud to the fitted plane, the accessibility score is calculated. This score is negatively correlated with the local flatness. When the root mean square error exceeds 0.002 meters, the score approaches zero. At the same time, from Starting from the current camera optical center position, ray visibility analysis is performed on each point: sampling is conducted along the direction from the point to the optical center at 0.001-meter intervals, detecting whether the sampled point is intercepted by other spatial points, and the proportion of uninterrupted rays is used as the visibility confidence score; if a point is intercepted by more than 30% of the ray length, it is determined to be in an occlusion area, and the corresponding occlusion feature value is set to low confidence. The accessibility score and visibility confidence score are combined to obtain the joint index of contact suitability for each point; after completing the independent calculation of features in each dimension, the category semantic vector (C-dimensional), local geometric descriptor (including the three-dimensional components of the normal, curvature scalar, and edge distance scalar, a total of 5 dimensions), and accessibility score of each point are combined. The 1D feature and occlusion confidence feature are concatenated along the channel dimension to form an original joint feature vector of dimension C+7. Then, the vector is spatially normalized: in a spherical neighborhood with a radius of 0.02 meters in the point cloud space, the local mean and standard deviation are calculated for each feature channel. Z-score normalization is performed to eliminate the scale difference of feature distribution between different regions. The normalized joint feature field is densely distributed in space. The feature vector corresponding to each spatial location uniformly encodes the semantic category attribute, local geometry, gripper contact feasibility, and visibility status under the current viewpoint of that point, which constitutes the decision basis for subsequent grab point scoring and filtering. Step 3: In the dynamic semantic-geometric feature field, establish semantic-geometric graspability scores for different candidate regions of the same target, and generate several candidate grasping points and corresponding end poses for each target. Optimal grasping points and end poses are selected to improve the rationality of grasping point selection and enhance the success rate of grasping complex-shaped objects. Within the same target point cloud set in a unified coordinate system, a neighborhood sphere query is performed along the local normal direction. For each candidate grasping region, the adaptation probability of semantic category and gripper contact surface material, the flatness represented by the eigenvalue ratio of the local point cloud covariance matrix, and the occlusion rate of the ray from the current camera viewpoint to the region being intercepted by other objects are calculated to generate three basic scoring components, ensuring that the grasping points possess both semantic and geometric graspability. Semantic matching, surface flatness, and clear field of view are used to evaluate the variance of the angle between the point cloud normal and the gripper approach direction in the candidate region to quantify low damage risk. The pressure deformation sensitivity is calculated by local thickness estimation and curvature change rate. At the same time, the dynamic grasping timing adaptability component is generated by combining the time alignment deviation between the conveyor belt speed and the target motion trajectory to reduce the gripping damage risk and improve the dynamic grasping timing matching degree. The five scoring components are nonlinearly weighted and fused through an adaptive weight adjustment mechanism based on the grasping history success rate. The weight coefficient is updated online according to the error backpropagation value in the previous grasping feedback to obtain the comprehensive semantic-geometric graspability score of each candidate grasping region. The accuracy of the comprehensive score is continuously optimized by using historical feedback. It should be noted that three basic scoring components are calculated for each candidate grasping region. The semantic fit probability is obtained by performing a dot product operation between the semantic category vector of the candidate region and a predefined gripper contact surface material matching matrix. This matrix is constructed based on prior knowledge of the friction coefficient of different materials and grasping stability. The flatness index is quantified using the eigenvalue ratio λ2 / λ1 of the local point cloud covariance matrix. The closer the ratio is to 1, the flatter the surface is, and the more suitable it is for parallel gripper contact. The occupancy rate is calculated by emitting rays from the camera optical center to the candidate region and counting the proportion of rays intercepted by other point clouds. A dense sampling strategy with a step interval of 0.001 meters is adopted. Regions with an occupancy rate exceeding 35% will have their scoring weight automatically reduced. The three components characterize the basic grasping suitability of the candidate region from the perspectives of semantic fit, geometric regularity, and field of view integrity, respectively. Based on the basic scores, the low damage risk and dynamic adaptation capability of the candidate region are further evaluated. The variance of the normal angle is quantified by calculating its standard deviation by statistically analyzing the angle deviation between the normal of each point in the candidate region and the preset gripper approach direction. The smaller the variance, the flatter and more uniform the contact surface is. The lower the risk of local stress concentration when the gripper closes, the more uniform the gripper is. The compressive deformation sensitivity is calculated based on the local thickness estimation and the curvature change rate. The thickness is obtained by searching the candidate region bidirectionally along the normal direction to the point cloud boundary. The curvature change rate is represented by the gradient magnitude of the combination of neighborhood feature values. The two are weighted and fused to form the deformation sensitivity index. The dynamic grasping temporal adaptability component is obtained by comparing the time alignment deviation between the real-time speed of the conveyor belt and the target motion trajectory. When the deviation exceeds 0.05 seconds, the weight of this component will be linearly reduced to ensure that the grasping region with high temporal matching degree can still be selected first in high-speed motion scenarios. The five scoring components are nonlinearly weighted and fused through an adaptive weight adjustment mechanism based on the grasping history success rate to obtain the comprehensive semantic-geometric graspability score of each candidate grasping region. The weight coefficients are updated according to the error backpropagation value in the previous grasping feedback. The error signal comes from the judgment result of grasping success or failure and the deviation record of each execution link. The gradient descent method is used to perform small step-size iterative correction near the preset weight base value. The nonlinear weighted fusion uses the Softmax normalized weight vector to perform product aggregation on the five components. The expression for the comprehensive semantic-geometric crawlable score is as follows: ; in, Candidate crawl regions or candidate crawl points, Indicates semantic fit. Indicates local geometric stability. Indicates current visibility and reachability. This indicates low risk of damage and deformation. This indicates adaptability to dynamic data capture timing. to As a weighting coefficient, this scoring mechanism enables the system to avoid capturing flexible wrinkled areas, edge areas with high slippage risk, or severely occluded areas. Furthermore, step 3 also includes: within candidate regions where the comprehensive score exceeds the dynamic threshold, based on the umbrella-shaped normal projection distribution of the local point cloud and the geometric constraints of the parallel opening of the gripper, a multi-resolution voxel mesh sampling strategy is used to discretize and generate multiple candidate gripping points and their corresponding quaternion end poses and estimated contact widths constrained by near-axis and normal alignment, ensuring that the gripping points satisfy geometric alignment and pose unambiguity at different resolutions. For flexible bagged products, semantic edge detection is used to extract the point cloud skeleton of the sealing area and prioritize nodes with uniform stress on the skeleton line. For fragile packaging boxes, the side wall label output by semantic segmentation and the flatness filtering results are used to screen for crease-free areas. For irregularly shaped parts, normal change detection and curvature peak positioning are used to actively avoid protruding tips, through holes and semantically labeled functional assembly surfaces, realizing semantically guided collision avoidance gripping point selection, reducing the risk of gripping damage and improving the gripping safety of complex-shaped targets. A path cost function is introduced, which combines the shortest collision-free trajectory length from the candidate gripping point to the current end of the robot arm, the displacement offset of the dynamic target within the gripping time window and the contact angle change rate during the gripper closing process into the cost evaluation, and outputs the optimal gripping point, end pose and approach path with the highest comprehensive score and the lowest path cost, thereby improving the gripping execution efficiency and gripper closing stability in dynamic scenarios. It should be noted that coarse sampling is performed within the candidate region with an initial voxel edge length of 0.005 meters to identify the main peak direction where the normal distribution is concentrated. Subsequently, fine sampling is performed within the angular neighborhood centered on the main peak direction, with the voxel edge length reduced to 0.002 meters. Each sampling point corresponds to a candidate gripping point position. For each candidate gripping point, the unambiguous attitude of the end effector is solved by aligning the approach axis direction with the local weighted normal direction and using quaternion interpolation to ensure that the gripper opening plane is parallel to the local tangent plane. At the same time, based on the consistency between the point cloud distribution boundary and the normal direction within the candidate region, the interaction between the gripper and the target surface is predicted. The contact width; for target objects of different semantic categories, semantically guided collision avoidance-type grasping point screening is performed. For flexible bagged goods, the point cloud skeleton of the sealing area is extracted by semantic edge detection operator, and nodes with curvature gradient less than 0.02 are searched on the skeleton line as candidate points with uniform force. Such nodes are retained first while filtering out the wrinkled area in the middle of the bag body. For fragile packaging boxes, based on the side wall category label output by semantic segmentation, combined with flatness filtering, point cloud areas with root mean square error less than 0.0015 meters are retained, and sampling points with normal abrupt change of more than 12 degrees on both sides of the crease line are removed. For irregularly shaped parts, the contact width is used. Normal mutation detection is used to identify the location of curvature peaks, with a curvature threshold of 0.15 set to locate the protrusion tip and the boundary of the through hole. Simultaneously, semantically annotated functional assembly surface masks are integrated to prevent the generation of any candidate gripping points within the masked area, ensuring that the gripping process does not damage critical functional parts of the workpiece. Furthermore, a path cost function is introduced to comprehensively select the best candidate gripping points. The path cost function is composed of three weighted indicators: the first is the shortest collision-free trajectory length from the candidate gripping point to the current end effector of the robotic arm. A fast expanding random tree algorithm is used to search for collision-free paths in the robotic arm configuration space, and the path length exceeds... Candidate points with a distance of 0.3 meters will be penalized by a penalty coefficient; the second term is the displacement offset of the dynamic target within the grasping time window, which is calculated based on the product of the target's instantaneous velocity vector and the estimated duration of the grasping action. When the offset exceeds 0.01 meters, the cost increases linearly; the third term is the contact angle change rate during the gripper closing process, which is the rate of change of the angle between the contact normal and the gripper approach axis during the closing stroke. When this rate exceeds 8 degrees per second, it is judged as having a slip risk and the cost is increased accordingly. Finally, the system outputs the optimal grasping point, end effector attitude, and approach path with the highest comprehensive score and the lowest path cost. Step 4: Based on the semantic category, material estimation, weight estimation, local contact geometry features, and task type of the target object, generate clamping force control parameters to match the rigid-flexible heterogeneous grasping requirements, avoid the slippage of rigid objects and the compression deformation of flexible objects, and achieve adaptive force control. Based on the semantic category index of the target object, a pre-constructed semantic-material association mapping table is indexed. The minimum anti-slip clamping force is back-calculated by combining local normal and friction cone constraints. The weight estimation is integrated to generate the initial values of basic clamping force and closing velocity to ensure that the clamping force is sufficient to prevent slippage and does not damage the object surface. A dynamic stiffness estimation model based on local contact geometry is introduced. Based on the curvature change of the grasping area and the compliance coefficient of the contact surface, the deformation risk adjustment factor is calculated to adaptively correct the clamping force rise slope and peak amplitude to avoid local indentation or structural damage to the target due to force control overshoot. A semantic category-driven clamping strategy decision tree is constructed. For rigid objects, position-force hybrid control is adopted; for flexible objects, force closed-loop progressive pressing is adopted; and for vulnerable objects, touch perception segmented approximation clamping is adopted to achieve non-destructive and stable grasping of rigid-flexible heterogeneous objects. It should be noted that the reference friction coefficient and recommended clamping force range for the corresponding category of the target object are queried. Based on this, and combined with the angle relationship between the local normal of the candidate gripping area and the contact surface of the gripper, the minimum normal clamping force to meet the anti-slip condition is calculated using the Coulomb friction cone model. In the specific calculation process, the gravity estimate is projected along the contact normal and divided by the product of the friction coefficient and the contact surface utilization coefficient to obtain the minimum clamping force threshold required to prevent slippage caused by the tangential component of gravity. At the same time, the weight estimate is multiplied by a preset safety factor and then incorporated into the baseline clamping force. The superposition of terms ultimately generates the basic clamping force command and the corresponding initial closing velocity of the gripper. The initial closing velocity is positively correlated with the magnitude of the basic clamping force to ensure that the gripper can approach the target surface at a reasonable speed without impact during the gripping initiation phase. The rate of curvature change is obtained by spatially differencing the gradient magnitude of the combination of neighborhood feature values. The compliance coefficient is read from a preset material property table according to the semantic category. The compliance coefficient of flexible objects is significantly higher than that of rigid objects. The deformation risk adjustment factor is mapped to the 0-1 region by a weighted sum of the rate of curvature change and the compliance coefficient via an S-shaped function. It was found that this factor directly affects the clamping force rise slope and peak limit parameter. When the deformation risk adjustment factor is close to 1, the clamping force rise slope is limited to below 0.5 N per millisecond, and the peak limit is reduced to 60% of the base clamping force to avoid local indentation or structural damage to the target due to force control overshoot. Based on the semantic category of the object, the grasping task is divided into three branches: rigid objects, flexible objects, and vulnerable objects. For the rigid object branch, a position-force hybrid control strategy is adopted, with position control as the main method to complete rapid approach. After contact detection, it switches to force control holding mode. For flexible object branches, a force-closed-loop progressive pressing strategy is adopted. During the clamping process, the force sensor feedback value is read in real time, and the clamping force is gradually increased in a step-by-step manner. After each incremental step, a stable phase of a preset time is maintained to allow the internal stress of the target to redistribute. For vulnerable object branches, a touch-sensing segmented approximation clamping strategy is adopted. A fast empty stroke closure is performed with a low torque threshold. After the initial contact is detected, the closure is paused and the current clamping opening width is recorded. Then, the micro-stepping mode is switched to gradually increase the clamping force by reducing the closing speed by 50% until the target holding force is reached. The clamping force control parameters can be in the following form: ; in, The basic clamping force is related to the category, material, and task. For friction estimation, To capture regional stability indicators, As a deformation risk indicator, Represents the amplitude limiting function. and These are the minimum and maximum clamping forces, respectively. The weighting coefficients for the comprehensive semantic-geometric crawlable score. These are the weighting coefficients for the friction estimator. To capture the weighting coefficients of regional stability indicators, The weighting coefficients for the deformation risk index; Step 5: Integrate target motion prediction, FR3 joint status, end-effector pose error and visual feedback to construct a position-attitude-temporal joint compensation term, correct the grasping point and execution trajectory online, reduce the position deviation and attitude accumulation error in dynamic grasping, and improve tracking accuracy. Step 6: When the end approaches the target and a contact event is detected, the gripper displacement change, contact force change and gripping success criteria are integrated to dynamically adjust the gripping force increment and closing speed to prevent slippage or excessive squeezing, achieve a smooth force control transition after contact, and ensure the safety and stability of the gripping process. Step 7: After the grasping is completed, the force control parameters and scoring weights of various objects are updated based on the error data during the grasping process. A transfer learning mechanism is introduced to generalize the correction strategy to similar unlabeled objects, forming a semantic-force control adaptive closed-loop database that can be iteratively optimized across tasks. This enables the system to have cross-task self-optimization capabilities and accelerates the grasping and deployment of unknown objects.
[0022] Example 2, as Figures 1 to 5 As shown, based on Embodiment 1, the present invention provides a technical solution: Step 5 specifically includes: using a Kalman predictor to estimate the target point cloud pose at the future grasping moment based on the position sequence of the dynamic target at continuous moments, while simultaneously collecting the current joint angle, end-effector actual pose, and trajectory tracking error of the FrankaFR3 robotic arm, effectively suppressing visual measurement noise and improving the target pose prediction accuracy; constructing a position compensation term that is a weighted combination of visual observation error, target velocity prediction error, and robotic arm tracking error; and generating an attitude compensation quaternion correction term based on the end-effector approach direction and local surface normal deviation to reduce end-effector positioning deviation and enhance grasping alignment reliability; estimating the timing deviation based on the target movement speed and gripper closing delay, generating a timing compensation term for the grasping start moment and gripper closing moment, forming a position-attitude-timing joint online correction command to compensate for dynamic lag and ensure that the grasping action is synchronized with the target movement; It should be noted that during the real-time dynamic compensation process, based on the continuous position sequence of the dynamic target, a Kalman predictor is used to recursively estimate the target point cloud pose at future grasping moments. The state vector of the Kalman filter includes the target's three-dimensional spatial coordinates, instantaneous velocity, and acceleration components in the robot arm's base coordinate system. The observation update frequency is synchronized with the visual sampling frequency at 30 Hz. Through iterative updates of the state transition matrix and covariance matrix, the filter suppresses high-frequency noise in visual measurements and outputs the optimal pose prediction value of the target within the next 50 to 150 millisecond time window. Simultaneously, Franka data is acquired. The current joint angles of the FR3 robotic arm, the actual pose of the end effector in the base coordinate system, and the trajectory tracking error calculated from forward kinematics are all uploaded to the main control unit via a real-time Ethernet bus at a rate of 1 kHz. Based on the target's predicted pose and the current state of the robotic arm, a correction command combining position compensation, attitude compensation, and timing compensation is constructed. The position compensation is a weighted combination of visual observation error, target velocity prediction error, and robotic arm tracking error. The visual observation error originates from the pixel reprojection deviation between the target centroid and the predicted position in the current frame, converted to Euclidean distance in three-dimensional space by the hand-eye calibration matrix. The target velocity prediction error is the component of the Kalman filter's innovation vector in the position channel. The robotic arm tracking error is the deviation between the desired trajectory point and the actual end effector position. The three errors are multiplied by preset weighting coefficients and then summed to form the position correction vector. The attitude compensation generates a quaternion correction term based on the deviation between the end effector's approach direction and the local surface normal. The end effector's attitude is smoothly adjusted through spherical linear interpolation to make the gripper opening level. The surface is kept parallel to the local tangent plane of the target. The timing compensation term estimates the timing deviation based on the product of the target's motion speed and the gripper closing delay. The gripper closing delay includes three parts: command transmission delay, actuator response delay, and mechanical motion lag. After the joint correction command is generated, the position compensation term, attitude compensation term, and timing compensation term are superimposed on the original grasping trajectory planning result to achieve online correction. The position compensation term is injected into the inverse kinematics solver of the robotic arm in the form of the desired pose increment, driving the joint angles to adjust in real time. The attitude compensation term updates the target pose of the end effector through quaternion multiplication to ensure that the gripper approach path always remains consistent with the local surface normal. The timing compensation term corrects the time alignment between the grasping start time and the gripper closing time. Specifically, it triggers the gripper closing command in advance within the complete control cycle before the target's predicted pose arrives at the grasping point, or dynamically delays the grasping start time when the target speed is too high to compensate for the spatial misalignment caused by the delay. The three compensation terms are updated online at a control frequency of 1 kHz to form a closed-loop correction link to ensure Franka The FR3 robotic arm can track the target's movement in real time during dynamic grasping; Estimate the target's capture point position in future time steps based on the target's position, velocity, and acceleration at consecutive time points. At the same time, combined with the current joint status of FR3 Actual end-effector pose Tracking error and visual observation error Construct compensation terms , and The grab point compensation model can be expressed as: ; in, For location compensation items, For attitude compensation, For timing compensation terms; Step 6 specifically includes: when the end effector approaches the target and the gripper's built-in six-dimensional force / torque sensor detects that the contact force exceeds the preset dynamic contact threshold, a contact event response is triggered. The gripper displacement change rate, contact force rise gradient, and end attitude angle deviation are simultaneously latched to ensure that the contact state is traceable and to provide a precise benchmark for subsequent adaptive control. Based on the contact pressure distribution and displacement-force ratio fed back by the distributed tactile array of the gripper's two-sided finger sleeves, the target's local compression, creep slip, or rigid contact state is determined. If the displacement continues to increase but the contact force does not increase synchronously, the gripping force increment is increased stepwise. If the contact force exceeds the limit, the closing speed is decelerated exponentially and the end pose is locally fine-tuned to achieve intelligent identification and differentiated response of the contact state, effectively preventing slippage or overload damage. A multimodal grasping success criterion matrix is adopted, which integrates the gripper opening and closing stroke, contact force holding time, optical flow displacement detection results of the target relative to the conveyor belt, and visual servo locking signal to perform real-time secondary closed-loop PID fine-tuning of the gripping force until the grasping is stable. The comprehensive multi-source information closed-loop control ensures the reliability of stable grasping and locking. It should be noted that when the end effector approaches the dynamic target under real-time dynamic compensation guidance, and the gripper's built-in six-dimensional force / torque sensor detects that the contact force exceeds the preset dynamic contact threshold (this threshold is dynamically set according to the target semantic category and deformation risk adjustment factor), a contact event response is immediately triggered. The response program simultaneously latches three key physical quantities: the gripper displacement change rate (obtained through encoder differential), the contact force rise gradient, and the end effector attitude angle deviation (Euler angle deviation relative to the preset approach direction). The latched data serves as the reference feature vector for subsequent contact state judgment. At the same time, the recursion of the current planned trajectory is paused, and the contact response control subroutine is entered to ensure a smooth transition from free space motion to contact interaction control, avoiding overshoot or contact instability caused by response delay. Contact state classification and adaptive adjustment are performed. When a continuous increase in displacement is detected without a synchronous increase in contact force, it is determined that the target has undergone local compression or creep slippage. The clamping force increment is gradually increased in a step-by-step manner, with each step being 0.5N and a 50ms stability window maintained between steps. When the contact force exceeds the limit, it is determined to be rigid contact or impending damage. An exponential deceleration algorithm is immediately executed to reduce the closing speed, with a deceleration time constant set to 20ms. At the same time, local fine-tuning of the end effector pose is initiated. Incremental adjustment is used to realign the gripper opening plane with the local tangential plane, disperse contact stress, and eliminate stress concentration points. A multimodal grasping success criterion matrix is used to comprehensively evaluate the grasping stability. This matrix integrates four independent criteria: whether the gripper opening and closing stroke converges to the allowable error range of the target opening width; whether the contact force holding time exceeds the preset stability window; whether the optical flow displacement detection result of the target relative to the conveyor belt tends to zero (indicating that the target has been effectively constrained); and whether the visual servo locking signal confirms that there is no relative movement between the end effector and the target. When all four criteria are met, the grasping is determined to be stable. Then, real-time secondary closed-loop PID fine-tuning is performed on the gripping force. A small correction amount (the correction amount does not exceed 10% of the reference value) is applied based on the current gripping force until the standard deviation of the force signal fluctuation is lower than the threshold, completing the stable locking of the dynamic grasping process. Step 7 specifically includes: After the grasping is completed, the visual observation residual sequence, joint torque error trajectory, contact force-time curve feature points and grasping success / failure binary labels of the grasping process are structured and recorded. The data is stored in the force control parameter library according to the three-level index of semantic category-material type-deformation level, realizing fine-grained structured storage of grasping experience, which is convenient for quick retrieval and reuse according to object attributes. Based on the weighted sliding window of historical grasping success rate, Bayesian inference and maximum a posteriori estimation are performed to iteratively update the semantic-geometric scoring weight coefficient and basic clamping force parameters, so that the parameters converge to the optimal value of the task one by one and abnormal grasping samples are automatically removed, improving the parameter convergence speed and estimation robustness, effectively avoiding the pollution of the model by extreme samples. A feature space similarity transfer learning mechanism based on Siamese network is introduced to map the semantic-force control joint parameter vector of the optimized object to the unlabeled but semantically or geometrically similar object through domain adaptive mapping, realizing zero-sample cross-task rapid generalization and parameter library cold start acceleration, realizing zero-sample rapid deployment of unknown objects, and significantly reducing the parameter tuning cost in new scenarios. It should be noted that the key data in this grasping process includes the visual observation residual sequence, joint torque error trajectory, contact force-time curve feature points, and grasping success / failure binary labels. The visual observation residual sequence records the reprojection error of the target centroid on the image plane for each frame. The joint torque error trajectory stores the curves showing the difference between the actual and expected torques of the seven joints during the grasping phase over time. The contact force-time curve feature points extract key statistics such as the peak value of the first derivative during the rising segment of the force signal and the mean and standard deviation during the steady-state holding phase. All data is stored in the force control parameter library according to a three-level index: semantic category, material type, and deformation level. The semantic category is derived from semantic... The output of the segmentation network reads the material type from a pre-defined material property table, and the deformation level is quantified based on a calculated deformation risk adjustment factor. This database uses a structured storage format and supports fast retrieval and batch export by index field. A weighted sliding window based on historical capture success rates iteratively updates the semantic-geometric scoring weight coefficients and basic clamping force parameters using Bayesian inference and maximum a posteriori estimation. The length of the weighted sliding window is set to the most recent 100 capture records. The weight of each capture within the window is determined by its success rate confidence score: a successful capture has a weight of 0.9, and a failed capture has a weight of 0.1, effectively suppressing the contribution of failed samples to parameter updates. To retain negative feedback information, the Bayesian inference process uses the current weight coefficients and clamping force parameters as random variables to be estimated, with their historical distributions as the prior distribution. The likelihood function is calculated using the captured observation data within the window. The optimal estimate of the parameters is obtained by maximizing the posterior probability. The updated parameters converge to the task's optimal value successively. Simultaneously, abnormal captured samples deviating from the mean by more than three standard deviations are automatically removed to prevent extreme data from contaminating the parameter estimation results, ensuring the robustness and convergence stability of the parameter iteration process. The Siamese network receives the semantic-force control joint parameter vector of the optimized object and the semantic and geometric feature vector of the new object through two branches, respectively. The system outputs a distance metric between the two objects in the embedded feature space. When the distance is less than a preset threshold of 0.25, the objects are considered to be similar in features, triggering parameter transfer. During the transfer process, the parameter vector of the optimized object is transformed into the feature space of the new object through a domain adaptive mapping network. This mapping network uses an adversarial training method to eliminate the distribution difference between the source and target domains, and finally generates the initial parameters of the new object. This enables the system to achieve zero-sample rapid deployment for unlabeled targets that are similar to the optimized object in semantic or geometric features, without the need to re-collect training data or perform a complete parameter iteration process. This accelerates the cold start phase of the parameter library and expands the system's adaptability to unknown objects.
[0023] Example 3, as Figures 1 to 5As shown, based on Embodiments 1 and 2, the present invention also provides an adaptive grasping system for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, which is used to implement an adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation, including the following modules: The multimodal perception and acquisition module is used to simultaneously acquire RGB images, depth maps, point cloud data of the target scene, as well as conveyor belt speed, target motion information, and robotic arm status information, ensuring the spatiotemporal consistency of multi-source data and providing complete input for dynamic grasping; The coordinate calibration and scene reconstruction module is used to complete the hand-eye calibration between the camera and the FR3 robotic arm, coordinate system transformation and 3D reconstruction of the target point cloud, establish a unified spatial reference, and achieve accurate positioning and reconstruction of the target point cloud; The Dynamic Semantic Geometric Feature Field (DSGF) construction module is used to map RGB-D and point cloud data into a dense 3D feature field containing category semantics, local geometry and operable attributes, and to integrate semantic and geometric information to improve the intelligence level of grasping decision-making. The candidate grab point generation and filtering module is used to generate candidate grab points based on information such as local normal, edge distribution, contact accessibility, occlusion degree and flexibility risk in the semantic feature field, and output the optimal grab point and end pose. It intelligently filters the optimal grab point to improve the grab success rate and security. The semantic-driven clamping force control module is used to adaptively calculate the clamping force, closing speed and holding strategy of the gripper based on the semantic category, material estimation, rigidity and flexibility characteristics, local surface contact features and task requirements of the target object. It also adaptively adjusts the force control parameters according to the object properties to avoid slippage or damage. The FR3 real-time dynamic compensation module is used to correct the gripping point position, gripping posture and execution time online based on dynamic target motion prediction, FR3 current joint status, end effector error, visual feedback error and contact feedback error, thereby correcting gripping errors online and ensuring dynamic gripping accuracy and stability. The execution and closed-loop feedback module is used to control the FR3 robotic arm and gripper to perform grasping actions, and continuously receive visual, force / torque or gripper status feedback to make closed-loop adjustments to the grasping process, realize closed-loop control of the entire grasping process, and ensure reliable execution and self-adaptation.
[0024] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An adaptive grasping method for robotic arms based on semantic-geometric fusion perception and real-time dynamic compensation, characterized in that, Includes the following steps: Step 1: Simultaneously collect multimodal data including RGB images, depth maps, point clouds, conveyor belt speed, robotic arm joint states, and end-effector poses. Through hand-eye calibration and spatiotemporal alignment, unify the multimodal data to the robotic arm base coordinate system and reconstruct the dynamic target point cloud under the unified coordinate system. Step 2: Construct a dynamic semantic geometric feature field, assigning each point in the target point cloud set with category semantics, local geometry, accessibility, and occlusion / visibility features to form a dense joint representation for grasping decisions; Step 3: In the dynamic semantic-geometric feature field, establish semantic-geometric graspability scores for different candidate regions of the same target, and generate several candidate grasp points and corresponding end poses on each target, and select the optimal grasp point and end pose. Step 4: Generate clamping force control parameters based on the semantic category, material estimation, weight estimation, local contact geometric features, and task type of the target object to match the rigid-flexible heterogeneous grasping requirements. Step 5: Integrate target motion prediction, FR3 joint status, end-effector pose error and visual feedback to construct a position-attitude-temporal joint compensation term and correct the grasp point and execution trajectory online. Step 6: When the end approaches the target and a contact event is detected, the gripper displacement change, contact force change and grasp success criteria are integrated to dynamically adjust the gripping force increment and closing speed. Step 7: After the crawling is completed, the force control parameters and scoring weights of various objects are updated based on the error data during the crawling process. A transfer learning mechanism is introduced to generalize the correction strategy to similar unlabeled objects, forming a semantic-force control adaptive closed-loop database that can be iteratively optimized across tasks.
2. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 1 specifically includes: Synchronously trigger the RGB-D camera, conveyor belt encoder, and Franka FR3 robotic arm's joint status feedback device to collect multimodal data including RGB images, depth maps, point clouds, conveyor belt linear velocity, joint angles, and end-effector six-dimensional pose; Based on the pre-calibrated hand-eye transformation matrix and depth camera intrinsic parameters, the RGB image and depth map are back-projected pixel by pixel to the robotic arm base coordinate system, and the spatiotemporal alignment of the conveyor belt motion state and the robotic arm state is achieved through timestamp interpolation and Kalman filtering. Adaptive voxel mesh downsampling, statistical outlier removal based on local density distribution, and region growing segmentation based on normal consistency constraints are sequentially performed on the fused point cloud under a unified coordinate system to reconstruct a dynamic target point cloud set that preserves geometric edge features and instantaneous motion vectors.
3. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 2 specifically includes: The dynamic target point cloud set is input into a pre-trained semantic segmentation network and a local geometric feature extractor to generate a category semantic vector and a local geometric descriptor containing normal, curvature and edge distance for each point; Based on the contact surface parameters of the gripper and the local flatness of the point cloud, the accessibility score of each point is calculated. At the same time, combined with the ray visibility analysis under the current viewpoint, the occlusion / visibility confidence feature is generated. Semantic vectors, geometric descriptors, accessibility scores, and occlusion confidence are channel-concatenated and spatially normalized to construct a dense semantic-geometric joint feature field corresponding to each spatial location.
4. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 3 specifically includes: Within the same target point cloud set in a unified coordinate system, a neighborhood sphere query is performed along the local normal direction. For each candidate grabbing region, the matching probability of semantic category and gripper contact surface material, the flatness represented by the ratio of eigenvalues of the local point cloud covariance matrix, and the occlusion rate of the ray from the current camera viewpoint to the region being intercepted by other objects are calculated to generate three basic scoring components. The variance of the angle between the point cloud normal and the gripper approach direction within the candidate region is evaluated to quantify the low damage risk. The compressive deformation sensitivity is calculated by local thickness estimation and curvature change rate. At the same time, dynamic gripping timing adaptation components are generated by combining the time alignment deviation between the conveyor belt speed and the target motion trajectory. The five scoring components are non-linearly weighted and fused using an adaptive weight adjustment mechanism based on the historical success rate of crawling. The weight coefficients are updated online based on the backpropagation error value in the previous round of crawling feedback to obtain the comprehensive semantic-geometric crawlability score for each candidate crawling region.
5. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 4, characterized in that: Step 3 also includes: Within the candidate region where the comprehensive score exceeds the dynamic threshold, based on the umbrella-shaped normal projection distribution of the local point cloud and the geometric constraints of the parallel opening of the gripper, a multi-resolution voxel mesh sampling strategy is used to discretize and generate multiple candidate gripping points and corresponding quaternion end poses and estimated contact widths constrained by the alignment of the proximity axis and normal. For flexible bagged products, semantic edge detection is used to extract the point cloud skeleton of the sealing area and prioritize nodes with uniform stress on the skeleton line. For fragile packaging boxes, the side wall label output by semantic segmentation and flatness filtering results are used to screen out areas without creases. For irregular parts, normal change detection and curvature peak positioning are used to actively avoid protruding tips, through holes and semantically labeled functional assembly surfaces. A path cost function is introduced, which combines the shortest collision-free trajectory length from the candidate gripping point to the current end effector of the robotic arm, the displacement of the dynamic target within the gripping time window, and the rate of change of the contact angle during the gripper closing process into the cost evaluation, and outputs the optimal gripping point, end effector posture, and approach path with the highest comprehensive score and the lowest path cost.
6. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 4 specifically includes: Based on the semantic category index of the target object, a pre-constructed semantic-material association mapping table is used to back-calculate the minimum anti-slip clamping force by combining the local normal and friction cone constraints, and the weight estimation is fused to generate the basic clamping force and initial value of the closing velocity. A dynamic stiffness estimation model based on local contact geometry is introduced. The deformation risk adjustment factor is calculated based on the curvature change of the gripping area and the compliance coefficient of the contact surface, and the clamping force rise slope and peak amplitude are adaptively corrected. A semantic category-driven clamping strategy decision tree is constructed, employing position-force hybrid control for rigid objects, force closed-loop progressive pressing for flexible objects, and touch-sensing segmented approximation clamping for fragile objects.
7. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 5 specifically includes: Based on the position sequence of the dynamic target at continuous time, a Kalman predictor is used to estimate the target point cloud pose at the future grasping time, while simultaneously collecting the current joint angle, end-effector actual pose, and trajectory tracking error of the FrankaFR3 robotic arm. A position compensation term is constructed by weighting the visual observation error, target velocity prediction error and robotic arm tracking error, and an attitude compensation quaternion correction term is generated based on the deviation between the end-effector approach direction and the local surface normal. Based on the estimated timing deviation between the target's motion speed and the gripper closing delay, timing compensation terms are generated for the gripping start time and the gripper closing time, forming a joint online correction command for position, attitude, and timing.
8. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 6 specifically includes: When the end effector approaches the target and the gripper’s built-in six-dimensional force / torque sensor detects that the contact force exceeds the preset dynamic contact threshold, a contact event response is triggered, and the gripper displacement change rate, contact force rise gradient and end attitude angle deviation are simultaneously latched. Based on the contact pressure distribution and displacement-force ratio fed back by the distributed tactile array of the two-sided finger sleeves of the gripper, the target's local compression, creep slippage or rigid contact state is determined. If the displacement continues to increase but the contact force does not increase synchronously, the clamping force increment is increased stepwise. If the contact force exceeds the limit, the closing speed is decelerated exponentially and the end pose is locally fine-tuned. A multimodal grasping success criterion matrix is adopted, which integrates the gripper opening and closing stroke, contact force holding time, optical flow displacement detection results of the target relative to the conveyor belt, and visual servo locking signal to perform real-time secondary closed-loop PID fine-tuning of the gripping force until the grasping is stable.
9. The adaptive grasping method for a robotic arm based on semantic-geometric fusion perception and real-time dynamic compensation according to claim 1, characterized in that: Step 7 specifically includes: After the capture is completed, the visual observation residual sequence, joint torque error trajectory, contact force-time curve feature points and capture success / failure binary labels of this capture process are recorded in a structured manner and stored in the force control parameter library according to the three-level index of semantic category-material type-deformation level. Based on the historical success rate of crawling, a weighted sliding window is used to perform Bayesian inference and maximum a posteriori estimation to iteratively update the semantic-geometric scoring weight coefficient and the basic clamping force parameter, so that the parameters converge to the optimal value of the task and abnormal crawling samples are automatically removed. A feature space similarity transfer learning mechanism based on Siamese networks is introduced to adaptively map the semantic-force control joint parameter vector of the optimized object to an unlabeled object that has similar semantic or geometric features.
10. A robotic arm adaptive grasping system based on semantic-geometric fusion perception and real-time dynamic compensation, used to implement the robotic arm adaptive grasping method based on semantic-geometric fusion perception and real-time dynamic compensation as described in any one of claims 1-9, characterized in that, Includes the following modules: The multimodal perception and acquisition module is used to simultaneously acquire RGB images, depth maps, point cloud data of the target scene, as well as conveyor belt speed, target motion information, and robotic arm status information; The coordinate calibration and scene reconstruction module is used to complete the hand-eye calibration, coordinate system transformation and 3D reconstruction of the target point cloud between the camera and the FR3 robotic arm; The dynamic semantic geometric feature field construction module is used to map RGB-D and point cloud data into a three-dimensional dense feature field containing category semantics, local geometry and operable attributes; The candidate grab point generation and filtering module is used to generate candidate grab points based on local normals, edge distribution, contact accessibility, occlusion degree and flexibility risk information in the semantic feature field, and output the optimal grab point and end pose. The semantic-driven clamping force control module is used to adaptively calculate the clamping force, closing speed and holding strategy of the gripper based on the semantic category, material estimation, rigidity and flexibility characteristics, local surface contact features and task requirements of the target object. The FR3 real-time dynamic compensation module is used to correct the gripping point position, gripping posture and execution time online based on dynamic target motion prediction, FR3 current joint status, end effector error, visual feedback error and contact feedback error. The execution and closed-loop feedback module is used to control the FR3 robotic arm and gripper to perform grasping actions and continuously receive visual, force / torque, or gripper status feedback to make closed-loop adjustments to the grasping process.