Robot grabbing confrontation attack method based on physical information
By constructing a joint optimization method for cage structures and control points, the problem in existing technologies that robot grasping attacks rely solely on the perception level is solved, and the grasping stability and lifting ability are weakened at the physical level. It is suitable for a variety of grasping systems and can effectively attack in real scenarios.
Patent Information
- Application Number
- CN202510847943.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the existing technology, adversarial attack methods for robotic grasping mainly focus on the perception level, ignoring the importance of physical interaction, resulting in incomplete attack effects and difficulty in affecting the physical robustness of the grasping system.
By constructing a cage structure and control points, generating a deformation map, establishing a joint optimization objective function, optimizing and adjusting the cage control points, outputting adversarial objects, and weakening the grasping stability and lifting ability.
Without significantly changing the appearance of the object, the physical performance of the grasping configuration is systematically weakened. It is applicable to various grasping systems, has good versatility and transferability, and can reproduce the attack effect in real physical scenarios.
Smart Images

Figure CN120807620A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of robots, and particularly relates to a robot grasping adversarial attack method based on physical information. BACKGROUND
[0002] With the wide application of robot technology in industrial manufacturing, warehouse logistics, home service and other fields, robot grasping, as one of the key capabilities of its interaction with the physical world, has increasingly become the focus of research and application. The existing work on adversarial attacks on robot grasping mainly focuses on attacking neural network models based on images or point clouds, for example, by adding tiny perturbations to mislead the grasping quality evaluation network and make it produce incorrect score judgments.
[0003] However, such methods only act on the perception level and ignore the fact that robot grasping is a behavior that highly depends on physical interaction, and its success or failure depends not only on the prediction results of the network, but also on a series of physical mechanisms such as gravity, friction, contact position and attitude, torque balance, etc. In other words, even if the attack makes the neural network output an incorrect grasping decision, if the grasping process itself has good physical stability, the robot may still successfully complete the grasping action. Therefore, attack strategies that rely only on the perception level are difficult to fully measure and affect the physical robustness of the grasping system. SUMMARY
[0004] In view of the above defects of the prior art, the application provides a robot grasping adversarial attack method based on physical information, which is used to effectively destroy the physical grasping process without significantly changing the overall structure or visual features of the object. The technical solution designed by the application comprises the following steps: S1: obtaining the cage structure of the target adversarial object and generating cage control points; S2: establishing a deformation mapping between the cage control points and the vertices of the target adversarial object; S3: constructing a joint optimization objective function depending on physical indicators; S4: optimizing and adjusting the cage control points based on the deformation mapping and the joint optimization objective function, and then outputting the target adversarial object after optimization and adjustment.
[0005] Preferably, the cage structure of the target adversarial object in S1 comprises: The target adversarial object is determined, and the representation of the target adversarial object in the form of a triangular mesh includes a vertex set and a sheet set. A three-dimensional Cartesian grid volume is constructed based on the bounding box of the target adversarial object. The triangular mesh of the target adversarial object is divided into a plurality of equidistant cells in each dimension to form a regular three-dimensional voxel structure. An external surface of the three-dimensional voxel structure is extracted to form a three-dimensional closed cage-shaped geometry. The surface of the cage-shaped geometry is composed of a plurality of rectangular or triangular sheets, and the cage-shaped structure tightly wraps the target adversarial object.
[0006] Preferably, the generating cage-shaped control points in S1 comprises: Each sheet in the sheet set corresponding to the cage-shaped structure is divided into a uniform grid array according to a preset resolution. All grid intersection points located on the surface of the cage-shaped structure in the division result constitute a cage-shaped control point set.
[0007] Preferably, S2 comprises: For any target adversarial object vertex, the new position under the perturbation of the cage-shaped control point is represented as a weighted linear combination of all cage-shaped control points, and the formula is as follows:
[0008] In the formula, is the new position of the i-th target adversarial object vertex, is the average value coordinate weight of the i-th target adversarial object vertex, is the i-th target adversarial object vertex, is the three-dimensional coordinate of the j-th cage-shaped control point, is the new position of the j-th cage-shaped control point, is the total number of cage-shaped control points, and for each j, , , . .
[0009] Preferably, the average value coordinate weight is as follows:
[0010]
[0011] In the formula, is the included angle between the i-th target adversarial object vertex and its adjacent cage-shaped control point in the perspective view, is the relative distance between the i-th target adversarial object vertex and its adjacent cage-shaped control point, is the relative distance between the j-th cage-shaped control point and the i-th target adversarial object vertex, is the relative distance between the j-th cage-shaped control point and the i-th target adversarial object vertex, is the relative distance between the j-th cage-shaped control point and the i-th target adversarial object vertex. vector.
[0012] Preferably, the joint optimization objective function in S3 is as follows:
[0013] In the formula, is the target adversarial object generated after the cage control point disturbance, is the lifting ability loss of the target adversarial object under the target grasping configuration, is the grasping stability index, is the regularization term based on Laplacian smoothing, and is the balance coefficient.
[0014] Preferably, the regularization term based on Laplacian smoothing is as follows:
[0015] In the formula, is all vertices of the target adversarial object, is the vertex adjacent vertex set.
[0016] Advantages: 1. The application breaks through the limitation of existing methods that only act on the perception layer of the neural network or rely on significant geometric deformation by introducing an adversarial attack strategy based on physical mechanisms. By constructing a controllable cage structure and an efficient control point disturbance mechanism, the application can realize perturbation of the local geometry of the object surface while keeping the appearance of the object basically unchanged, thereby systematically weakening the lifting ability and grasping stability of the object under the target grasping configuration. The application is suitable for various types of grasping systems and does not rely on specific network structures or input formats, thus having good universality and transferability. 2. The application effectively balances the attack strength and shape realizability by jointly optimizing multiple physical indicators and combining a geometric regularization term. The generated adversarial samples not only exhibit significant grasping interference effects in the simulation environment, but also can be reproduced in real physical scenarios through 3D printing and other means, thus having practical engineering application value and being widely applicable to robustness testing, safety evaluation and adversarial training of robot grasping systems, thus having good research promotion potential and commercialization prospects. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flowchart of a preferred embodiment of the application. DETAILED DESCRIPTION
[0018] The embodiments of the present application are described in detail below, and the following embodiments are implemented on the premise of the technical solutions of the present application, and detailed implementation modes and specific operation processes are given, but the protection scope of the present application is not limited to the following embodiments.
[0019] The present application designs a robot grasping anti-attack method based on physical information, and the technical scheme comprises the following steps, as shown in the figure, specifically comprising: Figure 1 S1: obtaining the cage structure of the target adversarial object and generating the cage control points; S2: establishing the deformation mapping of the cage control points and the vertexes of the target adversarial object; S3: constructing a joint optimization objective function depending on physical indicators; S4: optimizing and adjusting the cage control points based on the deformation mapping and the joint optimization objective function, and then outputting the target adversarial object after optimization and adjustment.
[0020] Preferably, the cage structure of the target adversarial object in S1 comprises: determining the target adversarial object, representing the target adversarial object in the form of a triangular mesh, containing a vertex set and a patch set, constructing a three-dimensional Cartesian grid volume based on the bounding box of the target adversarial object, dividing the triangular mesh of the target adversarial object into a plurality of equidistant cells in each dimension to form a regular three-dimensional voxel structure, extracting the external surface of the three-dimensional voxel structure to form a three-dimensional closed cage geometry, the surface of the cage geometry is composed of a plurality of rectangular or triangular patches, and the cage structure tightly wraps the target adversarial object.
[0021] Preferably, the generation of the cage control points in S1 comprises: dividing each patch in the patch set corresponding to the cage structure according to a predetermined resolution to form a uniform grid array, and all grid intersection points located on the surface of the cage structure in the division result constitute a cage control point set.
[0022] Specifically, for S1, a structured control framework, i.e. cage, is constructed around the target adversarial object, which serves as the basis carrier for shape optimization and provides spatial constraints and computational convenience for the subsequent arrangement and disturbance propagation of control points.
[0023] In addition, to ensure the flexibility and computational efficiency of the perturbation, the resolution of the cage structure can be parameterized according to the application requirements. For example, by setting the initial cage unit edge length (such as 0.04 meters) to control the initial control accuracy, the smaller the edge length, the more refined the cage structure, and the higher the controllable deformation ability. In the initial stage, the cage structure maintains a low resolution for global coarse adjustment, and then further refines in the optimization iteration process, gradually increasing the resolution of the cage, thereby gradually improving the local accuracy of the shape change.
[0024] The completed cage structure not only covers the entire geometric boundary of the object, but also avoids significant overlap or penetration with the object, ensuring the physical realizability and topological continuity of the shape change in the subsequent perturbation process. In addition, the geometric representation of the cage structure needs to support efficient spatial interpolation operations, such as using the "mean value coordinates" or "cage-based deformation" method, so that the control point perturbation can be transmitted to the original mesh vertices in a weighted manner, thereby producing natural and continuous geometric deformation.
[0025] In addition, after the cage structure (cage) is initialized, a set of control points (control points) need to be set on its surface, which serve as the operation unit of geometric perturbation and guide the local deformation of the three-dimensional object surface. The setting of control points and their deformation propagation mechanism is a key step in realizing fine and controllable perturbation.
[0026] Preferably, S2 comprises: For any target adversarial object vertex, the new position under the action of cage control point perturbation is represented as a weighted linear combination of all cage control points, as follows:
[0027] In the formula, is the new position of the th target adversarial object vertex, is the average value coordinate weight of , and is the th target adversarial object vertex, is the three-dimensional coordinate of the th cage control point, is the new position of the th cage control point, is the total number of cage control points, and for each , , .
[0028] Preferably, the mean value coordinate weight , is calculated as follows:
[0029]
[0030] wherein, is the angle between the adjacent cage control point and the vector . is the angle between the adjacent cage control point and the vector . is the vector from the cage control point to the object vertex.
[0031] Specifically, for S2, in order to realize the conduction of control point perturbation to the object surface, an interpolation relationship between the control point and the original vertex set of the object is established, and the Mean Value Coordinates (MVC) method is used to model the mapping relationship. This method has good smoothness and continuity, and can effectively avoid geometric folding or mutation.
[0032] In addition, for the mean value coordinate weight , let be located inside the cage, and form a polygon boundary with the control point . For each control point , calculate the vector relative to , and then calculate the weight using the above formula, wherein all are normalized and can be used for position interpolation. Through the above method, the slight perturbation of the control point in space will be transmitted to the entire object grid in a smooth and continuous manner, thereby realizing fine deformation of the geometric structure. Since the number of control points is much smaller than the number of grid vertices, this mechanism not only improves the computational efficiency, but also makes the optimization process more controllable and stable.
[0033] Preferably, the joint optimization objective function in S3 is calculated as follows:
[0034] wherein, is the target adversarial object generated after perturbation of the cage control point, is the lifting ability loss of the target adversarial object under the target grasping configuration, is the grasping stability index, is the regularization term based on Laplacian smoothing, and are balance coefficients.
[0035] Preferably, the regularization term based on Laplacian smoothing , the formula is as follows:
[0036] In the formula, all vertices of the target adversarial object, vertices adjacent vertex set.
[0037] Specifically, for S3, after the setting of the control points and the establishment of the mapping relationship thereof to the object geometry are completed, a key stage, iterative optimization and shape disturbance, is entered, aiming to systematically weaken the physical grasping performance of the object under the given grasping configuration through continuous and slight adjustment of the positions of the control points, two core physical indexes, lift capability and grasp stability, are specifically included, the above physical indexes are encoded into a loss function, and a regularization term is introduced to control the smoothness of the geometric deformation, a joint optimization objective function is constructed, and wherein is the lift capability loss of the target adversarial object under the target grasping configuration, and the lower the value is, the more difficult it is to resist gravity; is the grasp stability index, and the lower the value is, the more unstable it is; is the regularization term based on Laplacian smoothing, which is used to punish discontinuous or unnatural deformation of the geometric shape; and are balance coefficients, which are used to balance the grasping physical indexes and the degree of geometric deformation; the regularization term based on Laplacian smoothing , the regularization term ensures the smoothness of the disturbance and the consistency of the appearance of the object.
[0038] In addition, for S4, the optimization process adopts a multi-resolution strategy, in the initial stage, for an object with a unit size , a cage size of 0.04 is used to generate control points, and each control point is disturbed to quickly converge to the effective direction; then the control point grid is gradually refined (the cage size is halved at each step), to perform local fine adjustment with higher accuracy, and gradually approach the geometric state with the weakest performance. Each round of iteration uses optimization methods such as simulated annealing or gradient descent to update the positions of the control points, and evaluates the change of the loss function in real time to guide the optimization direction.
[0039] Finally, after multiple rounds of iterative optimization, the obtained adversarial object is highly similar in visual appearance to the original object, but greatly weakens the feasibility of the original grasping configuration at the physical level, thereby achieving the precise attack target of the grasping behavior of the present application.
[0040] The preferred embodiments of the application have been described above in detail. It should be understood that modifications and variations to the preferred embodiments could be made by those skilled in the art without departing from the spirit and scope of the application. Accordingly, it is intended that there be included within the scope of the application, all such modifications and variations as would be apparent to those skilled in the art upon reading this disclosure. It is intended to obtain for the inventors such patent rights as are available in any country on the world.
Claims
1. A robot grasping anti-attack method based on physical information, characterized in that: include: S1: Obtain the cage structure of the target adversarial object and generate cage control points; S2: Establish deformation mapping between cage control points and vertices of target adversarial object; S3: Constructing a joint optimization objective function that depends on physical indicators; S4: Optimize and adjust the cage control points based on deformation mapping and joint optimization objective function, and then output the optimized and adjusted target adversarial object.
2. The method for robot grasping against attacks based on physical information according to claim 1, characterized in that: The step of obtaining the cage structure of the target adversarial object in S1 includes: The target adversarial object is determined and represented in the form of a triangular mesh, including a vertex set and a face set. A three-dimensional Cartesian mesh volume is constructed based on the bounding box of the target adversarial object. The triangular mesh of the target adversarial object is divided into a number of equidistant cells in each dimension to form a regular three-dimensional voxel structure. The outer surface of the three-dimensional voxel structure is extracted to form a three-dimensional closed cage-shaped geometric body. The surface of the cage-shaped geometric body is composed of a number of rectangular or triangular facets. The cage-shaped structure, which serves as the target adversarial object, tightly wraps the target adversarial object.
3. The method for robot grasping against attacks based on physical information according to claim 2, characterized in that: The generation of cage control points in S1 includes: Each facet in the facet set corresponding to the cage structure is divided equally according to a preset resolution to form a uniform grid lattice. All grid intersections located on the surface of the cage structure in the division result constitute a cage control point set.
4. The method for robot grasping against attacks based on physical information according to claim 1, characterized in that: The S2 includes: For any target adversarial object vertex, its new position under the perturbation of the cage control points is expressed as a weighted linear combination of all cage control point motions, as follows: Where, For the The new positions of the vertices of the target adversarial object, for about The average coordinate weight of For the target adversarial object vertices, For the The three-dimensional coordinates of the cage control points, For the The new positions of the cage control points, is the total number of cage control points, and for each have , .
5. The method for robot grasping against attacks based on physical information according to claim 4, characterized in that: The average coordinate weight , the formula is as follows: Where, for Its adjacent cage control point is The angle of view, for Relative to vector.
6. The method for robot grasping against attacks based on physical information according to claim 1, characterized in that: The joint optimization objective function in S3 is formulated as follows: Where, is the target adversarial object generated after the cage-shaped control point perturbation, The lifting ability loss of the target adversarial object under the target grasping configuration, To capture stability indicators, is a regularization term based on Laplace smoothing, and is the balance coefficient.
7. The method for robot grasping against attacks based on physical information according to claim 6, characterized in that: The regularization term based on Laplace smoothing , the formula is as follows: Where, are all vertices of the target adversarial object, Vertex A set of adjacent vertices.
Citation Information
Patent Citations
Method and device for generating confrontation point cloud based on local information, medium and equipment
CN118823510A
Hybrid-mode anti-attack method and system based on laser radar point cloud
CN119312871A
Target capturing method based on 3D confrontation point cloud, electronic equipment and medium
CN119723031A
Intelligent competition system and robot
WO2018205102A1
Adversarial attack model training method and apparatus, adversarial image generation method and apparatus, electronic device, and storage medium
WO2021164334A1