Robot target grabbing method and system based on physical neural network

An improved SE(3)-Transformer network integrating multimodal perception and 3D physical modeling was used to construct a robot grasping system, which solved the problems of unstable grasping decisions and high computational burden in existing methods, and achieved efficient and stable grasping in complex environments.

CN120816490AInactive Publication Date: 2025-10-21YOUCHUANG POWER (WUXI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511139686.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing robot grasping methods lack a systematic understanding of physical laws when facing complex scenarios, resulting in a lack of stability and adaptability in grasping decisions. Furthermore, they impose a heavy computational burden and deployment requirements, making it difficult to reduce computational costs while ensuring grasping quality.

Method used

A physical neural network-based approach is adopted, which integrates multimodal perception modeling, three-dimensional physical modeling and an improved SE(3)-Transformer network structure to build a system with grasping strategy generation and iterative optimization capabilities. Stable grasping is achieved through joint training and feedback-driven updates of the physical modeling module and the neural network module.

Benefits of technology

It improves the stability and adaptability of robot grasping, reduces the computational burden, enables efficient grasping in complex environments, and has adaptive adjustment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120816490A_ABST
    Figure CN120816490A_ABST
Patent Text Reader

Abstract

The invention discloses a robot target grabbing method based on a physical neural network, and the method comprises the following steps: S1, collecting visual, tactile and force data, and generating fusion perception data; s2, constructing a model comprising physical modeling and a neural network module, and outputting physical state features; s3, inputting fusion perception data and physical state characteristics, and generating grabbing strategy parameters; s4, a joint loss function is constructed based on the grabbing error, and model training is completed; s5, generating grabbing strategy parameters in real time in an application stage; s6, executing a grabbing action and collecting feedback data; and S7, feeding back the data input model, and iteratively updating the grabbing strategy until stable grabbing. Physical modeling and a neural network structure are fused, and the grabbing stability and adaptability of the robot are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control and artificial intelligence technology, and in particular to a robot target grasping method and system based on physical neural network. Background Art

[0002] In the research and application of intelligent robot grasping control, most existing methods are based on classical control models or rely purely on neural network structures. Although such methods perform well in industrial environments with standard structures and clear task settings, they often appear powerless in complex scenarios, especially when faced with irregular targets, unstable support surfaces, or dynamic disturbances during grasping. A major problem is that traditional neural networks lack the systematic introduction of physical laws in model design, and the generation of strategies relies on fitting a large number of samples. In essence, they are "talking by pictures" and there is no clear modeling of the target's force state and motion trend, resulting in a lack of stability in grasping decisions.

[0003] In addition, although some current grasping models that integrate perceptual data attempt to integrate visual, tactile or force information, there is a lack of structural correlation between the data. The data are often input into the network in a simple splicing or stacking manner, which results in a lot of information redundancy and limited expressive ability. The action feedback after the grasping is completed is only passively recorded in many systems. There are few mechanisms to effectively introduce feedback into the strategy iteration, that is, there is a lack of "looking back" on the grasping results. This makes it impossible for the strategy to self-correct in real time. Once a deviation occurs, it can only rely on external intervention and lacks adaptive capabilities.

[0004] Existing methods also basically ignore the structural redundancy problem in the strategy generation process. Although neural networks have powerful representation capabilities, when deployed on edge computing platforms or actual robot ends, the model's computational complexity and structural complexity are still key bottlenecks. Existing models lack support for structural reconfiguration mechanisms, making it difficult to reduce computational consumption and deployment burden while ensuring grasping quality.

[0005] Therefore, how to provide a robot target grasping method and system based on physical neural network is a problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0006] One purpose of the present invention is to propose a robot target grasping method based on physical neural network. The present invention integrates multimodal perception modeling, three-dimensional physical modeling and improved SE(3)-Transformer network structure, constructs a physical neural network system with grasping strategy generation and iterative optimization capabilities, and describes in detail the complete process from perception information fusion, force state modeling, grasping strategy parameter generation, feedback drive update to stable grasping judgment. The method has the advantages of strong adaptability, high grasping stability and high deployment efficiency.

[0007] A robot target grasping method based on a physical neural network according to an embodiment of the present invention includes the following steps:

[0008] S1, collects visual data, tactile data and force data of the target object and the environment, performs filtering, normalization and feature extraction operations, and generates fused perception data;

[0009] S2. Constructing a physical neural network model, including a physical modeling module and a neural network module. The physical modeling module models the force and motion state of the fused perception data according to physical laws and outputs physical state characteristics.

[0010] S3, input the fused perception data and physical state characteristics into the neural network module to generate grasping strategy parameters;

[0011] S4. Based on the error between the grasping strategy parameters and the actual grasping action, a joint loss function is constructed and backpropagation is performed to optimize the parameters of the physical modeling module and the neural network module to complete model training.

[0012] S5. In the application phase, the real-time collected fusion perception data is input into the trained physical neural network model to output the grasping strategy parameters for the target grasping;

[0013] S6. Control the robot to perform a grasping action according to the grasping strategy parameters and collect feedback data during the grasping process, wherein the feedback data includes contact force changes, object posture changes, and stability status;

[0014] S7. Input the feedback data and the grasping strategy parameters into the physical neural network model, and update and generate new grasping strategy parameters until the target object is grasped stably.

[0015] Optionally, the neural network module adopts an improved SE(3)-Transformer, specifically including:

[0016] A multimodal embedding structure is set up to encode the visual features, tactile signals, and force data in the fused perception data into a unified spatial vector form as the input representation of the neural network module;

[0017] Introducing mechanical constraints into the attention calculation process, building an attention weight adjustment mechanism based on physical state characteristics to generate an attention map that includes physical attribute perception;

[0018] Construct a time embedding structure to encode the time series feedback information in the grasping process into time series node features, and input them together with the fused perception data into the improved SE(3)-Transformer structure to perform time-space joint modeling operations;

[0019] A reconfigurable control unit is introduced to perform pruning operations on high-order tensor channels in the improved SE(3)-Transformer structure. The reconfigurable control unit includes a response amplitude analysis module, a dynamic pruning module, and a structure adjustment module. Channel retention selection is performed based on the response amplitude of the corresponding channel of the physical state characteristics, and a combined tensor structure is generated through the structure adjustment module to replace the original tensor output path.

[0020] Set up the decoder branch and input the combined tensor structure into the mapping module, mapping it into three types of grasping strategy parameters: grasping position, grasping posture, and grasping force.

[0021] Optionally, the S2 specifically includes:

[0022] S21, constructing a physical neural network model, wherein the physical neural network model includes a physical modeling module and a neural network module;

[0023] S22. Inputting the fused perception data into a physical modeling module. The physical modeling module constructs a force and motion state modeling mechanism based on the mechanical laws in three-dimensional space, jointly models the force changes, velocity changes, and displacement changes of the target object during the grasping process, and generates physical state characteristics.

[0024] S23. In the modeling process, the resultant force of the target object is formed by a weighted combination of the product term of mass and acceleration, the product term of velocity and damping coefficient, and the product term of displacement and elastic coefficient.

[0025] Optionally, the S3 specifically includes:

[0026] S31, concatenating the fused perception data with the physical state features to generate a joint input tensor, which serves as the input of the neural network module;

[0027] S32. Build an attention calculation mechanism in the neural network module to generate attention weights based on the spatial information and physical state features in the joint input tensor:

[0028]

[0029] Among them, α ij represents the attention weight of the i-th position to the j-th adjacent position, x i represents the fusion perception feature vector of the i-th position, f i represents the physical state eigenvector of the i-th position, ψ q (·,·) represents the mapping function used to generate the query vector, ψ k (·,·) represents the mapping function used to generate the key vector, <·,·> represents the inner product operation, represents the adjacent set of the i-th position, exp(·) represents the exponential function;

[0030] S33, generating a high-order representation tensor based on the attention weight, and inputting the high-order representation tensor into the reconfigurable control unit, and generating a combined tensor structure through the structure adjustment module in the reconfigurable control unit;

[0031] S34. Input the combined tensor structure into the decoder branch and map it into three types of grasping strategy parameters: grasping position, grasping posture and grasping force.

[0032] Optionally, the S33 specifically includes:

[0033] S331, performing a multi-head aggregation operation on the intermediate representation tensor within the neural network module based on the attention weight to generate a high-order representation tensor containing multi-channel spatial semantic features;

[0034] S332, inputting the high-order representation tensor into a reconfigurable control unit, wherein the reconfigurable control unit includes a response amplitude analysis module, a dynamic clipping module, and a structure adjustment module;

[0035] S333. Calculate the response amplitude vector of each tensor channel through the response amplitude analysis module:

[0036]

[0037] Among them, r c Indicates the response amplitude value of channel number c, T cij Represents the characteristic intensity value of the cth channel at the spatial position (i, j), H is the height dimension of the tensor, and W is the width dimension of the tensor;

[0038] S334: Perform channel screening based on the response threshold set in the dynamic clipping module based on the response amplitude vector, retain channels with response amplitudes higher than the threshold, and mark the channels to be clipped;

[0039] S335. The channel input structure adjustment module is retained to construct a combined tensor structure, where the combined tensor structure integrates the retained multi-channel tensors in the channel dimension and rearranges the order to generate a decoder input.

[0040] Optionally, the S4 specifically includes:

[0041] S41, obtaining grasping strategy parameters output by the neural network module, wherein the grasping strategy parameters are generated by inputting the fusion perception data and the physical state characteristics into the neural network module, and include three types of parameters: grasping position, grasping posture, and grasping force;

[0042] S42, executing the robot grasping action, recording the actual grasping position, execution posture and applied force, and generating real grasping action data;

[0043] S43, comparing the grasping strategy parameters with the actual grasping action data, calculating the differences in the three parameters in terms of spatial position, posture direction, and force value, and generating an error vector;

[0044] S44. Construct a joint loss function. The joint loss function takes the error vector as input and includes a policy error term and a modeling error term. The policy error term is calculated by the difference between the grasping policy parameters and the actual grasping action data, and the modeling error term is calculated by the deviation between the physical state characteristics and the feedback data. The two are weightedly combined to form an overall loss expression.

[0045] S45. Taking the joint loss function as the optimization target, the internal parameters of the physical modeling module and the neural network module are updated through the back-propagation mechanism to complete the training process of the physical neural network model.

[0046] Optionally, the S6 specifically includes:

[0047] S61, controlling the robot to perform a grasping action according to grasping strategy parameters output by the physical neural network model, wherein the grasping strategy parameters include grasping position, grasping posture, and grasping force;

[0048] S62. Collect feedback data during the grasping action, where the feedback data includes contact force change, object posture change, and stability state. The contact force change is the force change vector of the end effector during the grasping contact process, the object posture change is the offset representation of the target object's posture in three-dimensional space, and the stability state is a scalar or label information indicating the stability of the grasping result state.

[0049] S63, performing normalization and structuring processing operations on the feedback data, encoding the contact force change, the object posture change, and the stability state into standard vector expressions, and constructing a feedback state vector set;

[0050] S64. Establish a mapping relationship between the feedback state vector set and the grasping strategy parameters, generate a binding data record, and send it as input to the physical neural network model to update the grasping strategy parameters and optimize the structural weight configuration of the physical modeling module and the neural network module.

[0051] Optionally, the S7 specifically includes:

[0052] S71, combining the collected contact force change, object posture change, and stability state with the grasping position, grasping posture, and grasping force generated in the previous round to form grasping strategy parameters, and jointly encoding them to form a joint input vector;

[0053] S72. Input the joint input vector into the physical neural network model, wherein the physical modeling module regenerates the physical state features based on the joint input vector, and the neural network module performs attention map construction, structure cutting and strategy mapping processes based on the physical state features, and outputs the updated grasping position, grasping posture and grasping force;

[0054] S73, controlling the robot to perform a grasping action according to the updated grasping position, grasping posture, and grasping force, and re-collecting contact force changes, object posture changes, and stability status;

[0055] S74. Construct a stable grasping criterion function based on the re-collected data. The criterion function includes three dimensions: the end contact force change rate, the object posture angle offset value, and the stability classification state, and is used to determine whether the current grasping state meets the stability condition.

[0056] S75. If all output values ​​of the stable grasping criterion function meet the preset threshold conditions, the final grasping position, grasping posture and grasping force are output. If not, repeat steps S71 to S74 until the grasping strategy parameters that meet the stability conditions are generated.

[0057] A robot target grasping system based on a physical neural network according to an embodiment of the present invention includes:

[0058] The sensory data acquisition module is used to collect visual data, tactile data, and force data of the target object and the environment, perform filtering, normalization, and feature extraction operations, and generate fused sensory data;

[0059] The physical modeling module is used to build a force and motion state modeling mechanism in three-dimensional space based on the fused perception data, generate physical state features, and calculate the modeling error during the model update process;

[0060] The neural network module adopts an improved SE(3)-Transformer structure, which takes as input the fusion perception data and physical state features, generates a combined tensor structure through a multimodal embedding structure, an attention weight generation mechanism, and a reconfigurable control unit, and outputs the grasping position, grasping posture, and grasping force through the decoder branch.

[0061] The grasping execution module is used to control the robot to execute grasping actions according to the grasping strategy parameters, collect contact force changes, object posture changes and stability status, and construct feedback data;

[0062] The feedback encoding module is used to normalize and structure the contact force changes, object posture changes, and stability states, generate a feedback state vector set, and bind it with the grasping strategy parameters to form a joint input vector;

[0063] The strategy iteration update module is used to input the joint input vector into the physical neural network model, regenerate the physical state characteristics and grasping strategy parameters, and judge whether the grasping state meets the preset conditions based on the stable grasping criterion function. If not, it will continue to iterate until a stable grasp is achieved;

[0064] The joint training module is used to construct a joint loss function that includes the policy error term and the modeling error term, perform backpropagation operations using the error vector, and optimize the parameter configuration of the physical modeling module and the neural network module.

[0065] The beneficial effects of the present invention are:

[0066] This paper addresses the uncertainty and multi-source sensory data fusion challenges faced by robots in actual grasping tasks, and designs a grasping decision-making method that integrates physical modeling and neural network calculations. By integrating the results of mechanical modeling with multimodal sensory information such as vision, touch, and force into the neural network model, not only is the accuracy of understanding the object's state improved, but the physical rationality of strategy generation is also enhanced. In particular, by collecting feedback data after the grasping action is executed and participating in the strategy update process in real time, the model has adaptive adjustment capabilities and can dynamically modify grasping parameters based on changes in the environment and object state.

[0067] In addition, the system adopts an improved SE(3)-Transformer network structure to embed temporal feedback modeling, and improves the compactness and response efficiency of policy output through channel pruning and tensor structure reconstruction. The design of the joint loss function ensures that the model optimizes both the policy execution error and the physical prediction error during the training phase, maintaining the closed-loop and stability of the prediction logic. In terms of grasping judgment, a comprehensive judgment function including the contact force change rate, posture angle offset and stability label is constructed, providing a clear and objective standard for judging whether the grasp is completed. Overall, this method effectively makes up for the shortcomings of existing deep learning grasping models in physical consistency, feedback utilization and policy iteration, and provides solid algorithmic support for robots to achieve stable and efficient grasping in practical environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0069] Figure 1 This is a flow chart of a robot target grasping method based on physical neural network proposed by the present invention;

[0070] Figure 2 This is a schematic diagram of the network structure of a robot target grasping method based on physical neural network proposed in the present invention;

[0071] Figure 3 This is the experimental system deployment diagram of the robot target grasping method based on physical neural network proposed in this invention. DETAILED DESCRIPTION

[0072] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0073] refer to Figure 1-3 , a robot target grasping method based on physical neural network, comprising the following steps:

[0074] S1, collects visual data, tactile data and force data of the target object and the environment, performs filtering, normalization and feature extraction operations, and generates fused perception data;

[0075] S2. Constructing a physical neural network model, including a physical modeling module and a neural network module. The physical modeling module models the force and motion state of the fused perception data according to physical laws and outputs physical state characteristics.

[0076] S3, input the fused perception data and physical state characteristics into the neural network module to generate grasping strategy parameters;

[0077] S4. Based on the error between the grasping strategy parameters and the actual grasping action, a joint loss function is constructed and backpropagation is performed to optimize the parameters of the physical modeling module and the neural network module to complete model training.

[0078] S5. In the application phase, the real-time collected fusion perception data is input into the trained physical neural network model to output the grasping strategy parameters for the target grasping;

[0079] S6. Control the robot to perform a grasping action according to the grasping strategy parameters and collect feedback data during the grasping process, wherein the feedback data includes contact force changes, object posture changes, and stability status;

[0080] S7. Input the feedback data and the grasping strategy parameters into the physical neural network model, and update and generate new grasping strategy parameters until the target object is grasped stably.

[0081] This invention uses a multimodal perception method that integrates visual, tactile, and force data, combined with a structural design that combines physical modeling with neural network decision-making, to achieve intelligent generation and dynamic adjustment of target object grasping strategies. This method not only improves the model's responsiveness to changes in mechanical states in complex environments, but also enhances the stability and reliability of the robot's grasping actions.

[0082] In this embodiment, the neural network module adopts an improved SE(3)-Transformer, which specifically includes:

[0083] A multimodal embedding structure is set up to encode the visual features, tactile signals, and force data in the fused perception data into a unified spatial vector form as the input representation of the neural network module;

[0084] Introducing mechanical constraints into the attention calculation process, building an attention weight adjustment mechanism based on physical state characteristics to generate an attention map that includes physical attribute perception;

[0085] Construct a time embedding structure to encode the time series feedback information in the grasping process into time series node features, and input them together with the fused perception data into the improved SE(3)-Transformer structure to perform time-space joint modeling operations;

[0086] A reconfigurable control unit is introduced to perform pruning operations on high-order tensor channels in the improved SE(3)-Transformer structure. The reconfigurable control unit includes a response amplitude analysis module, a dynamic pruning module, and a structure adjustment module. Channel retention selection is performed based on the response amplitude of the corresponding channel of the physical state characteristics, and a combined tensor structure is generated through the structure adjustment module to replace the original tensor output path.

[0087] Set up the decoder branch and input the combined tensor structure into the mapping module, mapping it into three types of grasping strategy parameters: grasping position, grasping posture, and grasping force.

[0088] This paper achieves a deep fusion of spatial information and mechanical features through an improved SE(3)-Transformer structure, introduces mechanical constraints into the attention mechanism, and strengthens the model's ability to respond to changes in physical state. At the same time, the model's computational burden is effectively reduced through structural pruning and combined tensor construction mechanisms, making the grasping strategy parameter generation process more efficient and deployable.

[0089] In this embodiment, S2 specifically includes:

[0090] S21, constructing a physical neural network model, wherein the physical neural network model includes a physical modeling module and a neural network module;

[0091] S22. Inputting the fused perception data into a physical modeling module. The physical modeling module constructs a force and motion state modeling mechanism based on the mechanical laws in three-dimensional space, jointly models the force changes, velocity changes, and displacement changes of the target object during the grasping process, and generates physical state characteristics.

[0092] S23. In the modeling process, the resultant force of the target object is formed by a weighted combination of the product term of mass and acceleration, the product term of velocity and damping coefficient, and the product term of displacement and elastic coefficient.

[0093] The present invention introduces a joint modeling mechanism based on the force and motion laws of three-dimensional space in the physical modeling module, and clarifies the expression of the resultant force as a weighted combination of mechanical terms such as mass, velocity, damping and elastic coefficient, making the generated physical state characteristics more physically interpretable and providing more realistic and reliable physical semantic support for neural network input.

[0094] In this embodiment, S3 specifically includes:

[0095] S31, concatenating the fused perception data with the physical state features to generate a joint input tensor, which serves as the input of the neural network module;

[0096] S32. Build an attention calculation mechanism in the neural network module to generate attention weights based on the spatial information and physical state features in the joint input tensor:

[0097]

[0098] Among them, α ij represents the attention weight of the i-th position to the j-th adjacent position, x i represents the fusion perception feature vector of the i-th position, f i represents the physical state eigenvector of the i-th position, ψ q (·,·) represents the mapping function used to generate the query vector, ψ k (·,·) represents the mapping function used to generate the key vector, <·,·> represents the inner product operation, represents the adjacent set of the i-th position, exp(·) represents the exponential function;

[0099] S33, generating a high-order representation tensor based on the attention weight, and inputting the high-order representation tensor into the reconfigurable control unit, and generating a combined tensor structure through the structure adjustment module in the reconfigurable control unit;

[0100] S34. Input the combined tensor structure into the decoder branch and map it into three types of grasping strategy parameters: grasping position, grasping posture and grasping force.

[0101] This method combines sensory information and physical state features into a neural network input tensor. Attention weights are generated based on spatial and physical relationships, effectively guiding the network's focus on key force areas and posture trends during the grasping process. The resulting combined tensor structure supports the refined output of multiple types of strategy parameters, providing structural support for grasping control strategies.

[0102] In this embodiment, the S33 specifically includes:

[0103] S331, performing a multi-head aggregation operation on the intermediate representation tensor within the neural network module based on the attention weight to generate a high-order representation tensor containing multi-channel spatial semantic features;

[0104] S332, inputting the high-order representation tensor into a reconfigurable control unit, wherein the reconfigurable control unit includes a response amplitude analysis module, a dynamic clipping module, and a structure adjustment module;

[0105] S333. Calculate the response amplitude vector of each tensor channel through the response amplitude analysis module:

[0106]

[0107] Among them, r c Indicates the response amplitude value of channel number c, T cij Represents the characteristic intensity value of the cth channel at the spatial position (i, j), H is the height dimension of the tensor, and W is the width dimension of the tensor;

[0108] S334: Perform channel screening based on the response threshold set in the dynamic clipping module based on the response amplitude vector, retain channels with response amplitudes higher than the threshold, and mark the channels to be clipped;

[0109] S335. The channel input structure adjustment module is retained to construct a combined tensor structure, where the combined tensor structure integrates the retained multi-channel tensors in the channel dimension and rearranges the order to generate a decoder input.

[0110] This paper optimizes the structure of high-order tensors in neural networks through mechanisms such as response amplitude analysis, dynamic pruning, and structural reconstruction. This not only improves the ability to focus on significant channels, but also effectively eliminates redundant information and optimizes the input representation of the decoder branches. This structural pruning mechanism enhances the model's generalization ability and improves operational efficiency.

[0111] In this embodiment, the S4 specifically includes:

[0112] S41, obtaining grasping strategy parameters output by the neural network module, wherein the grasping strategy parameters are generated by inputting the fusion perception data and the physical state characteristics into the neural network module, and include three types of parameters: grasping position, grasping posture, and grasping force;

[0113] S42, executing the robot grasping action, recording the actual grasping position, execution posture and applied force, and generating real grasping action data;

[0114] S43, comparing the grasping strategy parameters with the actual grasping action data, calculating the differences in the three parameters in terms of spatial position, posture direction, and force value, and generating an error vector;

[0115] S44. Construct a joint loss function. The joint loss function takes the error vector as input and includes a policy error term and a modeling error term. The policy error term is calculated by the difference between the grasping policy parameters and the actual grasping action data, and the modeling error term is calculated by the deviation between the physical state characteristics and the feedback data. The two are weightedly combined to form an overall loss expression.

[0116] S45. Taking the joint loss function as the optimization target, the internal parameters of the physical modeling module and the neural network module are updated through the back-propagation mechanism to complete the training process of the physical neural network model.

[0117] This paper integrates the dual constraints of policy error and modeling error into the loss function design, and constructs a joint loss expression based on the error vector. This improves the training process's ability to simultaneously correct for deviations in both action execution and mechanical modeling. Through the backpropagation mechanism, the physical modeling module and the neural network module are collaboratively optimized, achieving a more stable parameter convergence process.

[0118] In this embodiment, S6 specifically includes:

[0119] S61, controlling the robot to perform a grasping action according to grasping strategy parameters output by the physical neural network model, wherein the grasping strategy parameters include grasping position, grasping posture, and grasping force;

[0120] S62. Collect feedback data during the grasping action, where the feedback data includes contact force change, object posture change, and stability state. The contact force change is the force change vector of the end effector during the grasping contact process, the object posture change is the offset representation of the target object's posture in three-dimensional space, and the stability state is a scalar or label information indicating the stability of the grasping result state.

[0121] S63, performing normalization and structuring processing operations on the feedback data, encoding the contact force change, the object posture change, and the stability state into standard vector expressions, and constructing a feedback state vector set;

[0122] S64. Establish a mapping relationship between the feedback state vector set and the grasping strategy parameters, generate a binding data record, and send it as input to the physical neural network model to update the grasping strategy parameters and optimize the structural weight configuration of the physical modeling module and the neural network module.

[0123] This paper designs a structured feedback data collection and encoding mechanism that maps contact force, posture change, and stability state into standard vectors. These vectors are then bound to strategy parameters to form a complete data input pathway. This mechanism enables the physical neural network to continuously modify the strategy generation path based on actual grasping feedback, improving the system's grasping execution stability.

[0124] In this embodiment, the S7 specifically includes:

[0125] S71, combining the collected contact force change, object posture change, and stability state with the grasping position, grasping posture, and grasping force generated in the previous round to form grasping strategy parameters, and jointly encoding them to form a joint input vector;

[0126] S72. Input the joint input vector into the physical neural network model, wherein the physical modeling module regenerates the physical state features based on the joint input vector, and the neural network module performs attention map construction, structure cutting and strategy mapping processes based on the physical state features, and outputs the updated grasping position, grasping posture and grasping force;

[0127] S73, controlling the robot to perform a grasping action according to the updated grasping position, grasping posture, and grasping force, and re-collecting contact force changes, object posture changes, and stability status;

[0128] S74. Construct a stable grasping criterion function based on the re-collected data. The criterion function includes three dimensions: the end contact force change rate, the object posture angle offset value, and the stability classification state, and is used to determine whether the current grasping state meets the stability condition.

[0129] S75. If all output values ​​of the stable grasping criterion function meet the preset threshold conditions, the final grasping position, grasping posture and grasping force are output. If not, repeat steps S71 to S74 until the grasping strategy parameters that meet the stability conditions are generated.

[0130] This paper establishes an iterative update process based on the combined encoding of feedback data and grasping parameters. It uses a stability criterion function to comprehensively evaluate the effectiveness of grasping actions and optimizes the strategy based on the model output. This feedback-driven mechanism ensures that the robot continuously adjusts its strategy until stable grasping is achieved in a dynamic environment, significantly improving its operational adaptability.

[0131] A robot target grasping system based on physical neural network, comprising:

[0132] The sensory data acquisition module is used to collect visual data, tactile data, and force data of the target object and the environment, perform filtering, normalization, and feature extraction operations, and generate fused sensory data;

[0133] The physical modeling module is used to build a force and motion state modeling mechanism in three-dimensional space based on the fused perception data, generate physical state features, and calculate the modeling error during the model update process;

[0134] The neural network module adopts an improved SE(3)-Transformer structure, which takes as input the fusion perception data and physical state features, generates a combined tensor structure through a multimodal embedding structure, an attention weight generation mechanism, and a reconfigurable control unit, and outputs the grasping position, grasping posture, and grasping force through the decoder branch.

[0135] The grasping execution module is used to control the robot to execute grasping actions according to the grasping strategy parameters, collect contact force changes, object posture changes and stability status, and construct feedback data;

[0136] The feedback encoding module is used to normalize and structure the contact force changes, object posture changes, and stability states, generate a feedback state vector set, and bind it with the grasping strategy parameters to form a joint input vector;

[0137] The strategy iteration update module is used to input the joint input vector into the physical neural network model, regenerate the physical state characteristics and grasping strategy parameters, and judge whether the grasping state meets the preset conditions based on the stable grasping criterion function. If not, it will continue to iterate until a stable grasp is achieved;

[0138] The joint training module is used to construct a joint loss function that includes the policy error term and the modeling error term, perform backpropagation operations using the error vector, and optimize the parameter configuration of the physical modeling module and the neural network module.

[0139] The proposed robotic object grasping system has a clear overall structure and efficient inter-module collaboration, enabling a closed-loop control process from sensory input to grasping action execution and feedback update. The system's physical modeling module, SE(3)-Transformer neural network module, and policy iteration update mechanism work together to enhance the generalization capability, execution accuracy, and policy robustness of object grasping tasks in complex physical environments.

[0140] Example 1:

[0141] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the sorting scenario of special-shaped parts on the assembly line. In this scenario, the robot needs to flexibly adjust the grasping strategy according to the real-time perception results to cope with challenges such as changes in object position, complex shapes, and environmental interference.

[0142] The system hardware is based on a six-degree-of-freedom robotic arm with three types of sensors integrated at its end: a depth camera (640×480 resolution, 30fps) for acquiring spatial information about the target object; a set of flexible tactile sensors for sensing surface pressure distribution; and a three-dimensional force sensor with an accuracy of 0.1N and a range of 0-100N for real-time monitoring of force changes during the grasping process. The entire system is controlled by an industrial computer equipped with an Intel Core i7 processor, 16GB of memory, and an NVIDIA RTX 3060 graphics card, which performs data processing and model inference.

[0143] In terms of software, we first built a multimodal perception module. Using Python and OpenCV tools, we capture, filter, and extract features from depth images, extracting visual information such as contours, edges, and center of mass. We also call on sensor interfaces to collect tactile and force data, normalize them, and fuse them with visual features to form the model input.

[0144] In terms of model architecture, a dual-module structure of physical neural networks is adopted. The physical modeling part uses basic physical laws such as Newton's laws of motion and friction models, combined with sensor data, to infer the acceleration, velocity, and displacement changes of the target object during the force process, and output physical feature quantities reflecting the real dynamic state. The neural network part adopts an improved SE(3)-Transformer structure, which encodes the fusion perception data and physical state characteristics into spatial vectors, combines mechanical constraints and time embedding mechanisms, and learns the grasping strategy parameters. This structure further uses a reconfigurable control unit to cut and combine intermediate tensor channels, improve network computing efficiency, and output control parameters in three dimensions: grasping position, grasping posture, and grasping force.

[0145] During model training, a dataset containing 10,000 data samples was constructed, covering objects of varying shapes and materials (such as metal blocks, plastic balls, and wooden prisms). Various lighting and background environments were simulated to enhance model robustness. The training and test sets were split 8:2. The Adam optimizer was used with a learning rate of 0.001, 50 training epochs, and the mean squared error (MSE) loss function. Evaluation and parameter adjustments were performed on the test set every five epochs.

[0146] After the model was deployed, three test scenarios were set up in the lab: static grasping in a simple background, static grasping in a complex background, and grasping a dynamic target. By comparing the model with traditional geometric model methods, visual servoing methods, and common CNN methods, the average grasping success rate, grasping time, and posture error of each method in the same scenario were recorded. The results are shown in Table 1:

[0147] Table 1 Performance comparison of different crawling methods

[0148]

[0149] Experimental results show that the proposed method outperforms comparable solutions in key metrics such as grasping success rate, time consumption, and accuracy. In particular, the introduction of a physical model and feedback mechanism effectively addresses the issues of object slippage and gripper posture deviation in grasping dynamic objects. For example, in the grasping of a metal sphere, the average success rate remained above 90% despite complex backgrounds and lighting interference.

[0150] Therefore, the present invention demonstrates stability and adaptability in multimodal perception fusion, grasping strategy generation and dynamic adjustment, and can meet the requirements of practical applications such as flexible manufacturing, intelligent logistics and unmanned sorting.

[0151] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A robot target grasping method based on physical neural network, characterized in that: The steps include: S1, collects visual data, tactile data and force data of the target object and the environment, performs filtering, normalization and feature extraction operations, and generates fused perception data; S2. Constructing a physical neural network model, including a physical modeling module and a neural network module. The physical modeling module models the force and motion state of the fused perception data according to physical laws and outputs physical state characteristics. S3, input the fused perception data and physical state characteristics into the neural network module to generate grasping strategy parameters; S4. Based on the error between the grasping strategy parameters and the actual grasping action, a joint loss function is constructed and backpropagation is performed to optimize the parameters of the physical modeling module and the neural network module to complete model training. S5. In the application phase, the real-time collected fusion perception data is input into the trained physical neural network model to output the grasping strategy parameters for the target grasping; S6. Control the robot to perform a grasping action according to the grasping strategy parameters and collect feedback data during the grasping process, wherein the feedback data includes contact force changes, object posture changes, and stability status; S7. Input the feedback data and the grasping strategy parameters into the physical neural network model, and update and generate new grasping strategy parameters until the target object is grasped stably.

2. A robot target grasping method based on physical neural network according to claim 1, characterized in that: The neural network module adopts an improved SE(3)-Transformer, which specifically includes: A multimodal embedding structure is set up to encode the visual features, tactile signals, and force data in the fused perception data into a unified spatial vector form as the input representation of the neural network module; Introducing mechanical constraints into the attention calculation process, building an attention weight adjustment mechanism based on physical state characteristics to generate an attention map that includes physical attribute perception; Construct a time embedding structure to encode the time series feedback information in the grasping process into time series node features, and input them together with the fused perception data into the improved SE(3)-Transformer structure to perform time-space joint modeling operations; A reconfigurable control unit is introduced to perform pruning operations on high-order tensor channels in the improved SE(3)-Transformer structure. The reconfigurable control unit includes a response amplitude analysis module, a dynamic pruning module, and a structure adjustment module. Channel retention selection is performed based on the response amplitude of the corresponding channel of the physical state characteristics, and a combined tensor structure is generated through the structure adjustment module to replace the original tensor output path. Set up the decoder branch and input the combined tensor structure into the mapping module, mapping it into three types of grasping strategy parameters: grasping position, grasping posture, and grasping force.

3. The method for grasping a target by a robot based on a physical neural network according to claim 1, wherein: The S2 specifically includes: S21, constructing a physical neural network model, wherein the physical neural network model includes a physical modeling module and a neural network module; S22. Inputting the fused perception data into a physical modeling module. The physical modeling module constructs a force and motion state modeling mechanism based on the mechanical laws in three-dimensional space, jointly models the force changes, velocity changes, and displacement changes of the target object during the grasping process, and generates physical state characteristics. S23. In the modeling process, the resultant force of the target object is formed by a weighted combination of the product term of mass and acceleration, the product term of velocity and damping coefficient, and the product term of displacement and elastic coefficient.

4. The method for grasping a target by a robot based on a physical neural network according to claim 1, wherein: The S3 specifically includes: S31, concatenating the fused perception data with the physical state features to generate a joint input tensor, which serves as the input of the neural network module; S32. Build an attention calculation mechanism in the neural network module to generate attention weights based on the spatial information and physical state features in the joint input tensor: Among them, α ij represents the attention weight of the i-th position to the j-th adjacent position, x i represents the fusion perception feature vector of the i-th position, f i represents the physical state eigenvector of the i-th position, ψ q (·,·) represents the mapping function used to generate the query vector, ψ k (·,·) represents the mapping function used to generate the key vector, <·,·> represents the inner product operation, represents the adjacent set of the i-th position, exp(·) represents the exponential function; S33, generating a high-order representation tensor based on the attention weight, and inputting the high-order representation tensor into the reconfigurable control unit, and generating a combined tensor structure through the structure adjustment module in the reconfigurable control unit; S34. Input the combined tensor structure into the decoder branch and map it into three types of grasping strategy parameters: grasping position, grasping posture and grasping force.

5. The method for robot target grasping based on physical neural network according to claim 4, characterized in that: The S33 specifically includes: S331, performing a multi-head aggregation operation on the intermediate representation tensor within the neural network module based on the attention weight to generate a high-order representation tensor containing multi-channel spatial semantic features; S332, inputting the high-order representation tensor into a reconfigurable control unit, wherein the reconfigurable control unit includes a response amplitude analysis module, a dynamic clipping module, and a structure adjustment module; S333. Calculate the response amplitude vector of each tensor channel through the response amplitude analysis module: Among them, r c Indicates the response amplitude value of channel number c, T cij Represents the characteristic intensity value of the cth channel at the spatial position (i, j), H is the height dimension of the tensor, and W is the width dimension of the tensor; S334: Perform channel screening based on the response threshold set in the dynamic clipping module based on the response amplitude vector, retain channels with response amplitudes higher than the threshold, and mark the channels to be clipped; S335. The channel input structure adjustment module is retained to construct a combined tensor structure, where the combined tensor structure integrates the retained multi-channel tensors in the channel dimension and rearranges the order to generate a decoder input.

6. The method for grasping a target by a robot based on a physical neural network according to claim 1, wherein: The S4 specifically includes: S41, obtaining grasping strategy parameters output by the neural network module, wherein the grasping strategy parameters are generated by inputting the fusion perception data and the physical state characteristics into the neural network module, and include three types of parameters: grasping position, grasping posture, and grasping force; S42, executing the robot grasping action, recording the actual grasping position, execution posture and applied force, and generating real grasping action data; S43, comparing the grasping strategy parameters with the actual grasping action data, calculating the differences in the three parameters in terms of spatial position, posture direction, and force value, and generating an error vector; S44. Construct a joint loss function. The joint loss function takes the error vector as input and includes a policy error term and a modeling error term. The policy error term is calculated by the difference between the grasping policy parameters and the actual grasping action data, and the modeling error term is calculated by the deviation between the physical state characteristics and the feedback data. The two are weightedly combined to form an overall loss expression. S45. Taking the joint loss function as the optimization target, the internal parameters of the physical modeling module and the neural network module are updated through the back-propagation mechanism to complete the training process of the physical neural network model.

7. The method for grasping a target by a robot based on a physical neural network according to claim 1, characterized in that: The S6 specifically includes: S61, controlling the robot to perform a grasping action according to grasping strategy parameters output by the physical neural network model, wherein the grasping strategy parameters include grasping position, grasping posture, and grasping force; S62. Collect feedback data during the grasping action, where the feedback data includes contact force change, object posture change, and stability state. The contact force change is the force change vector of the end effector during the grasping contact process, the object posture change is the offset representation of the target object's posture in three-dimensional space, and the stability state is a scalar or label information indicating the stability of the grasping result state. S63, performing normalization and structuring processing operations on the feedback data, encoding the contact force change, the object posture change, and the stability state into standard vector expressions, and constructing a feedback state vector set; S64. Establish a mapping relationship between the feedback state vector set and the grasping strategy parameters, generate a binding data record, and send it as input to the physical neural network model to update the grasping strategy parameters and optimize the structural weight configuration of the physical modeling module and the neural network module.

8. The method for robot target grasping based on physical neural network according to claim 1, characterized in that: The S7 specifically includes: S71, combining the collected contact force change, object posture change, and stability state with the grasping position, grasping posture, and grasping force generated in the previous round to form grasping strategy parameters, and jointly encoding them to form a joint input vector; S72. Input the joint input vector into the physical neural network model, wherein the physical modeling module regenerates the physical state features based on the joint input vector, and the neural network module performs attention map construction, structure cutting and strategy mapping processes based on the physical state features, and outputs the updated grasping position, grasping posture and grasping force; S73, controlling the robot to perform a grasping action according to the updated grasping position, grasping posture, and grasping force, and re-collecting contact force changes, object posture changes, and stability status; S74. Construct a stable grasping criterion function based on the re-collected data. The criterion function includes three dimensions: the end contact force change rate, the object posture angle offset value, and the stability classification state, and is used to determine whether the current grasping state meets the stability condition. S75. If all output values ​​of the stable grasping criterion function meet the preset threshold conditions, the final grasping position, grasping posture and grasping force are output. If not, repeat steps S71 to S74 until the grasping strategy parameters that meet the stability conditions are generated.

9. A robot target grasping system based on a physical neural network, executing the robot target grasping method based on a physical neural network according to any one of claims 1 to 8, characterized in that: include: The sensory data acquisition module is used to collect visual data, tactile data, and force data of the target object and the environment, perform filtering, normalization, and feature extraction operations, and generate fused sensory data; The physical modeling module is used to build a force and motion state modeling mechanism in three-dimensional space based on the fused perception data, generate physical state features, and calculate the modeling error during the model update process; The neural network module adopts an improved SE(3)-Transformer structure, which takes as input the fusion perception data and physical state features, generates a combined tensor structure through a multimodal embedding structure, an attention weight generation mechanism, and a reconfigurable control unit, and outputs the grasping position, grasping posture, and grasping force through the decoder branch. The grasping execution module is used to control the robot to execute grasping actions according to the grasping strategy parameters, collect contact force changes, object posture changes and stability status, and construct feedback data; The feedback encoding module is used to normalize and structure the contact force changes, object posture changes, and stability states, generate a feedback state vector set, and bind it with the grasping strategy parameters to form a joint input vector; The strategy iteration update module is used to input the joint input vector into the physical neural network model, regenerate the physical state characteristics and grasping strategy parameters, and judge whether the grasping state meets the preset conditions based on the stable grasping criterion function. If not, it will continue to iterate until a stable grasp is achieved; The joint training module is used to construct a joint loss function that includes the policy error term and the modeling error term, perform backpropagation operations using the error vector, and optimize the parameter configuration of the physical modeling module and the neural network module.