Surface displacement-force conversion method and device based on graph neural network
Through the surface displacement-force conversion method based on graph neural network, the problems of poor generalization of force perception and high computational cost in the prior art are solved, and high precision and low computational force estimation is achieved, which is suitable for complex scenarios.
Patent Information
- Application Number
- CN202510144371.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-24
AI Technical Summary
When realizing force perception, the prior art is difficult to maintain high generalization in unknown scenarios, and the modeling is complex and the calculation cost is high.
The surface displacement-force transformation method based on graph neural network is adopted, and the displacement-force transformation is achieved through finite element segmentation and graph establishment, combined with graph attention sub-network and full-connected sub-network.
It realizes high-precision, low computational volume and high generalization force estimation, and is suitable for complex scenarios such as visual haptic sensors and surgical robots.
Smart Images

Figure CN120199377A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optical sensors, and particularly relates to a surface displacement-force conversion method and device based on a graph neural network. Background Art
[0002] In the field of robot operation, real-time estimation of contact force is the basis of robot operation. Precise force perception enables robots to perform complex operations. However, realizing force perception often relies on complex hardware devices, resulting in high costs. Existing technologies collect images through high-spatial-resolution camera devices, and displacement fields and force fields can be obtained from pictures of the contact surface. However, due to the nonlinearity and heterogeneity of elastic materials on the contact surface, it is still challenging to achieve accurate and generalizable force estimation methods through displacement fields.
[0003] Finite element analysis and other numerical calculation methods realize force reconstruction by constructing physical models of materials, which are currently relatively common displacement-force conversion methods. These methods model the relationship between displacement fields and force fields and have the advantage of high generalization in the same application scenario, but require complex modeling and a large amount of computational cost. Considering the powerful nonlinear fitting ability of neural networks, some methods train networks based on a large amount of displacement-force paired data, and this method has high accuracy on the training data. However, due to the lack of physical constraints in these methods and their heavy dependence on the scale and diversity of the training dataset, the generalization in unknown scenarios is limited. Summary of the Invention
[0004] The object of the present invention is to overcome the deficiencies of the existing technologies and propose a surface displacement-force conversion method and device based on a graph neural network. The present invention has the advantages of high precision, low computational amount and high generalization, and has high application value.
[0005] The first aspect of the present invention provides a surface displacement-force conversion method based on a graph neural network, including:
[0006] Before applying a surface force to an object to be contacted with an elastomeric surface, perform finite element segmentation on the elastomeric surface, and based on the segmentation result, establish a graph corresponding to the elastomeric surface composed of finite element nodes and edges of the finite elements;
[0007] Apply a surface force to the object to be contacted, and obtain a three-dimensional feature vector of each finite element node, where the three-dimensional feature vector includes the displacement feature and the position feature of the finite element node;
[0008] Input the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, and in combination with the edge features included in the graph, the network outputs a predicted value of the three-dimensional force of each finite element node; wherein, the displacement-force conversion network uses a graph neural network;
[0009] Calculate the resultant force of the predicted values of the three-dimensional forces of each finite element node to obtain the predicted result of the surface force.
[0010] In a specific embodiment of the present invention, the obtaining of the three-dimensional feature vector of each finite element node includes:
[0011] 1) Set surface landmark points on the surface of the elastomer;
[0012] 2) After applying a surface force to the contacted object, obtain the positions and displacements of the surface landmark points;
[0013] 3) Convert the displacements of the surface landmark points into the displacements of the finite element nodes, including: the vertical displacement and the horizontal displacement of the finite element nodes;
[0014] Wherein, in the non-contact state, the coordinate field of the i-th finite element node is denoted as s i =(x i , y i , z i ) T , where s i is the coordinate field of the i-th finite element node, and x i , y i , z i are the coordinates of the i-th finite element node in the x, y, and z directions respectively; the coordinate field of the i-th surface landmark point is denoted as s mi =(x mi , y mi , z mi ) T , where m i represents the i-th surface landmark point, and s mi is the coordinate field of the i-th surface landmark point, and x mi , y mi , z mi are the coordinates of the i-th surface landmark point in the x, y, and z directions respectively; the displacement of the i-th surface landmark point is denoted as γ i =(u mi , v mi , w mi ) T , where γ i is the displacement field of the i-th surface landmark point, and u mi , v mi , w mi are the displacements of the i-th landmark point in the x, y, and z directions respectively;
[0015] Use photometric stereo method to calculate the vertical displacement w i of the i-th finite element node;
[0016] The horizontal displacement (u mi , v mi ) of the surface landmark points is obtained by using the marker tracking method, and then the horizontal displacement (u i , v i ) of the i-th finite element node is calculated by bilinear interpolation. For the i-th finite element node, the four surface landmark points closest to the node are denoted as mi0, mi1, mi2, and mi3 respectively, and then the interpolation parameters λ i and μ i of the i-th finite element node are obtained by solving the following equations:
[0017]
[0018] Then
[0019]
[0020] where M k and N k are the projections of the cross product of the relevant vectors d mik in the vertical direction obtained from the surface landmark point coordinates x mik , y i , y i and the finite element node coordinates x k . Here, k represents the serial numbers 0, 1, 2, 3, corresponding to the four surface landmark points respectively. The update calculation processes of d k , M k and N k are as follows:
[0021]
[0022]
[0023] d k = (x mik , y mik ) T - (x i , y i ) T
[0024] Then the horizontal displacement of the i-th finite element node is:
[0025]
[0026] 4) Based on the displacements of the finite element nodes and combined with the positions of the finite element nodes, the three-dimensional feature vectors of the finite element nodes are obtained;
[0027] Based on the result of step 3), the displacement position characteristic variable δ of the i-th finite element node is obtained: i =(u i ,v i ,w i ,x i ,y i ,z i ) T , is a six-dimensional vector; then the encoding function transforms δ i Encoded as a three-dimensional feature vector in represents the initial characteristics of the i-th finite element node, a i ,b i ,c i Represents the three-dimensional eigenvector of the i-th finite element node;
[0028]
[0029] in, Represents the encoding function.
[0030] In a specific embodiment of the present invention, the encoding function is to regularize the three-dimensional displacement and three-dimensional position of the finite element node and then superimpose them on each other.
[0031] In a specific embodiment of the present invention, the displacement-force conversion network is composed of a graph attention subnetwork and a fully connected subnetwork connected in sequence, wherein the graph attention subnetwork includes two layers of graph attention layers connected in sequence, and the fully connected subnetwork includes four layers of fully connected layers connected in sequence.
[0032] In a specific embodiment of the present invention, before inputting the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, the method further includes:
[0033] training the displacement-force conversion network;
[0034] The training of the displacement-force conversion network comprises:
[0035] 1) applying surface force to the contacted object in different scenarios, obtaining the three-dimensional feature vector of each finite element node after each contact, and collecting the three-dimensional force of the contact as the true force label to construct a data set;
[0036] 2) Divide the data set obtained in step 1) into a training set, a test set, and a validation set;
[0037] 3) constructing the displacement-force conversion network;
[0038] 4) Use the training set to train the displacement-force conversion network. During the training process, use the validation set to check the error, and use the test set to evaluate the performance after the training ends. When the error and the performance reach the preset conditions for ending the training, the training is completed, and the final displacement-force conversion network is obtained.
[0039] In a specific embodiment of the present invention, it further includes:
[0040] After the three-dimensional feature vectors of each finite element node are input into a preset displacement-force conversion network, the network considers the edge features included in the graph features, that is, the relationship of pairwise node connections, and then constructs the adjacency matrix H of the graph adj , the adjacency matrix is composed of 0 and 1, 1 indicates that two nodes are on the same finite element unit, and 0 indicates that two nodes are not on the same finite element unit; then input the adjacency matrix and the three-dimensional feature vectors of each node into the graph attention sub-network.
[0041] In a specific embodiment of the present invention, it further includes:
[0042] The graph attention sub-network adopts a self-attention mechanism, and the update process is as follows:
[0043]
[0044] α ij = softmax(LeakyReLU(aT[Wh i ||Wh j ))
[0045] where i and j are node numbers, t is the number of steps for node feature update, and t takes 0, 1, 2; a and W are learnable weight matrices, and α ij represents the weight relationship between nodes i and j; represents the updated feature of each node at the t-th step, N i is the set of adjacent nodes of the i-th finite element node; MLP θ represents a fully connected sub-network with θ layers;
[0046] The expression of the loss function during the training of the displacement-force conversion network is as follows:
[0047]
[0048] where, represents the predicted value of the output force of the i-th finite element node, n node represents the number of nodes; f total represents the true value of the cohesive force, that is, the true force label.
[0049] The second aspect of the embodiments of the present invention provides a surface displacement-force conversion device based on a graph neural network, including:
[0050] A finite element segmentation module, configured to perform finite element segmentation on the surface of the elastomer before applying a surface force to the object to be contacted with the elastomer surface, and based on the segmentation result, establish a graph corresponding to the elastomer surface composed of finite element nodes and edges of the finite element;
[0051] A node feature acquisition module, configured to apply a surface force to the object to be contacted and acquire a three-dimensional feature vector of each finite element node, where the three-dimensional feature vector includes a displacement feature and a position feature of the finite element node;
[0052] A node force prediction module, configured to input the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, and in combination with the edge features included in the graph, the network outputs a predicted value of the three-dimensional force of each finite element node; wherein, the displacement-force conversion network uses a graph neural network;
[0053] A surface force prediction module, configured to calculate the resultant force of the predicted values of the three-dimensional forces of each finite element node to obtain a predicted result of the surface force.
[0054] The third aspect of the embodiments of the present invention provides an electronic device, including:
[0055] At least one processor; and a memory communicatively connected to the at least one processor;
[0056] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the above-mentioned surface displacement-force conversion method based on a graph neural network.
[0057] The fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the above-mentioned surface displacement-force conversion method based on a graph neural network.
[0058] The features and beneficial effects of the present invention are as follows:
[0059] The present invention can be applied to fields such as visual tactile sensors and surgical robots where it is difficult to add force sensors for collaborative control, etc., to solve the problems of poor generalization of existing networks, difficult modeling of existing models, and large computational complexity. For the finite element mesh of the elastomer, the present invention uses node and edge mapping to construct a graph. The architecture based on the attention mechanism allows different weights to be implicitly assigned to nodes within the neighborhood, effectively dealing with the heterogeneous effects caused by Young's modulus and Poisson's ratio in different parts of the elastomer. By combining physical priors with neural networks, it can exhibit strong generalization ability on unfamiliar data. In the field of visual tactile sensors, since a high-resolution image is used as the input, its displacement field can be easily obtained, and most sensors have a force sensing requirement. Therefore, the present invention realizes a general force measurement mode for visual tactile sensors. In the field of surgical robots, it is difficult to model soft tissues such as organs and materials with complex structures, and it is also difficult to directly apply force sensors to surgical knives. The present invention can provide an indirect real-time force measurement method for the field of surgical robots, facilitating complex surgical robot operations. Description of the Drawings
[0060] Figure 1 is the overall flowchart of a surface displacement-force conversion method based on a graph neural network according to an embodiment of the present invention.
[0061] Figure 2 is a schematic diagram of a data acquisition device according to a specific embodiment of the present invention;
[0062] In the figure, 1 is a three-coordinate translation stage, 2 is a six-axis force sensor, 3 is an indenter of different shapes, and 4 is a visual tactile sensor.
[0063] Figure 3 is a schematic diagram of different indenters according to a specific embodiment of the present invention.
[0064] Figure 4 is the working principle diagram of a displacement-force conversion network according to a specific embodiment of the present invention. Detailed Embodiments
[0065] The present invention proposes a surface displacement-force conversion method and device based on a graph neural network. The following further details the present invention in conjunction with the drawings and specific embodiments. The following embodiments are used to illustrate the present invention, but are not limited to the scope of the present invention.
[0066] In a first aspect embodiment of the present invention, a surface displacement-force conversion method based on a graph neural network is proposed. The overall process is as Figure 1 shown, including:
[0067] Before applying a surface force to an object with an elastomeric surface, perform finite element segmentation on the elastomeric surface. Based on the segmentation results, construct a graph corresponding to the elastomeric surface composed of finite element nodes and the edges of the finite elements.
[0068] Apply a surface force to the object in contact, and obtain the three-dimensional feature vectors of each finite element node. The three-dimensional feature vectors include the displacement features and position features of the finite element nodes.
[0069] Input the three-dimensional feature vectors of each finite element node into a preset displacement-force conversion network. Combining the edge features included in the graph, the network outputs the predicted values of the three-dimensional forces of each finite element node; wherein, the displacement-force conversion network uses a graph neural network.
[0070] Calculate the resultant force of the predicted values of the three-dimensional forces of each finite element node to obtain the predicted result of the surface force.
[0071] In a specific embodiment of the present invention, the surface displacement-force conversion method based on a graph neural network includes the following steps:
[0072] 1) Training stage.
[0073] 1-1) Construct a data set.
[0074] 1-1-1) Finite element segmentation.
[0075] This embodiment can be applied to an object with an elastomeric contact surface to estimate the force received by the object. In a specific embodiment of the present invention, the object in contact is a visual tactile sensor, and its surface has a layer of silicone. For the elastomeric surface of an object, it is often difficult to obtain the true force distribution on its surface. In this embodiment, the elastomeric surface of the object is modeled as a graph through finite element, where the finite element nodes correspond to the nodes of the graph, and the edges of the finite element cells correspond to the edges of the graph. Specifically, for the elastomeric surface of the object, first perform finite element segmentation on the entire elastomeric part. In this embodiment, it is assumed that the part outside the elastomer is fixed with a rigid body or is a free surface, and corresponding displacement boundary conditions are given during the finite element segmentation process. Based on the segmented elastomeric surface, construct a graph G=(V, E), where V represents the finite element nodes and E represents the edges of the finite elements.
[0076] 1-1-2) Based on the finite element segmentation results in step 1-1-1), construct a data set.
[0077] In this embodiment, the data set includes: a training set, a validation set, and a test set. A single sample in this data set uses the displacement position feature of each finite element node as the input and the force corresponding to this node as the output. During the training process, the true value of the force needs to be used as a label, so it is necessary to collect force-displacement data pairs. Generally, it is necessary to create different displacement fields through different contacts, and then collect the true values during contact through a force sensor connected to the elastomer.
[0078] Here, a specific embodiment of data acquisition of the present invention is given. The following process can be referred to during the data acquisition process of different visual and tactile sensors:
[0079] 1-1-2-1) Build a data acquisition device.
[0080] The structure of the data acquisition device in a specific embodiment of the present invention is as Figure 2 shown. Figure 2 As shown in the figure, the tactile sensor 4 is fixed on the optical platform, and the surface of the tactile sensor 4 has an elastomer surface made of silica gel; the indenter 3 is installed on a platform 1 with 4 degrees of freedom (the 4 degrees of freedom include 3 translational degrees of freedom and 1 rotational degree of freedom), and a 6-axis force sensor 2 (in a specific embodiment of the present invention, SRI M3813A is used) is installed between the platform 1 and the indenter 3. The platform 1 is fixed on the optical platform by bolts, so that the indenter 3 is directly facing the effective detection area of the visual and tactile sensor 4.
[0081] 1-1-2-2) Based on the device built in step 1-1-2-1), use the indenter to press the silica gel layer on the surface of the visual and tactile sensor, and then collect the contact displacement field of the preset surface landmark points of the visual and tactile sensor and the corresponding true force label during each press.
[0082] In this embodiment, during each acquisition, the platform 1 controls the indenter 3 to approach the visual and tactile sensor 4 and press on the silica gel layer on the surface of the visual and tactile sensor 4 to produce a certain deformation. At this time, the three-dimensional force data on the force sensor 2 is recorded, and the image sensor built into the visual and tactile sensor 4 captures the image at this time and obtains the displacement of the surface landmark points preset on the visual and tactile sensor 4 through a reconstruction algorithm, thereby obtaining the contact displacement field corresponding to this press (including the three-dimensional position and displacement of the surface landmark points of the visual and tactile sensor) and the corresponding true force label to form a data sample. The reconstruction algorithm part here can be seen in the patent "A Visual and Tactile Sensor Based on Lensless Imaging and Its Measurement Method" (Patent No.: CN202410567802.4).
[0083] Furthermore, for the elastomer disposed on the surface of the visual-tactile sensor, after being pressed, the deformation of the contact surface can be collected by using a 3D scanner or a depth camera externally to obtain the displacement of the surface landmark points. At the same time, the three-dimensional force collected by the force sensor corresponding to this press is obtained as the true force to form a set of data pairs. In this embodiment, 1000 sets of data pairs are collected under a single indenter, and the data pairs with poor pressing quality are filtered. Finally, about 700 sets of data pairs are obtained under each indenter.
[0084] Furthermore, in order to create different contact scenarios, a specific embodiment of the present invention designs 8 different-shaped indenters, such as Figure 3 shown, among which 6 indenters are used to construct the training set, 1 indenter is used to construct the validation set, and the remaining 1 indenter is used to construct the test set. For each indenter, the translation size and rotation direction are randomly set in this embodiment, the indentation depth is set between 0 and 1 mm, and about 700 sets of data pairs composed of the three-dimensional positions and displacements of the surface landmark points of the visual-tactile sensor and the corresponding true force labels are collected under each indenter.
[0085] 1-1-2-3) Convert the displacement of the surface landmark points in the data pairs collected in step 1-1-2-2) into the displacement of the finite element nodes.
[0086] In this embodiment, after obtaining the data samples composed of the contact displacement field of the surface landmark points and the corresponding force labels, in order to obtain the initial features required for the data set, that is, the displacement field of the finite element nodes, it is necessary to convert the discrete displacement of the landmark points into a denser surface node displacement. In a specific embodiment of the present invention, when not in contact, the coordinate field of the i-th finite element node of the visual-tactile sensor is denoted as s i =(x i , y i , z i ) T , where s i is the coordinate field of the i-th finite element node, and x i , y i , z i are the coordinates of the i-th finite element node in the x, y, and z directions respectively. The coordinate field of the i-th surface landmark point of the visual-tactile sensor is denoted as s mi =(x mi , y mi , z mi ) T , where, m i represents the i-th surface landmark point, s mi is the coordinate field of the i-th surface landmark point, and x mi , y mi , z miThey are the coordinates of the i-th surface landmark point in the x, y, and z directions respectively. Denote the displacement of the i-th surface landmark point as γ i =(u mi , v mi , w mi ), T where γ i is the displacement field of the i-th surface landmark point, and u mi , v mi , w mi are the displacements of the i-th landmark point in the x, y, and z directions respectively.
[0087] Calculate the vertical displacement of the finite element node. In this embodiment, the photometric stereo method is used to calculate the vertical displacement w i of the i-th finite element node. The photometric stereo method converts the surface gradient into color intensity, and then calculates the depth using the Poisson equation with the color intensity. The mapping from color intensity to depth gradient can be realized by the look-up table method or neural network. In the specific embodiment of the present invention, the look-up table method is adopted to calculate the depth deformation.
[0088] Calculate the horizontal displacement of the finite element node. In this embodiment, first use the marker tracking method to obtain the horizontal displacement (u mi , v mi ) of the surface landmark point, where m i represents the i-th surface landmark point, u mi represents the displacement of the i-th landmark point in the x direction, i.e., the horizontal horizontal axis direction, and v mi represents the displacement of the i-th surface landmark point in the y direction, i.e., the horizontal vertical axis direction. Then calculate the horizontal displacement (u i , v i ) of the i-th finite element node by bilinear interpolation, where u i represents the displacement of the i-th finite element node in the x direction, i.e., the horizontal horizontal axis direction, and v i represents the displacement of the i-th finite element node in the y direction, i.e., the horizontal vertical axis direction. The specific implementation process is as follows: for the i-th finite element node, first find the four surface landmark points closest to this node, denoted as mi0, mi1, mi2, mi3 respectively, and then obtain the interpolation parameters λ i , μ i by solving the following equations:[[]]
[0089]
[0090] Furthermore, it can be deduced that:[[]]
[0091]
[0092] where, M kand N k is the projection of the cross product in the vertical direction obtained from the surface landmark point coordinates x mik , y mik and the finite element node coordinates x i , y i to obtain the relevant vector d k where k represents the sequence numbers 0, 1, 2, 3, corresponding to four surface landmark points respectively. Here, d k , M k and N k are updated as intermediate parameters with the change of nodes, and the calculation process is as follows:
[0093]
[0094] d k = (x mik , y mik ) T - (x i , y i ) T
[0095] Then, calculate the horizontal displacement of the i-th finite element node:
[0096]
[0097] 1 - 1 - 2 - 4) Based on the displacements of the finite element nodes and combined with the positions of the finite element nodes, obtain the three-dimensional feature vectors of the finite element nodes to construct a data set.
[0098] In this embodiment, based on the results of step 1 - 1 - 2 - 3), obtain the information of six dimensions of each finite element node after each contact of the indenter, δ i = (u i , v i , w i , x i , y i , z i ) T , that is, the displacement position feature. δ i represents the displacement position feature variable of the i-th finite element node, which is a six-dimensional vector. The first three dimensions represent the three-dimensional displacement of the node, and the last three dimensions represent the position of the node. In this embodiment, the node displacement feature variable δ i is encoded into a three-dimensional feature vector as the input feature of the subsequent network where represents the initial feature of the i-th finite element node, and a i , b i , c i represent the three-dimensional feature vectors of the i-th finite element node.
[0099]
[0100] wherein represents an encoding function, and in this embodiment, the encoding function is to superpose the displacement field (i.e., the three-dimensional displacement of the nodes) and the position field (i.e., the three-dimensional position of the nodes) after regularization.
[0101] Finally, the three-dimensional feature vector and the true force label of the finite element nodes corresponding to each press form a sample, and all the samples form a data set. In this embodiment, the samples generated by the 6 types of indenter used to construct the training set form the training set, the samples generated by the indenter used to construct the validation set form the validation set, and the samples generated by the indenter used to construct the test set form the test set.
[0102] 1-2) Construct a displacement-force conversion network.
[0103] In this embodiment, the displacement-force conversion network can reflect the conversion relationship from the displacement field to the force field. This network is composed of a graph attention sub-network and a fully connected sub-network connected in sequence, wherein the graph attention sub-network includes two layers of graph attention layers connected in sequence, and the fully connected sub-network includes four layers of fully connected layers connected in sequence. The displacement-force conversion network of this embodiment inputs the three-dimensional feature vector of the finite element nodes into the graph attention sub-network, and then through the fully connected sub-network, finally outputs the predicted force field.
[0104] Specifically, for the graph constructed in step 1-1), this embodiment constructs a graph neural network (Graph Neural Network, GNN) as the displacement-force conversion network, and updates the finite element node features and predicts the force of each finite element node according to the label of the true force in the data set obtained in step 1-1). Figure 4 is the working principle diagram of the displacement-force conversion network of a specific embodiment of the present invention. As Figure 4 shown, the displacement-force conversion network constructed in this embodiment takes the three-dimensional displacement of the finite element nodes as the input. This network considers the features of the graph established in step 1), wherein the edge features included in the features of the graph, that is, the relationship of pairwise node connections, are used to construct the adjacency matrix H of the graph adj . This adjacency matrix is a matrix composed of 0 and 1. 1 indicates that two nodes are on the same finite element unit, and 0 indicates that two nodes are not on the same finite element unit. The edge features of the graph can be encoded through the adjacency matrix. The node features of the nodes are also included in the graph features. Each node feature (including: displacement, position, node belonging area attribute such as each node belonging to different part labels of the object, and feature vectors that can be added for specific tasks. In this embodiment, displacement and position are used as specific feature vectors) is converted into a three-dimensional feature vector through the encoding function i = 1, 2, …, n node , where i represents the node number, and n node represents the number of nodes. In this embodiment, the adjacency matrix and the three-dimensional feature vector of each node are used as parameters and input into the graph attention sub-network. During the feature update process of the graph attention sub-network, a self-attention mechanism is introduced to learn the relationships between different adjacent nodes. It should be noted that the graph attention sub-network in this embodiment adopts a graph attention mechanism network (Graph attention network), and this graph attention sub-network includes two layers of graph attention layers in total.
[0105] Among them, in the first graph attention layer, the trainable parameters are the trainable parameters α between pairwise nodes ij and the weight parameter W. Each node obtains the surrounding node information according to the adjacency matrix, and then trains the weight of each node for the central node respectively. Due to the existence of the multi-head training mechanism, the same operation is repeated n hesd times, and n head represents the number of heads of the multi-head. Finally, the updated node features of this layer are obtained. At this time, the node feature dimension should be the feature dimension n of the hidden layer hid .
[0106] Next, the feature vectors of different heads are concatenated together to obtain the input vector returned to the second graph attention layer of the graph attention sub-network, and then a similar operation of aggregating node features is performed, and the output is a hidden feature dimension.
[0107] Next, the output of the second graph attention layer is passed through a fully connected sub-network (MLP) to obtain the output 3D features, and then the three-dimensional resultant force and the label force in the same direction are compared, and finally the operation of backpropagation is performed.
[0108] After reaching the training termination condition, the displacement-force conversion network of this embodiment outputs the three-dimensional force vector of the i-th finite element node where T represents the number of iteration steps. represents the final feature of the i-th finite element node, which is a three-dimensional vector, represents the predicted value of the output force of the i-th finite element node, respectively represent the predicted values of the output forces of the i-th node in the x, y, and z directions.
[0109] 1 - 3) Use the training set to train the network constructed in step 1 - 2).
[0110] In this embodiment, a self-attention mechanism is incorporated into the displacement-force conversion network, which can reflect the different influences of adjacent nodes on the central node. For the initial node features in this embodiment Feature updates are performed in the graph attention subnetwork. The specific process is as follows:
[0111]
[0112] α ij = softmax(LeakyReLU(a T [Wh i ||Wh j ))
[0113] where i and j are node numbers, t is the number of steps for node feature update. In this embodiment, the graph attention subnetwork includes two graph attention layers, that is, t takes 0, 1, 2, a and W are learnable weight matrices, and α ij represents the weight relationship between nodes i and j. Among them, the aggregation function learns the features of nodes and their neighborhoods. By establishing trainable weights between node pairs, the LeakyReLU function is used for activation and the softmax function is applied for normalization. represents the updated feature of each node at the t-th step, N i is the set of adjacent nodes of the i-th finite element node, which can be obtained from the adjacency matrix H adj Here, the adjacency matrix is transformed into an undirected one and normalized. MLP θ represents a fully connected subnetwork with θ layers. In this embodiment, θ is 4.
[0114] In this embodiment, the displacement-force conversion network calculates the predicted value of the resultant force by summing the forces of all finite element nodes on the surface of the elastic body, and then takes the true force label as the true value f total of the aggregated force, and uses the L1 loss to train the network. The expression is as follows:
[0115]
[0116] In a specific embodiment of the present invention, the code is written using the PyTorch library to implement this network. This embodiment trains the network on an RTX4080 GPU. During the training process, the error is checked on the validation set, and the performance evaluation of force reconstruction is performed on the test set after the training is completed. The network is tested on the validation set once every 10,000 iterations to obtain the error. When the error does not decrease after 4,000,000 iterations or 200 tests, it is considered that the network has converged and the training process is terminated.
[0117] Then, a series of hyperparameters of the network are adjusted according to the evaluation results. By testing the hyperparameters within a reasonable range, the best performance hyperparameters in this embodiment are as follows: the learning rate is set to 0.0001, the batch size is 4, the hidden layer dimensions of the graph attention layer and the fully connected layer are both 1024, the number of multi-head attention heads is 4, the random seed is 4, and LeakyReLU with α=0.2 is used as the activation function.
[0118] 2) Application phase.
[0119] 2-1) Obtain the three-dimensional feature vector of each finite element node on the elastic surface of the contacted object.
[0120] In this embodiment, the surface deformation of the visual tactile sensor acquired in real time can be converted from a displacement field to a force field, and the force field of the surface of the visual tactile sensor can be obtained in real time. By touching different surfaces with the visual tactile sensor, an image can be obtained by the built-in image sensor of the visual tactile sensor, and then the displacement field of the surface marker point can be obtained through the reconstruction algorithm, δ i =(u i ,v i ,w i ,x i ,y i ,z i ) T , i=1,2,…,n node As input, it is finally transformed into the three-dimensional feature vector of the finite element node.
[0121] 2-2) The three-dimensional feature vector of each finite element node obtained in step 2-1) is input into the displacement-force conversion network trained in step 1), and the network outputs the predicted value of the three-dimensional force of each finite element node i=1,2,…,n node Then, the predicted value of the resultant force is obtained by adding the predicted values of the forces in the same direction of all nodes and serving as the predicted result of the surface force of the visual tactile sensor. The method described in this embodiment can obtain the surface distribution force of the visual tactile sensor, and its distribution conforms to the laws of physics, which indicates that it has a certain degree of accuracy.
[0122] To implement the above embodiment, the second aspect of the present invention proposes a surface displacement-force conversion device based on a graph neural network, comprising:
[0123] A finite element segmentation module, used for performing finite element segmentation on the elastic surface before applying a surface force to the contacted object having the elastic surface, and establishing a graph consisting of finite element nodes and finite element edges corresponding to the elastic surface based on the segmentation result;
[0124] A node feature acquisition module, configured to apply a surface force to the object in contact, and acquire a three-dimensional feature vector of each finite element node, where the three-dimensional feature vector includes a displacement feature and a position feature of the finite element node;
[0125] A node force prediction module, configured to input the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, and combine the edge features included in the graph, and the network outputs a predicted value of the three-dimensional force of each finite element node; wherein, the displacement-force conversion network uses a graph neural network;
[0126] A surface force prediction module, configured to calculate the resultant force of the predicted values of the three-dimensional forces of each finite element node to obtain a predicted result of the surface force.
[0127] It should be noted that the foregoing explanation of the embodiments of a surface displacement-force conversion method based on a graph neural network is also applicable to a surface displacement-force conversion device based on a graph neural network in this embodiment, and will not be elaborated here. A surface displacement-force conversion device based on a graph neural network proposed in an embodiment of the present invention, before applying a surface force to an object in contact having an elastomeric surface, performs finite element segmentation on the elastomeric surface, and based on the segmentation result, establishes a graph corresponding to the elastomeric surface composed of finite element nodes and finite element edges; applies a surface force to the object in contact, and acquires a three-dimensional feature vector of each finite element node, where the three-dimensional feature vector includes a displacement feature and a position feature of the finite element node; inputs the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, and combines the edge features included in the graph, and the network outputs a predicted value of the three-dimensional force of each finite element node; wherein, the displacement-force conversion network uses a graph neural network; calculates the resultant force of the predicted values of the three-dimensional forces of each finite element node to obtain a predicted result of the surface force. Thus, it can be applied to fields such as visual tactile sensors and surgical robots where it is difficult to add force sensors for collaborative control, etc., to solve the problems of poor generalization of existing networks, difficult modeling of existing models, and large computational complexity.
[0128] To implement the above embodiments, a third aspect embodiment of the present invention proposes an electronic device, including:
[0129] At least one processor; and a memory communicatively connected to the at least one processor;
[0130] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the above-mentioned surface displacement-force conversion method based on a graph neural network.
[0131] To implement the above embodiments, an embodiment of the fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the above-described method for surface displacement-force conversion based on a graph neural network.
[0132] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0133] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the above-described method for surface displacement-force conversion based on a graph neural network in the above embodiments.
[0134] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0135] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0136] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0137] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a manner that is not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0138] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0139] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0140] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0141] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0142] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.
Claims
1. A surface displacement-force conversion method based on graph neural network, characterized in that: include: Before applying a surface force to a contacted object having an elastic surface, finite element segmentation is performed on the elastic surface, and based on the segmentation result, a graph consisting of finite element nodes and finite element edges corresponding to the elastic surface is established; Applying a surface force to the contacted object to obtain a three-dimensional feature vector of each finite element node, wherein the three-dimensional feature vector includes a displacement feature and a position feature of the finite element node; The three-dimensional feature vector of each finite element node is input into a preset displacement-force conversion network, and the network outputs a predicted value of the three-dimensional force of each finite element node in combination with the edge features contained in the graph; wherein the displacement-force conversion network adopts a graph neural network; The resultant force is calculated based on the predicted value of the three-dimensional force of each finite element node to obtain the predicted result of the surface force.
2. The method according to claim 1, characterized in that The step of obtaining the three-dimensional feature vector of each finite element node includes: 1) Setting surface marking points on the surface of the elastomer; 2) after applying a surface force to the contacted object, obtaining the position and displacement of the surface marker point; 3) converting the displacement of the surface marker point into the displacement of the finite element node, including: the vertical displacement and the horizontal displacement of the finite element node; In which, under non-contact conditions, the coordinate field of the i-th finite element node is recorded as Among them, s i is the coordinate field of the i-th finite element node, x i ,y i , z i are the coordinates of the i-th finite element node in the x, y, and z directions respectively; the coordinate field of the i-th surface marker point is recorded as Among them, m i represents the i-th surface landmark point, s mi is the coordinate field of the i-th surface landmark point, x mi ,y mi , z mi are the coordinates of the i-th surface marker point in the x, y, and z directions respectively; the displacement of the i-th surface marker point is recorded as where γ i is the displacement field of the i-th surface landmark point, u mi, v mi , w mi are the displacements of the i-th landmark point in the x, y, and z directions respectively; Calculate the vertical displacement w of the i-th finite element node using the photometric stereo method i ; The marker tracking method is used to obtain the horizontal displacement (u mi, v mi ), and then the horizontal displacement (u i , v i ); For the i-th finite element node, find the four surface landmarks closest to the node and record them as mi0, mi1, mi2, and mi3 respectively, and then obtain the interpolation parameter λ of the i-th finite element node by solving the following equation: i , μ i : but Among them, M k and N k is the coordinate x of the surface landmark point mik ,y mik And the finite element node coordinates x i ,y i The obtained correlation vector d k The projection of the cross product in the vertical direction, where k represents the serial number 0, 1, 2, 3, corresponding to the four surface landmarks respectively; where d k 、M k and N k The update calculation process is as follows: Then the horizontal displacement of the i-th finite element node is: 4) based on the displacement of the finite element node and in combination with the position of the finite element node, obtaining a three-dimensional feature vector of the finite element node; Among them, based on the result of step 3), the displacement position characteristic variable of the i-th finite element node is obtained: is a six-dimensional vector; then the encoding function transforms δ i Encoded as a three-dimensional feature vector in represents the initial characteristics of the i-th finite element node, a i , b i , c i Represents the three-dimensional eigenvector of the i-th finite element node; in, Represents the encoding function.
3. The method according to claim 2, characterized in that The encoding function is to regularize the three-dimensional displacement and three-dimensional position of the finite element node and then superimpose them on each other.
4. The method according to claim 2, characterized in that: The displacement-force conversion network consists of a graph attention subnetwork and a fully connected subnetwork connected in sequence, wherein the graph attention subnetwork includes two layers of graph attention layers connected in sequence, and the fully connected subnetwork includes four layers of fully connected layers connected in sequence.
5. The method according to claim 4, characterized in that Before inputting the three-dimensional characteristic vector of each finite element node into a preset displacement-force conversion network, the method further includes: training the displacement-force conversion network; The training of the displacement-force conversion network comprises: 1) applying surface force to the contacted object in different scenarios, obtaining the three-dimensional feature vector of each finite element node after each contact, and collecting the three-dimensional force of the contact as the true force label to construct a data set; 2) Divide the data set obtained in step 1) into a training set, a test set, and a validation set; 3) constructing the displacement-force conversion network; 4) Using the training set to train the displacement-force conversion network, using the validation set to check the error during the training process, and using the test set to evaluate the performance after the training is completed; when the error and the performance reach the preset training end conditions, the training is completed and the final displacement-force conversion network is obtained.
6. The method according to claim 5, characterized in that Also includes: After the three-dimensional feature vector of each finite element node is input into the preset displacement-force conversion network, the network considers the edge features contained in the graph features, that is, the relationship between the two nodes, and then constructs the adjacency matrix H of the graph adj , the adjacency matrix consists of 0 and 1, 1 indicates that the two nodes are on the same finite element unit, and 0 indicates that the two nodes are not on the same finite element unit; then the adjacency matrix and the three-dimensional feature vector of each node are jointly input into the graph attention subnetwork.
7. The method according to claim 6, characterized in that Also includes: The graph attention sub-network adopts the self-attention mechanism, and the update process is as follows: Among them, f, j are node numbers, t is the number of steps of node feature update, t takes 0, 1, 2; a and W are learnable weight matrices, α ij Represents the weight relationship between nodes i and j; Represents the updated features of each node at step t, N i is the set of neighboring nodes of the i-th finite element node; MLP θ represents a fully connected subnetwork of layer θ; The loss function expression during the training of the displacement-force conversion network is as follows: in, represents the predicted value of the output force of the i-th finite element node, n node represents the number of nodes; f total Represents the true value of the aggregation force, that is, the true force label.
8. A surface displacement-force conversion device based on graph neural network, characterized in that: include: A finite element segmentation module, used for performing finite element segmentation on the elastic surface before applying a surface force to the contacted object having the elastic surface, and establishing a graph consisting of finite element nodes and finite element edges corresponding to the elastic surface based on the segmentation result; A node feature acquisition module, used to apply surface force to the contacted object to obtain a three-dimensional feature vector of each finite element node, wherein the three-dimensional feature vector includes a displacement feature and a position feature of the finite element node; A node force prediction module, used for inputting the three-dimensional feature vector of each finite element node into a preset displacement-force conversion network, and combining the edge features contained in the graph, the network outputs the predicted value of the three-dimensional force of each finite element node; wherein the displacement-force conversion network adopts a graph neural network; The surface force prediction module is used to calculate the resultant force of the predicted value of the three-dimensional force of each finite element node to obtain the predicted result of the surface force.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Visual tactile sensor based on lensless imaging and measuring method thereof
CN118482656A