Intelligent Robot Safe Operation Method Based on 3D Gaussian and Tactile Image Fusion
By combining the three-dimensional Gaussian and tactile image fusion method, the problem of insufficient control stability of the gripper in traditional robot operations is solved, and the reliability and safety of the robotic arm in complex tasks is improved.
Patent Information
- Application Number
- CN202510391768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Traditional robot operation methods rely on single visual information, resulting in insufficient stability of the holder control, making it difficult to adapt to complex tasks and safe operation requirements.
Combining the method of fusion of three-dimensional Gaussian and tactile image, high-dimensional features are extracted through the RGB image sequence of the hand viewing angle of the robot arm and the third viewing angle, combined with the force sensor to obtain the contact force distribution of the gripper, the LSTM strategy head is used for action prediction, and a dynamic feedback optimization control algorithm for the safe area is introduced to ensure the stable movement and safe operation of the gripper.
It improves the reliability and efficiency of the robotic arm in complex tasks, enhances the understanding of three-dimensional space, and ensures the safety and stability of operations.
Smart Images

Figure CN119897874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and particularly relates to an intelligent robot safe operation method based on the fusion of three-dimensional Gaussian and tactile images. Background Art
[0002] Traditional robot working modes are single, with poor manipulation flexibility and lack of understanding ability. They cannot adapt to unplanned and random application scenarios and are difficult to meet the requirements of wide application and embodied intelligence.
[0003] In related technologies, with the rapid development of large-scale pre-trained models, embodied intelligence based on large models has achieved remarkable results in various tasks, demonstrating strong generalization ability and broad application prospects. An embodied intelligence system can perform complex tasks by perceiving, understanding, and interacting with the physical world, becoming an important direction in the field of artificial intelligence.
[0004] For example, the paper (Li P, Wu H, Huang Y, et al. GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal Conditioned Policy[J]. arXiv preprint arXiv:2408.14368, 2024.) proposed the GR-MG method, which combines language instructions and target images to achieve general robot operations. In the training stage, GR-MG uses trajectory data with text and images for conditional training and relies only on images when text is missing. In the inference stage, GR-MG generates a target image according to the language instruction through a diffusion-based image editing model, and then combines the text and the generated image to execute the task. To improve the accuracy of the target image, a progress-guided image generation model is designed, significantly improving the image fidelity and task execution performance. This method effectively utilizes a large amount of partially annotated data and enhances the generalization ability of the robot in different environments.
[0005] However, traditional multi-modal large model control schemes mainly rely on third-person perspective and hand perspective to extract RGB features. The single visual information leads to insufficient stability of gripper control, thus affecting the stability and efficiency of task execution. Summary of the Invention
[0006] (I) Technical Problems to be Solved
[0007] Aiming at the deficiencies of the prior art, the present invention provides an intelligent robot safe operation method based on the fusion of three-dimensional Gaussian and tactile images, solving the technical problem of insufficient stability of gripper control caused by single visual information.
[0008] (II) Technical Solutions
[0009] To achieve the above object, the present invention is realized by the following technical solutions:
[0010] An intelligent robot safe operation method based on the fusion of three-dimensional Gaussian and tactile images, wherein a gripper is installed at the end of the robotic arm of the intelligent robot; it includes:
[0011] Collect RGB image sequences from the perspective of the robotic arm's hand and the third perspective, and extract the image features of the RGB image sequences through a pre-trained ViT model;
[0012] Based on the pre-saved query vector and the weight parameters of cross-attention, perform cross-attention calculation on the image features to obtain high-dimensional features based on three-dimensional Gaussian;
[0013] The force sensor is used to sense the contact force distribution between the left and right jaws of the gripper and the object to be operated in real time, obtain a two-dimensional RGB tactile image sequence, and extract the tactile image features of the RGB tactile image sequence through the ViT model;
[0014] Use the high-dimensional features and the tactile image features as the inputs of the resampler respectively, and splice them after obtaining the corresponding feature sequences;
[0015] Use the splicing result and the language instruction as the inputs of the feature fusion decoder, and adopt the LSTM strategy head to perform action prediction on the output of the feature fusion decoder to obtain the pose change amount of the gripper and the next state of the gripper;
[0016] Based on the current pose of the gripper and the pose change amount, calculate the target pose to determine the moving direction and distance of the gripper.
[0017] Preferably, in the process of calculating the target pose, a dynamic feedback optimization control algorithm based on a safe area is introduced, including:
[0018] (1) Design a pose controller:
[0019] Based on the pose change amount, define the first position error between the target position and the current position:
[0020]
[0021] Among them, e1 is the first position error; is the pose change amount; is the target pose of the gripper to be calculated at the target position; is the current pose of the gripper at the current position;
[0022] A pose controller is used to calculate the control input to ensure that the gripper moves smoothly towards the target position:
[0023]
[0024] where u1 is the output of the pose controller that selects a PD controller; k p1 and k d1 are the first proportional and first derivative gains respectively; is the change rate of the first position error;
[0025] (2) Design a safety area controller:
[0026] Set the safety area as a planar area with a height of h. Outside this planar area, the free movement of the entire robotic arm is not restricted. Once the gripper approaches the area boundary within this planar area, the safety area controller outputs a control quantity away from the area boundary:
[0027]
[0028] where u2 is the output of the safety area controller; k p2 and k d2 are the second proportional and second derivative gains respectively; e2 is the second position error between the current position and the safety area, and the component of e2 in the Z direction , d z is the distance from the end of the gripper to the area boundary; is the change rate of the second position error;
[0029] (3) Specify the evaluation function of the control target:
[0030] Based on the square of the difference between the target pose and the current pose, design the first evaluation function of the pose controller:
[0031]
[0032] where t0 is the initial time point; T is the preset time span; ‖∙‖ is the norm of the vector; t is the time variable;
[0033] Based on the square of the difference between the distance from the end of the gripper to the area boundary and the height of the safety area, design the second evaluation function of the safety area controller:
[0034]
[0035] (4) Calculate the predicted evaluation gradient descent value and construct a state feedback mechanism:
[0036] Design a saturation suppression function :
[0037]
[0038] Among them, is the evaluation gradient, and i takes 1 or 2; α is the upper limit of the gradient amplitude;
[0039] The gradient descent method is used to obtain the input after single-objective optimization:
[0040]
[0041] Among them, u i is the basic control input; λ is the learning rate; K α is the gain coefficient corresponding to the saturation suppression function ;
[0042] The state feedback mechanism is constructed to enable the robot state to converge quickly. The single-objective controller after adding state feedback is expressed as:
[0043]
[0044] Among them, K fb is the state feedback gain matrix, and the subscript fb represents state feedback, is the target state, and x i is the current actual state;
[0045] The comprehensive output is:
[0046]
[0047] (5) Calculate the new safe target pose to ensure that the gripper avoids touching the boundary of the area:
[0048]
[0049] Among them, is the updated target pose.
[0050] Preferably, the process of obtaining the query vector and the weight parameters of the cross-attention includes:
[0051] Select a fixed world coordinate system within the manipulator operation scenario;
[0052] Use P three-dimensional Gaussian distributions to represent the manipulator operation scenario. Each three-dimensional Gaussian is represented by a d-dimensional vector in the form of , where d = 10 + |C sem | + 2, and m, s, r, and c are the mean, scale, rotation vector, and semantic category respectively, is the set of real numbers, and C sem is the number of semantic labels;
[0053] Initialization and learnable vectors ; where D is the dimension of the learnable vectors, and D >> d, used to represent the parameter attributes of the Gaussian, Q j g is a learnable three-dimensional spatial feature, with the superscript g representing the query vector related to the Gaussian representation, and the subscript j representing the index of the Gaussian kernel;
[0054] Voxelization is performed according to the initialized Gaussian kernel, and the occupancy prediction result of each voxel grid is calculated by the following formula:
[0055]
[0056] where, is the sum of the contributions of each Gaussian distribution to any position p in the selected world coordinate system, used to characterize the occupancy prediction result; is the three-dimensional Gaussian occupancy prediction distribution at point p; is the set of neighboring Gaussians of point p; exp is the exponential function; Σ is the covariance matrix; S = diag(s) represents the scale matrix, diag is the function to extract diagonal elements, and R = q2r(q) is the function to convert the quaternion q used to represent the three-dimensional space rotation into a rotation matrix;
[0057] If the occupancy prediction result of any voxel grid is greater than a certain threshold thresholds, it is considered that the voxel grid is occupied, that is, the position v of any voxel grid is expressed as:
[0058]
[0059] The set of occupied voxel grid positions is expressed as:
[0060]
[0061] where, |∙| is the size of the set; M is the number of voxel grid positions used;
[0062] The features corresponding to the occupied voxel grid positions are represented as a sparse tensor:
[0063]
[0064] where, F(v) is the sparse tensor of any voxel grid position v; F S is the set of sparse tensors corresponding to the set S;
[0065] P Gaussian kernels respectively correspond to P specific voxel grid positions {v1, v2,..., v pFor} ∈ S, define the set of reserved voxels P S as follows:
[0066]
[0067] For each voxel grid position v S in P j apply a sparse convolution kernel to obtain the output feature F'(v):
[0068]
[0069] where is the set of relative positions of the convolution kernel; W k is the convolution weight corresponding to the relative position k; is the relative position offset; σ is the non-linear activation function ReLU;
[0070] Collect RGB images from N perspectives in the robotic arm operation scenario. For each perspective's RGB image, obtain the corresponding image feature through encoding by the ViT model:
[0071]
[0072] where is the image feature of the nth perspective; I n is the RGB image of the nth perspective;
[0073] Perform cross-attention mechanism operation on the initialized query vector Query g and to obtain:
[0074]
[0075] where cross_attn is cross-attention;
[0076] Process the query vector Query g through a set of multi-layer perceptrons MLP to obtain the updated Gaussian kernel parameters :
[0077]
[0078] where are the updated mean, scale, rotation vector, and semantic category respectively;
[0079] Directly replace the original Gaussian kernel with the updated Gaussian kernel, and the new Gaussian kernel parameters are expressed as:
[0080]
[0081] Embed it into the target voxel grid of size X×Y×Z; For each new three-dimensional Gaussian, calculate the radius of its neighborhood according to its scale attribute, and recalculate the occupancy prediction result of each voxel grid;
[0082] For each re-determined occupied voxel grid, calculate the corresponding label C of its semantic category
[0083] The cross-entropy loss CE between label and the predicted value is as follows:
[0084]
[0085] where is the loss function; represents the set of occupied voxel grids; the subscript label represents the label;
[0086] Update the model parameters through the loss function, and save the query vector Query g and the weight parameters of the cross-attention Cross-Attn after training.
[0087] Preferably, the image features of the RGB image sequence are represented as:
[0088]
[0089] where is the image feature at time t, and the superscript represents vision; are the RGB images of the robotic arm hand view and the third view at time t respectively; ⊕ is the element-wise addition symbol;
[0090] Perform cross-attention calculation on the image features based on the pre-saved query vector and the weight parameters of the cross-attention to obtain the high-dimensional features based on the three-dimensional Gaussian; which is represented as:
[0091]
[0092] where is the high-dimensional feature based on the three-dimensional Gaussian at time t.
[0093] Preferably, extract the haptic image features of the RGB haptic image sequence through the ViT model; which is represented as:
[0094]
[0095] where is the haptic image feature at time t; They are the RGB tactile images of the left and right jaws of the gripper at time t. The superscripts lta and rta represent left and right tactile perceptions respectively.
[0096] Preferably, taking the high-dimensional feature and the tactile image feature as the inputs of the resampler respectively, after obtaining the corresponding feature sequences, they are concatenated; including:
[0097] Input the high-dimensional feature and the tactile image feature into the K and V fully connected layers of the resampler respectively, and input a group of learnable vectors q into the Q fully connected layer of the resampler:
[0098]
[0099] Among them, Q R ∈ is the query vector corresponding to the learnable parameters of the resampler, is the number of resampler query vectors, is the size of the hidden dimension, are the linear transformation matrices of the key and the value respectively, and d g / ta represents the dimension size of the high-dimensional feature or the tactile image feature; are the transformed key and value vectors generated by the input of the high-dimensional feature respectively; are the transformed key and value vectors generated by the input of the tactile image feature respectively; are the feature sequences corresponding to the high-dimensional feature and the tactile image feature output by the resampler respectively; softmax is the activation function; the superscript T is the transpose symbol;
[0100] Concatenate the two obtained feature sequences:
[0101]
[0102] Among them, is the concatenation result at time t, and the superscript g-ta represents the fusion of the Gaussian feature and the tactile feature; concat is the concatenation operation.
[0103] Preferably, taking the concatenation result and the language instruction as the inputs of the feature fusion decoder, and using the LSTM policy head to perform action prediction on the output of the feature fusion decoder to obtain the pose change amount of the gripper and the next state of the gripper; including:
[0104] Execute the feature fusion process through the feature fusion decoder, and integrate it with the language instruction to output ; where the feature fusion process is expressed as:
[0105]
[0106] Among them, is the output of the l-th cross-attention layer at time t; Tanh is the hyperbolic tangent function; α ∈ is a learnable parameter used to control the contributions of different features during the fusion process; A is the cross- or self-attention operation; are the learnable parameters of the cross-attention layer respectively; are the outputs of the l-th and l+1-th self-attention layers at time t respectively; are the learnable parameters of the self-attention layer respectively; L is the sum of the number of alternating self-attention layers and cross-attention layers, and is an even number; N l is the number of language instructions;
[0107] Using the LSTM policy head to perform action prediction on the output X t to obtain the pose change amount of the gripper and the next state of the gripper; expressed as:
[0108]
[0109] Among them, MaxPooling is the max pooling layer, is its output; h t and h t-1 are the hidden states at times t and t-1 respectively; are the predicted pose change amount of the gripper and the next state of the gripper respectively, and the state refers to the open or closed state of the left and right jaws of the gripper.
[0110] An intelligent robot safety operating system based on the fusion of three-dimensional Gaussian and tactile images, wherein a gripper is installed at the end of the robotic arm of the intelligent robot; including:
[0111] A collection module for collecting RGB image sequences from the perspective of the robotic arm hand and the third perspective, and extracting image features of the RGB image sequences through a pre-trained ViT model;
[0112] A calculation module for performing cross-attention calculation on the image features based on a pre-saved query vector and weight parameters of cross-attention to obtain high-dimensional features based on three-dimensional Gaussian;
[0113] An extraction module for sensing the contact force distribution between the left and right jaws of the gripper and the object to be operated in real time through a force sensor, obtaining a two-dimensional RGB tactile image sequence, and extracting tactile image features of the RGB tactile image sequence through the ViT model;
[0114] A sampling module for using the high-dimensional features and the tactile image features as inputs of a resampler respectively, obtaining corresponding feature sequences and then splicing them;
[0115] A prediction module, configured to use the splicing result and the language instruction as the input of a feature fusion decoder, and adopt an LSTM policy head to perform action prediction on the output of the feature fusion decoder, so as to obtain the pose change amount of the gripper and the next state of the gripper;
[0116] A determination module, configured to calculate a target pose based on the current pose of the gripper and the pose change amount, so as to determine the moving direction and distance of the gripper.
[0117] A storage medium stores a computer program for the safe operation of an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images, wherein the computer program enables a computer to execute the safe operation method of the intelligent robot as described above.
[0118] An electronic device includes:
[0119] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the safe operation method of the intelligent robot as described above.
[0120] (III) Beneficial effects
[0121] The present invention provides a safe operation method for an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images. Compared with the prior art, the following beneficial effects are achieved:
[0122] In the present invention, first, RGB image sequences from the perspective of the robotic arm hand and the third perspective are collected to obtain high-dimensional features based on three-dimensional Gaussian; then, the force sensor is used to sense the contact force distribution between the end of the gripper and the object to be operated in real time, and a two-dimensional RGB tactile image sequence is obtained to obtain tactile image features; then, the high-dimensional features and the tactile image features are respectively used as the input of a resampler, and after obtaining the corresponding feature sequences, they are spliced, and the splicing result and the language instruction are used as the input of a feature fusion decoder, and an LSTM policy head is adopted to perform action prediction on the output of the feature fusion decoder; finally, based on the current pose of the gripper and the pose change amount, a target pose for determining the moving direction and distance of the gripper is calculated. By combining visual information and force sensor data, the grasping strategy of the gripper is adjusted in real time, the operation process is optimized, and the reliability and efficiency of the robotic arm in complex tasks are improved. Description of the drawings
[0123] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0124] Figure 1 It is a block diagram of a method for safe operation of an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images provided by an embodiment of the present invention;
[0125] Figure 2 It is a flowchart of a method for safe operation of an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images provided by an embodiment of the present invention;
[0126] Figure 3 It is a flowchart of a method for three-dimensional space expression based on three-dimensional Gaussian representation provided by an embodiment of the present invention.
[0127] Figure 4 It is a schematic diagram of a scenario for safe operation of an intelligent robot provided by an embodiment of the present invention. Detailed implementation manners
[0128] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0129] The embodiments of the present application provide a method for safe operation of an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images, which solves the technical problem of insufficient control stability of the gripper caused by single visual information.
[0130] The technical solutions in the embodiments of the present application for solving the above technical problems are generally as follows:
[0131] (1) A method for three-dimensional space expression based on three-dimensional Gaussian representation and its application in precise operation of a robotic arm:
[0132] The embodiments of the present invention propose a method for three-dimensional space expression based on three-dimensional Gaussian representation, aiming to significantly improve the ability of the large model to understand three-dimensional scenes. Through this method, the model can more accurately perceive and express the three-dimensional space, thereby realizing the precise operation and control of the robotic arm and improving the adaptability and flexibility of the robotic arm in complex environments.
[0133] (2) Use tactile images to represent the forces on the grippers and fine-tune the model pre-trained based on RGB:
[0134] In the embodiments of the present invention, force sensors are introduced between the left and right grippers of the gripper to obtain two-dimensional images of the forces on the grippers in real time, enhancing the model's perception of the gripper state, improving the gripper control accuracy, and thus enhancing the stability and success rate of task execution. By combining force sensor data with visual information, the system can adjust the grasping strategy of the gripper in real time, optimize the operation process, and improve the reliability and efficiency of the robotic arm in complex tasks.
[0135] (3) Dynamic safety control algorithm for robotic arms based on multi-modal large models:
[0136] The embodiments of the present invention propose a dynamic feedback optimization control algorithm based on a safety area, aiming to solve the safety hazards in the operation instructions of machine learning models. By adjusting the actions of the end effector of the robotic arm (the gripper is selected in the embodiments of the present invention) in real time, obstacles are automatically identified and avoided to ensure operation safety and stability, significantly reducing risks, enhancing the adaptability, flexibility, and operation accuracy of the robotic arm, and further enhancing the reliability and intelligence level of the system.
[0137] To better understand the above technical solutions, the following will describe the above technical solutions in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0138] Embodiment 1:
[0139] As Figure 1 shown, the embodiments of the present invention provide an intelligent robot safety operation method based on the fusion of three-dimensional Gaussian and tactile images. A gripper is installed at the end of the robotic arm of the intelligent robot; the method includes:
[0140] S1. Collect RGB image sequences from the perspective of the robotic arm's hand and the third perspective, and extract the image features of the RGB image sequences through a pre-trained ViT model;
[0141] S2. Perform cross-attention calculation on the image features based on the pre-stored query vectors and the weight parameters of cross-attention to obtain high-dimensional features based on three-dimensional Gaussian;
[0142] S3. Use force sensors to sense the contact force distribution between the left and right grippers of the gripper and the object to be operated in real time, obtain two-dimensional RGB tactile image sequences, and extract the tactile image features of the RGB tactile image sequences through the ViT model;
[0143] S4. Use the high-dimensional features and the tactile image features as the inputs of the resampler respectively, obtain the corresponding feature sequences and then splice them;
[0144] S5. Use the splicing result and the language instruction as the input of the feature fusion decoder, and adopt an LSTM policy head to predict the actions of the output of the feature fusion decoder, so as to obtain the pose change amount of the gripper and the next state of the gripper;
[0145] S6. Based on the current pose of the gripper and the pose change amount, calculate the target pose to determine the moving direction and distance of the gripper.
[0146] In the embodiment of the present invention, by combining visual information and force sensor data, the grasping strategy of the gripper is adjusted in real time, the operation process is optimized, and the reliability and efficiency of the robotic arm in complex tasks are improved.
[0147] In an optional implementation manner, during the process of calculating the target pose in S6, a dynamic feedback optimization control algorithm based on a safety area is further introduced. Generally speaking: define a fixed-height plane as the safety area, combine an evaluation function, gradient optimization, and a feedback mechanism for dynamically adjusting the target position to adjust the action direction and amplitude of the robot in real time, effectively avoid exceeding the safety boundary, and ensure the safety and stability of task execution.
[0148] As Figure 2 shown, Figure 2 The flowchart of the provided intelligent robot safety operation method based on the fusion of three-dimensional Gaussian and tactile images is given.
[0149] Next, each step of the above solution will be introduced in detail in combination with Figure 2 :
[0150] In step S1, collect the RGB image sequences of the hand view and the third view of the robotic arm, and extract the image features of the RGB image sequences through a pre-trained ViT model.
[0151] It should be noted that in the traditional robotic arm operation driven by large models, many large models rely on image input for training and inference. However, the precise control of the robotic arm requires the model to have a full understanding of the three-dimensional space, and it is difficult to meet this requirement with only one or two views of RGB images. Although introducing depth maps can enhance the model's spatial perception ability, it will also significantly increase the computational burden and affect the overall efficiency.
[0152] Different from this, the embodiment of the present invention proposes a new method combining three-dimensional Gaussian, which improves the ability of the robotic arm large model to extract three-dimensional space features through phased training. First, independently train the three-dimensional space feature extraction module, and then only use RGB input to learn three-dimensional features, thereby enhancing the model's environmental perception ability while maintaining computational efficiency.
[0153] Correspondingly, for the robotic arm operation scenario, RGB image sequences from the perspective of the robotic arm's hand and the third perspective are collected, and the image features of the RGB image sequences are extracted through a pre-trained ViT model; expressed as:
[0154]
[0155] Wherein, is the image feature at time t, and the superscript represents vision; are the RGB images from the perspective of the robotic arm's hand and the third perspective at time t respectively; ⊕ is the element-wise addition symbol.
[0156] In step S2, based on the pre-saved query vector and the weight parameters of cross-attention, cross-attention calculation is performed on the image features to obtain high-dimensional features based on three-dimensional Gaussian.
[0157] As described above, the new method combining three-dimensional Gaussian proposed in the embodiment of the present invention improves the ability of the robotic arm large model to extract three-dimensional space features through staged training. First, the three-dimensional space feature extraction module is trained independently. Specifically, in this step, based on the pre-saved query vector and the weight parameters of cross-attention, cross-attention calculation is performed on the image features to obtain high-dimensional features based on three-dimensional Gaussian.
[0158] Therefore, it is necessary to supplement the detailed acquisition methods of the pre-saved query vector and the weight parameters of cross-attention. The relevant steps are as follows (where steps S400~S900 are as Figure 3 shown):
[0159] S100. Select a fixed world coordinate system within the robotic arm operation scenario.
[0160] S200. Use P three-dimensional Gaussian distributions to represent the robotic arm operation scenario. Each three-dimensional Gaussian is represented by a d-dimensional vector, in the form of , where d = 10 + |C sem +2|, and m, s, r, and c are the mean, scale, rotation vector, and semantic category respectively, is the set of real numbers, and C sem is the number of semantic labels.
[0161] S300. Initialize and as learnable vectors ; where D is the dimension of the learnable vector, and D >> d, is used to represent the parameter attributes of the Gaussian, and Q j gis a learnable three-dimensional spatial feature, where the superscript g represents the query vector related to the Gaussian representation, and the subscript j represents the index of the Gaussian kernel.
[0162] S400. Perform voxelization based on the initialized Gaussian kernel, and the occupancy prediction result of each voxel grid is calculated by the following formula:
[0163]
[0164] where is the sum of the contributions of each Gaussian distribution to an arbitrary position p in the selected world coordinate system, and is used to characterize the occupancy prediction result; is the three-dimensional Gaussian occupancy prediction distribution at point p; is the set of neighboring Gaussians of point p; exp is the exponential function; Σ is the covariance matrix; S = diag(s) represents the scale matrix, diag is the function to extract diagonal elements, and R = q2r(q) is the function to convert the quaternion q used to represent the three-dimensional space rotation into a rotation matrix.
[0165] If the occupancy prediction result of any voxel grid is greater than a certain threshold thresholds, it is considered that the voxel grid is occupied, that is, the position v of any voxel grid is expressed as:
[0166]
[0167] The set of occupied voxel grid positions is expressed as:
[0168]
[0169] where |∙| is the size of the set; M is the number of voxel grid positions used.
[0170] S500. Represent the features corresponding to the occupied voxel grid positions as a sparse tensor:
[0171]
[0172] where F(v) is the sparse tensor of any voxel grid position v; F S is the set of sparse tensors corresponding to the set S.
[0173] The P Gaussian kernels respectively correspond to P specific voxel grid positions {v1, v2,..., v p} ∈ S, and the set P S of reserved voxels is defined as:
[0174]
[0175] For P SEach voxel grid position v in j Apply a sparse convolution kernel to obtain the output feature F'(v):
[0176]
[0177] Where is the set of relative positions of the convolution kernel (for example, a 3×3×3 convolution kernel corresponds to ); W k is the convolution weight corresponding to the relative position k; is the relative position offset; σ is the non-linear activation function ReLU.
[0178] S600. Collect RGB images of N viewpoints in the robotic arm operation scenario. For each RGB image of each viewpoint, encode it through the ViT model to obtain the corresponding image feature:
[0179]
[0180] Where is the image feature of the nth viewpoint; I n is the RGB image of the nth viewpoint.
[0181] It should be noted that here N viewpoints generally refer to N third viewpoints.
[0182] S700. Perform cross-attention mechanism operation on the initialized query vector Query g and to obtain:
[0183]
[0184] Where cross_attn is cross-attention.
[0185] S800. Process the query vector Query g through a set of multi-layer perceptrons MLP to obtain the updated Gaussian kernel parameters :
[0186]
[0187] Where are the updated mean, scale, rotation vector, and semantic category respectively.
[0188] S900. Directly replace the original Gaussian kernel with the updated Gaussian kernel, and the new Gaussian kernel parameters are expressed as:
[0189]
[0190] Will Embed it into the target voxel grid with a size of X×Y×Z.
[0191] S1000. For each new three-dimensional Gaussian, calculate the radius of its neighborhood according to its scale attribute, and recalculate the occupancy prediction result of each voxel grid with reference to Formulas (2) and (4).
[0192] S1100. For each re-determined occupied voxel grid, calculate the corresponding label C of its semantic category label and the predicted value The cross-entropy loss CE between them is:
[0193]
[0194] where is the loss function; represents the set of occupied voxel grids; the subscript label represents the label.
[0195] S1200. Update the model parameters through the loss function, and save the query vector Query g and the weight parameters of the cross-attention Cross-Attn after training is completed.
[0196] On the basis of clarifying how to obtain the query vector Query g and the weight parameters of the cross-attention Cross-Attn, in this step, cross-attention calculation is performed on the image features to obtain high-dimensional features based on three-dimensional Gaussian; expressed as:
[0197]
[0198] where is the high-dimensional feature based on three-dimensional Gaussian at time t.
[0199] In step S3, the contact force distribution between the left and right jaws of the gripper and the object to be operated is sensed in real time by a force sensor, a two-dimensional RGB tactile image sequence is obtained, and the tactile image features of the RGB tactile image sequence are extracted by the ViT model.
[0200] It can be understood that in the embodiment of the present invention, the force sensor can be arranged at the end of the gripper in advance to support obtaining RGB values through the force sensor in this step, sensing the distribution of the contact force between the end of the gripper and the object to be operated in real time, and transmitting two-dimensional RGB tactile images in real time.
[0201] Next, in this step, the tactile image features of the RGB tactile image sequence are extracted by the ViT model; expressed as:
[0202]
[0203] Among them, is the haptic image feature at time t; are the RGB haptic images of the left and right gripper jaws at time t, respectively. The superscripts lta and rta represent left and right tactile perceptions, respectively.
[0204] In step S4, the high-dimensional feature and the haptic image feature are respectively used as inputs to a resampler. After obtaining the corresponding feature sequences, they are concatenated.
[0205] As Figure 2 shown, in this step, the high-dimensional feature and the haptic image feature are respectively used as inputs to a resampler. After obtaining the corresponding feature sequences, they are concatenated; the specific related steps include:
[0206] First, the high-dimensional feature and the haptic image feature are respectively input into the K and V fully connected layers of the resampler, and a set of learnable vectors q are input into the Q fully connected layer of the resampler:
[0207]
[0208] Among them, Q R ∈ is the query vector corresponding to the learnable parameters of the resampler, is the number of resampler query vectors, is the size of the hidden dimension, are the linear transformation matrices for the key and value respectively, and d g / ta represents the dimension size of the high-dimensional feature or the haptic image feature; are the transformed key and value vectors generated from the input of the high-dimensional feature respectively; are the transformed key and value vectors generated from the input of the haptic image feature respectively; are the feature sequences corresponding to the high-dimensional feature and the haptic image feature respectively output by the resampler; softmax is the activation function; the superscript T is the transpose symbol.
[0209] Next, the two obtained feature sequences are concatenated:
[0210]
[0211] Among them, is the concatenation result at time t. The superscript g-ta represents the fusion of the Gaussian feature and the tactile feature; concat is the concatenation operation.
[0212] In step S5, the splicing result and the language instruction are used as the input of the feature fusion decoder, and the LSTM policy head is used to perform action prediction on the output of the feature fusion decoder to obtain the pose change amount of the gripper and the next state of the gripper.
[0213] As described above, the resampler is used to combine the Gaussian features and the haptic image features, output the compressed features and input them into the feature fusion decoder.
[0214] In this step, the feature fusion decoder fuses the language instruction (which refers to: pick up the red block in Figure 2 ) with the splicing result of the encoded Gaussian features and haptic image features, and uses the LSTM policy head to perform action prediction on the output of the feature fusion decoder to obtain the pose change amount of the gripper and the next state of the gripper; the specific steps include:
[0215] First, the feature fusion process is executed by the feature fusion decoder and integrated with the language instruction to output ; where the feature fusion process is expressed as:
[0216]
[0217] Among them, is the output of the l-th cross-attention layer at time t; Tanh is the hyperbolic tangent function; α ∈ is a learnable parameter used to control the contribution of different features in the fusion process; A is the cross or self-attention operation; are the learnable parameters of the cross-attention layer respectively; are the outputs of the l-th and l+1-th self-attention layers at time t respectively (after being processed by the tokenizer and finding the corresponding word vectors, the first ) can be obtained; are the learnable parameters of the self-attention layer respectively; L is the sum of the number of alternating self-attention layers and cross-attention layers, and is an even number; N l is the number of language instructions.
[0218] It should be noted that during the fusion process, the type of attention layer to be input is judged according to the input of Q and K. If they are the same, it is the self-attention layer Self-Attn, otherwise it is the cross-attention layer Cross-Attn. Since the two are alternately stacked, it is exemplary to set both to 24 layers, a total of 48 layers, and the self-attention layer Self-Attn can use the pre-trained parameters of the large language model mpt-1b-redpajama-200b and freeze them, and the cross-attention layer Cross-Attn can use the pre-trained parameters of OpenFlamingo and fine-tune them.
[0219] Since the feature vector output by the feature fusion decoder needs to be converted into the underlying control instructions in the end, in order to facilitate the conversion, this step then uses the LSTM strategy head to output X t Perform motion prediction to obtain the position change of the gripper and the next state of the gripper, such as the 7 degrees of freedom (DoF) change of the gripper posture and the gripper state; specifically expressed as:
[0220]
[0221] Among them, MaxPooling is the maximum pooling layer, For its output; h t 、h t-1 are the hidden states at time t and t-1 respectively; are the predicted position change of the gripper and the next state of the gripper, respectively, and the state refers to whether the left and right jaws of the gripper are open or closed.
[0222] Specifically, to improve the pre-trained attention mechanism backbone and LSTM strategy head, this embodiment of the present invention employs a maximum likelihood objective for imitation learning. The goal is to optimize the relative pose using a regression loss (e.g., using mean squared error (MSE) loss) and classify the gripper state using a classification loss (using binary cross entropy (BCE) loss):
[0223]
[0224] in, is the total loss function, are the change in the gripper's posture at time t and the next state of the gripper, respectively, and λgripper is the weight factor of the gripper loss.
[0225] In step S6, a target posture is calculated based on the current posture of the gripper and the posture change amount to determine the moving direction and distance of the gripper.
[0226] As you can understand, this step determines the gripper's movement direction and distance to the target position through calculation, ensuring that the system completes the input language command. Moreover, because the model reasoning is performed in a time series, the gripper is always in the appropriate position at every moment.
[0227] Furthermore, considering that in reality, there is unpredictability and randomness in the output generation of multi-modal large models, which leads to disorder, uncertainty, and even potential danger in the operation instructions during robotic arm control. This uncertainty may generate inappropriate action sequences during critical mission phases, resulting in unintended or dangerous operations, limiting the application of large models in safety-sensitive scenarios, and reducing operation efficiency and increasing risks.
[0228] Correspondingly, the embodiment of the present invention innovatively proposes a dynamic feedback optimization control algorithm based on a safety region, aiming to solve the safety hazards in the operation instructions of the above-mentioned machine learning model.
[0229] Specifically, the proposed algorithm includes:
[0230] (1) Design a pose controller:
[0231] Based on the pose variation, define the first position error between the target position and the current position:
[0232]
[0233] where e1 is the first position error; is the pose variation; is the target pose of the gripper at the target position to be calculated; is the current pose of the gripper at the current position.
[0234] Use the pose controller to calculate the control input to ensure that the gripper moves smoothly towards the target position:
[0235]
[0236] where u1 is the output of the pose controller using a PD controller; k p1 、k d1 are the first proportional and first differential gains respectively; is the rate of change of the first position error.
[0237] (2) Design a safety region controller:
[0238] As Figure 4 shown, set the safety region as a planar region with a height of h. Outside this planar region, the free movement of the entire robotic arm is not restricted. Once the gripper approaches the region boundary within this planar region, the safety region controller outputs a control quantity away from the region boundary:
[0239]
[0240] where u2 is the output of the safety region controller; k p2、k d2 are the second proportional and second differential gains respectively; e2 is the second position error between the current position and the safety area, and the component of e2 in the Z direction , d z is the distance from the end of the gripper to the boundary of the region; is the rate of change of the second position error.
[0241] (3) Specify the evaluation function of the control objective:
[0242] Note that the expression for J is based on the state error, which is typically quadratic. This approach is crucial to ensure convexity and singularity of the solution to the optimization problem. It is assumed that every state of the system can be predicted based on a custom model.
[0243] Based on the above conditions, J for each target can be designed as follows:
[0244] Based on the square of the difference between the target pose and the current pose, the first evaluation function of the pose controller is designed:
[0245]
[0246] Where t0 is the initial time point, usually the time when the control process starts; T is the preset time span; ‖∙‖ is the modulus of the vector; and t is the time variable.
[0247] Based on the square of the difference between the distance from the end of the gripper to the boundary of the area and the height of the safety area, the second evaluation function of the safety area controller is designed:
[0248]
[0249] (4) Calculate the predicted evaluation gradient descent value and build a state feedback mechanism:
[0250] Design saturation suppression function :
[0251]
[0252] in, To evaluate the gradient (the gradient direction in the state space is the fastest convergent direction), i takes 1 or 2; α is the upper limit of the gradient amplitude.
[0253] Use gradient descent method to obtain the input after single objective optimization:
[0254]
[0255] Among them, u i is the basic control input; λ is the learning rate; K α is the saturation suppression function The corresponding gain factor.
[0256] Therefore, the state feedback mechanism is constructed to make the robot state converge quickly. The single-objective controller after adding state feedback is expressed as:
[0257]
[0258] Among them, K fb is the state feedback gain matrix, is the target state, x i The current actual status.
[0259] like Figure 4 As shown, the final comprehensive output is expressed as:
[0260]
[0261] (5) Calculate a new safe target pose to ensure that the gripper avoids touching the region boundary:
[0262]
[0263] in, is the updated target pose.
[0264] So far, the embodiment of the present invention has completed the entire process of the intelligent robot safe operation method based on the fusion of three-dimensional Gaussian and tactile images.
[0265] Example 2:
[0266] An embodiment of the present invention provides a secure operating system for an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images, wherein a gripper is installed at the end of the mechanical arm of the intelligent robot; the system comprises:
[0267] A collection module is used to collect RGB image sequences from the manipulator's hand perspective and the third perspective, and extract image features of the RGB image sequences using a pre-trained ViT model;
[0268] A calculation module, configured to perform a cross-attention calculation on the image features based on a pre-saved query vector and a cross-attention weight parameter to obtain a high-dimensional feature based on a three-dimensional Gaussian;
[0269] an extraction module, configured to sense in real time the contact force distribution between the left and right jaws of the gripper and the operated object using a force sensor, obtain a two-dimensional RGB tactile image sequence, and extract tactile image features of the RGB tactile image sequence using the ViT model;
[0270] a sampling module, configured to use the high-dimensional features and the tactile image features as inputs of a resampler, obtain corresponding feature sequences, and then perform splicing;
[0271] A prediction module is used to use the splicing result and the language instruction as the input of the feature fusion decoder, and use the LSTM strategy head to perform action prediction on the output of the feature fusion decoder to obtain the posture change of the gripper and the next state of the gripper;
[0272] The determination module is used to calculate the target posture based on the current posture of the gripper and the posture change amount to determine the moving direction and distance of the gripper.
[0273] Example 3:
[0274] An embodiment of the present invention provides a storage medium storing a computer program for safe operation of an intelligent robot based on fusion of three-dimensional Gaussian and tactile images, wherein the computer program enables a computer to execute the method for safe operation of an intelligent robot as described in Example 1.
[0275] Example 4:
[0276] An embodiment of the present invention provides an electronic device, including:
[0277] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing the intelligent robot safety operation method as described in Example 1.
[0278] It is understandable that the intelligent robot safety operating system, storage medium and electronic device based on the fusion of three-dimensional Gaussian and tactile images provided in the embodiments of the present invention correspond to the intelligent robot safety operation method based on the fusion of three-dimensional Gaussian and tactile images provided in the embodiments of the present invention. The explanations, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the intelligent robot safety operation method, and will not be repeated here.
[0279] In summary, compared with the existing technology, the present invention has the following beneficial effects:
[0280] 1. This embodiment of the present invention uses a 3D Gaussian model to represent the 3D space of the robotic arm's operating scenario, providing stable and accurate 3D spatial features for the robotic arm's multimodal large model. These features significantly enhance the robotic arm's ability to perceive and understand the operating environment, effectively improving the success rate and accuracy of the robotic arm in performing various operational tasks.
[0281] 2. In the embodiments of the present invention, force sensors are added to both ends of the gripper to obtain tactile feedback and fuse it with three-dimensional spatial features. The combination of force information and three-dimensional spatial features enables the robotic arm to more accurately perceive the force changes during the grasping process, optimize the grasping strategy, and improve the operation accuracy and stability. This technology enhances the ability of the robotic arm to interact with objects during operation.
[0282] 3. In the embodiments of the present invention, by introducing a control method with safety area constraints, the problem of uncontrollability of the output of large models is solved. Combining real-time environmental data with safety area constraints ensures that the robotic arm not only follows the inference trajectory during movement but also avoids entering dangerous areas, thereby improving safety. This method enhances the adaptability of the robotic arm in complex environments, avoids collision risks, and ensures the stability and success rate of task execution.
[0283] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0284] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent robot safety operation method based on the fusion of three-dimensional Gaussian and tactile images, characterized in that, A gripper is installed at the end of the robotic arm of the intelligent robot; it includes: Collect RGB image sequences from the perspective of the robotic arm's hand and the third perspective, and extract the image features of the RGB image sequences through a pre-trained ViT model; Based on the pre-saved query vector and the weight parameters of cross-attention, perform cross-attention calculation on the image features to obtain high-dimensional features based on a three-dimensional Gaussian; Real-time sense the contact force distribution between the left and right jaws of the gripper and the object to be operated through a force sensor, obtain a two-dimensional RGB tactile image sequence, and extract the tactile image features of the RGB tactile image sequence through the ViT model; Take the high-dimensional features and the tactile image features as the inputs of a resampler respectively, obtain the corresponding feature sequences and then splice them; Take the splicing result and the language instruction as the inputs of a feature fusion decoder, and use an LSTM policy head to perform action prediction on the output of the feature fusion decoder to obtain the pose change amount of the gripper and the next state of the gripper; Based on the current pose of the gripper and the pose change amount, calculate the target pose to determine the moving direction and distance of the gripper.
2. The intelligent robot safety operation method according to claim 1, wherein, During the process of calculating the target pose, introduce a dynamic feedback optimization control algorithm based on a safety region, including: (1) Design a pose controller: Based on the pose change amount, define the first position error between the target position and the current position: Among them, e1 is the first position error; is the pose change amount; is the target pose of the gripper at the target position to be calculated; is the current pose of the gripper at the current position; Use a pose controller to calculate the control input to ensure that the gripper moves smoothly towards the target position: Among them, u1 is the output of the pose controller that selects the PD controller; k p1 , k d1 are the first proportional and first derivative gains respectively; is the change rate of the first position error; (2) Design a safety region controller: Set the safety region as a planar region with a height of h. Outside this planar region, the free movement of the entire robotic arm is not restricted. Once the gripper approaches the region boundary within this planar region, the safety region controller outputs a control quantity away from the region boundary: Among them, u2 is the output of the safety area controller; k p2 , k d2 are the second proportional and second derivative gains respectively; e2 is the second position error between the current position and the safety area, and the component of e2 in the Z direction , d z is the distance from the end of the gripper to the area boundary; is the change rate of the second position error; (3) Specify the evaluation function of the control target: Based on the square of the difference between the target pose and the current pose, design the first evaluation function of the pose controller: Where, t0 is the initial time point; T is the preset time span; ‖∙‖ is the modulus of the vector; t is the time variable; Based on the square of the difference between the distance from the end of the gripper to the region boundary and the height of the safety region, design the second evaluation function of the safety region controller: (4) Calculate the predicted evaluation gradient descent value and construct a state feedback mechanism: Design saturation inhibition function : Among them, is the evaluation gradient, i takes 1 or 2; α is the upper limit of the gradient amplitude; Use the gradient descent method to obtain the input after single-target optimization: where, u i is the basic control input; λ is the learning rate; K α is the gain coefficient corresponding to the saturation inhibition function ; Construct the state feedback mechanism to enable the rapid convergence of the robot state. The single-target controller after adding state feedback is expressed as: Among them, K fb is the state feedback gain matrix, and the subscript fb represents state feedback, is the target state, and x i is the current actual state; The comprehensive output is: (5) Calculate a new safety target pose to ensure that the gripper avoids touching the region boundary: Among them, is the updated target pose.
3. The intelligent robot safe operation method according to claim 1 or 2, characterized in that, The acquisition process of the query vector and the weight parameters of cross-attention includes: Select a fixed world coordinate system within the robotic arm operation scenario; Adopt P three-dimensional Gaussian distributions represent the operation scenario of the robotic arm, and each three-dimensional Gaussian is represented by a d-dimensional vector in the form of , where d = 10 + |C sem +2|, and m, s, r, and c are the mean, scale, rotation vector, and semantic category respectively,[[]] is the set of real numbers, and C sem is the number of semantic labels; Initialization And it is a learnable vector ; where D is the dimension of the learnable vector, and D >> d Used to represent the parameter attributes of the Gaussian, Q j g Is a learnable three-dimensional space feature, the superscript g represents the query vector related to the Gaussian representation, and the subscript j represents the index of the Gaussian kernel; Perform voxelization processing according to the initialized Gaussian kernel, and the occupancy prediction result of each voxel grid is calculated by the following formula: Among them, is the sum of the contributions of each Gaussian distribution to any position p in the selected world coordinate system, and is used to characterize the occupancy prediction result; is the three-dimensional Gaussian occupancy prediction distribution at point p; is the set of neighboring Gaussians of point p; exp is the exponential function; Σ is the covariance matrix; S = diag(s) represents the scale matrix, diag is the function to extract the diagonal elements, and R = q2r(q) is the function to convert the quaternion q used to represent the three-dimensional space rotation into a rotation matrix; If the occupancy prediction result of any voxel grid is greater than a certain threshold thresholds, it is considered that the voxel grid is occupied, that is, the position v of any voxel grid is expressed as: Represent the set of occupied voxel grid positions as: where, |∙| is the size of the set; M is the number of voxel grid positions used; Represent the features corresponding to the occupied voxel grid positions as a sparse tensor: where F(v) is the sparse tensor at any voxel grid position v; F S is the set of sparse tensors corresponding to the set S; P Gaussian kernels respectively correspond to P specific voxel grid positions {v1, v2, …, v p} ∈ S, and define the set P of retained voxels S as follows: For each voxel grid position v in P S apply a sparse convolutional kernel to obtain the output feature F'(v): j Among them, is the set of relative positions of the convolution kernel; W k is the convolution weight corresponding to the relative position k; is the relative position offset; σ is the non-linear activation function ReLU; Collect RGB images from N perspectives in the robotic arm operation scenario. For each RGB image of each perspective, obtain the corresponding image features through encoding by the ViT model: Among them, is the image feature of the nth perspective; I n is the RGB image of the nth perspective; The initialized query vector Query g and perform cross-attention mechanism operations to obtain: where, cross_attn is cross-attention; The query vector Query g is processed by a set of multi-layer perceptrons MLP to obtain updated Gaussian kernel parameters : Among them, are the updated mean, scale, rotation vector, and semantic category respectively; Directly replace the original Gaussian kernel with the updated Gaussian kernel, and the parameters of the new Gaussian kernel are expressed as: Embed into a target voxel grid of size X×Y×Z; For each new 3D Gaussian, calculate the radius of its neighborhood according to its scale attribute, and recalculate the occupancy prediction result of each voxel grid; For each re-determined occupied voxel grid, calculate the label C corresponding to its semantic category label and the predicted value to obtain the cross-entropy loss CE between them: Among them, is the loss function; represents the set of occupied voxel grids; the subscript label represents the label; Update the model parameters through the loss function, and save the query vector Query in it after the training is completed. g And the weight parameters of the cross-attention Cross-Attn.
4. The intelligent robot safe operation method according to claim 3, characterized in that The image features of the RGB image sequence are represented as: Among them, is the image feature at time t, and the superscript represents vision; are the RGB images of the manipulator hand view and the third view at time t respectively; ⊕ is the element-wise addition symbol; Perform cross-attention calculation on the image features based on the pre-saved query vector and the weight parameters of cross-attention to obtain high-dimensional features based on 3D Gaussians; represented as: Among them, is the high-dimensional feature based on three-dimensional Gaussian at time t.
5. The intelligent robot safe operation method according to claim 4, characterized in that Extract the tactile image features of the RGB tactile image sequence through the ViT model; represented as: Among them, is the tactile image feature at time t; are the RGB tactile images of the left and right jaws of the gripper at time t, and the superscripts lta and rta represent left and right tactile perceptions respectively.
6. The intelligent robot safe operation method according to claim 5, wherein, Use the high-dimensional features and the tactile image features as the inputs of the resampler respectively, and splice them after obtaining the corresponding feature sequences; including: Input the high-dimensional features and the tactile image features into the K and V fully connected layers of the resampler respectively, and input a set of learnable vectors q into the Q fully connected layer of the resampler: Among them, Q R ∈ is a query vector corresponding to the learnable parameters of the resampler, is the number of resampler query vectors, is the size of the hidden dimension, are the linear transformation matrices of the key and value respectively, d g / ta represents the dimensionality size of the high-dimensional features or haptic image features; are the transformed key and value vectors generated from the high-dimensional feature inputs respectively; are the transformed key and value vectors generated from the haptic image feature inputs respectively; are the feature sequences corresponding to the high-dimensional features and haptic image features output by the resampler respectively; softmax is the activation function; the superscript T is the transpose symbol; Splice the two obtained feature sequences: Among them, is the splicing result at time t, where the superscript g-ta represents the fusion of Gaussian features and tactile features; concat is the splicing operation.
7. The intelligent robot safety operation method according to claim 6, characterized in that Use the splicing result and the language instruction as the input of the feature fusion decoder, and perform action prediction on the output of the feature fusion decoder using the LSTM policy head to obtain the pose change amount of the gripper and the next state of the gripper; including: The feature fusion process is performed by the feature fusion decoder and integrated with the language instruction for output ; where the feature fusion process is expressed as: Among them, is the output of the l-th cross-attention layer at time t; Tanh is the hyperbolic tangent function; α ∈ is a learnable parameter used to control the contributions of different features during the fusion process; A is the cross- or self-attention operation; are the learnable parameters of the cross-attention layer respectively; are the outputs of the l-th and (l + 1)-th self-attention layers at time t respectively; are the learnable parameters of the self-attention layer respectively; L is the sum of the number of alternating self-attention layers and cross-attention layers, and is an even number; N l is the number of language instructions; Use the LSTM policy head to predict the action for the output X t to obtain the pose change amount of the gripper and the next state of the gripper; which is expressed as: Among them, MaxPooling is the max pooling layer, is its output; h t , h t-1 are the hidden states at times t and t - 1 respectively; are respectively the predicted pose change amount of the gripper and the next state of the gripper, and the state refers to the open or closed state of the left and right jaws of the gripper.
8. An intelligent robot safety operating system based on the fusion of three-dimensional Gaussian and tactile images, characterized in that, A gripper is installed at the end of the robotic arm of the intelligent robot; including: A collection module, used to collect RGB image sequences from the perspective of the robotic arm hand and the third perspective, and extract the image features of the RGB image sequences through a pre-trained ViT model; A calculation module, used to perform cross-attention calculation on the image features based on the pre-saved query vector and the weight parameters of cross-attention to obtain high-dimensional features based on 3D Gaussians; An extraction module, used to sense the contact force distribution between the left and right jaws of the gripper and the object to be operated in real time through a force sensor, obtain a two-dimensional RGB tactile image sequence, and extract the tactile image features of the RGB tactile image sequence through the ViT model; A sampling module, used to use the high-dimensional features and the tactile image features as the inputs of the resampler respectively, and splice them after obtaining the corresponding feature sequences; A prediction module, used to use the splicing result and the language instruction as the input of the feature fusion decoder, and perform action prediction on the output of the feature fusion decoder using the LSTM policy head to obtain the pose change amount of the gripper and the next state of the gripper; A determination module, used to calculate the target pose based on the current pose of the gripper and the pose change amount to determine the moving direction and distance of the gripper.
9. A storage medium, characterized in that, It stores a computer program for the safe operation of an intelligent robot based on the fusion of three-dimensional Gaussian and tactile images, wherein the computer program causes a computer to execute the intelligent robot safe operation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Comprising: One or more processors; A memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include those for executing the intelligent robot safe operation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Mechanical arm autonomous operation strategy learning method based on vision-touch fusion
CN114660934A
Mechanical arm six-degree-of-freedom visual closed-loop grabbing method based on TSDF three-dimensional reconstruction
CN114851201A