A Robot Grasping Prediction Method Based on Triplet Contrast Network
Through the robot grasping prediction method based on triple comparison network, the ternary comparison loss function and self-attention mechanism are used to solve the problem of sample relationships being ignored and training in the existing technology, and achieve higher prediction accuracy and generalization.
Patent Information
- Application Number
- CN202211303570.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-10-24
AI Technical Summary
The existing robot crawling prediction algorithm ignores the relationship between samples during the training process, and based on data-driven, it requires a large amount of crawling data as support for the training set, which is expensive to train and has weak generalization.
Using a robot grab prediction method based on a triple comparison network, a training set is constructed, where each sample data reflects the correspondence between the tactile data at multiple moments during the robot arm grabs an object and whether the grab is successful. The encoder is trained using the ternary contrast loss function to generate high-dimensional features, and the classifier is trained in combination with the self-attention mechanism to improve the accuracy of network prediction.
This improves sample utilization, can obtain better network parameters from fewer samples, improves network prediction accuracy, alleviates overfitting problems, and improves the generalization of the model.
Smart Images

Figure CN115519579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and more specifically, to a robot grasping prediction method based on a triple contrast network. Background Art
[0002] With the development of the times, intelligent robot technology has been widely applied in many industries around the world, and the application scenarios of robots are becoming increasingly rich. For example, robots are used in industries such as automotive assembly, metal processing, 3C manufacturing, home appliance production, food packaging, and catering services. In these application scenarios, robot grasping is a common area that needs to be further studied and optimized. Before performing a grasp, accurately predicting the success rate of this grasp based on existing information is of great benefit for further generating high-quality and stable grasps.
[0003] In the prior art, patent application CN114700947A proposed a visual-tactile fusion grasping sliding probability prediction method. In this method, a visual camera and a tactile sensor are installed on the robot hand to obtain information in two modalities. The visual and tactile perception fusion algorithm uses the collected visual information to obtain the external contour information of the target object, and the tactile information to obtain information such as the surface topography and texture hardness of the target object. After preprocessing the collected information, a convolutional neural network is used to extract the features of the visual information and the features of the tactile information. The new features obtained after fusing the two modality feature information are used to predict the sliding probability during the object grasping process.
[0004] Patent application CN114083535A provides a physical measurement method and device for the grasping posture quality of a robot hand. The method includes: determining the flatness score of the contact point between the candidate grasping posture of the robot hand and the object; determining the centroid score of the candidate grasping posture for gripping the object; and evaluating the quality of the candidate grasping posture based on the flatness score and the centroid score. The physical measurement method for the grasping posture quality of the robot hand provided by this solution, based on the characteristics that humans tend to contact the flatter parts of an object and are used to contacting the central part of the object when picking up an object in daily life, evaluates the quality of the robot hand grasping posture from the flatness of the object itself and gravity analysis respectively through two physical measurement scores.
[0005] Although the above prior art has proposed solutions that can predict whether a grasp is successful from different perspectives, in terms of the data volume of the training set, some existing solutions are data-driven and require a large amount of grasping data as support for the training set, resulting in high training costs. From the perspective of feature extraction methods, in the training process of some existing solutions, the relationship between samples is not considered, and the sample information cannot be fully explored and utilized, or manual features are created and constructed, resulting in strong limitations, weak generalization ability, and low credibility.
[0006] In summary, when people grasp an object, they can sense and predict whether the current force is sufficient to stably grasp the object. Recent research has shown that by using the tactile information before the lifting stage of a robot grasping an object, it is also possible to effectively predict whether the robot's grasp is successful. However, the current robot grasping prediction algorithms ignore the relationship between samples during the training process and only train in a simple end-to-end manner. In addition, these methods are usually data-driven and it is difficult to effectively train the model when there is less grasping sample data available for training. Summary of the Invention
[0007] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide a robot grasping prediction method based on a triplet contrast network. The method includes the following steps:
[0008] Construct a training set, where each sample data reflects the corresponding relationship between the tactile data at multiple moments during the process of a robotic arm grasping an object and a classification label, and this classification label is used to indicate whether the grasp is successful;
[0009] Train an encoder based on a set loss function to obtain optimized parameters. During the training process, randomly select two samples with the same label from the training set as the anchor sample and the positive sample respectively, and randomly select a sample with the opposite label as the negative sample; input the sample data into the corresponding encoders respectively to encode and obtain the high-dimensional features of the anchor sample, the high-dimensional features of the positive sample, and the high-dimensional features of the negative sample;
[0010] Freeze the optimized parameters of the encoder, and use the trained encoder to encode the input sample data into high-dimensional features, and then input the high-dimensional features into a classifier for training;
[0011] Use the trained encoder and the trained classifier to predict the grasping result for the tactile data collected in real time.
[0012] Compared with the prior art, the advantages of the present invention are that the proposed robot grasping prediction method based on a triplet contrast network uses a deep neural network to train the encoder by comparing samples with each other, improving the sample utilization rate and being able to obtain better network parameters from fewer samples. Moreover, the present invention adds a self-attention mechanism to the network, which enables the network to more effectively capture the internal correlation of features, thereby further improving the network prediction accuracy.
[0013] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Brief Description of the Drawings
[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0015] Figure 1 is a flowchart of a robot grasping prediction method based on a triplet contrast network according to an embodiment of the present invention;
[0016] Figure 2 is a schematic diagram of an encoder network according to an embodiment of the present invention;
[0017] Figure 3 is a schematic diagram of encoder contrastive learning training according to an embodiment of the present invention;
[0018] Figure 4 is a schematic diagram of classifier training according to an embodiment of the present invention;
[0019] Figure 5 is a schematic diagram of the process of a robot grasping prediction method based on a triplet contrast network according to an embodiment of the present invention. Detailed Embodiments
[0020] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0021] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0022] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.
[0023] In all examples shown and discussed herein, any specific values should be construed as merely exemplary and not as limitations. Thus, other examples of the exemplary embodiments may have different values.
[0024] It should be noted that: Like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.
[0025] The present invention mainly relates to a robot grasping prediction method based on a triplet contrast network and belongs to the technical field of robots. The content of its technical solution is as follows:
[0026] See Figure 1 and Figure 5As shown, the provided robot grasping prediction method based on a triple contrast network includes the following steps:
[0027] Step S110, construct a training set, where each sample data reflects the correspondence between the tactile data at multiple moments during the process of the robotic arm grasping an object and whether the grasping is successful.
[0028] In practical applications, a robot grasping dataset based on touch or containing touch information can be obtained by self-collection or online downloading of public datasets.
[0029] For example, using the Calandra public dataset, a Weiss WSG-50 robotic arm with GELSIGHT tactile sensors installed on two fingers respectively is used in the dataset to collect tactile data, and a Microsoft Kinect 2 depth camera is placed in front of the grasping workbench to collect visual data. A total of 9,296 times of 106 different items are grasped. During the dataset collection process, when the robotic arm is in the initial position, it is marked as moment T a , when the robotic arm grasps the target item but has not yet lifted it, it is marked as moment T b , the moment 2 seconds after grasping the item and hovering in the air is T c , and the corresponding visual-tactile data is obtained for these three moments.
[0030] Furthermore, preprocess the dataset to enhance the data, including random horizontal flipping, random cropping, normalization processing, etc., and divide it into a training set and a test set.
[0031] For example, only use the tactile data at T a and T b moments. As data augmentation, randomly flip the tactile data images (either both tactile data from the two fingers are flipped or not), and randomly crop the 256×256 tactile data to 224×224.
[0032] In one embodiment, to improve the training efficiency and reduce the dependence on the sample size, divide the training set and the test set based on a small-sample scenario. For example, from the perspective of a small sample of the number of grasped item categories, randomly select the grasping data of 10 grasped items from the dataset to form the training set. From the perspective of a small sample of the grasping data, randomly select 20 data from the grasping data of each item as the training set samples. Thus, a training set with only 200 samples can be obtained. In addition, randomly select 1,000 times of grasping data from the grasping data of the remaining un-trained items as the test set.
[0033] In summary, the constructed training set has a small sample size. Each sample data reflects the correspondence between the tactile data (or tactile and visual data) at multiple moments during the process of the robotic arm grasping an object and the classification label, and the classification label is used to indicate whether the grasping is successful.
[0034] Step S120: Construct an encoder and train the encoder based on a large number of sample triples generated from a small number of samples in the training set.
[0035] Specifically, step S120 includes the following sub-steps:
[0036] Step S121: Construct an encoder
[0037] The encoder training framework adapts to different numbers of tactile sensor scenarios. For the case of only a single tactile sensor, the tactile data collected by the sensor at times T a and T b are respectively denoted as G a and G b . The difference between G a and G b is taken to highlight the changing part of the tactile information between the two moments, and G a and G b -G a are stacked in the channel direction to form a six-channel image and input into a pre-trained Resnet-50 backbone network without the last layer to extract high-dimensional features. For the case of multiple tactile sensors, after processing the data of different sensors in the above form, the high-dimensional features extracted from the data of different tactile sensors are concatenated to obtain high-dimensional features. Then, it is input into the self-attention module to obtain the fused high-dimensional features. For the schematic diagram of the encoder in the double-tactile sensor scenario, see Figure 2 . This encoder contains two backbone networks, and each backbone network corresponds to the tactile data of one sensor. It should be understood that when there are multiple sensors and the models of each sensor are the same and the contact situations with the target object are similar, the parameters of each backbone network can be shared to save the model space size, or the parameters can not be shared to fit a more complex mapping network to obtain better performance.
[0038] It should be understood that the backbone network adopted can be replaced by a network other than Resnet-50, such as VGG, AlexNet, etc. In addition, the input visual-based tactile modality information can theoretically also be replaced by other modality data, such as visual photographing information, etc.
[0039] Step S122: Train the encoder
[0040] See Figure 3As shown in the figure, during the training process, two samples with the same label (i.e., both captured successfully or failed) are randomly selected from the training set as anchor samples and positive samples, and a sample with the opposite label is randomly selected as a negative sample. The sample data are input into the encoder to encode the high-dimensional features of the anchor sample, positive sample, and negative sample, which are represented as f a 、f p 、f n Then, the triplet contrastive loss function is used as the evaluation index, which is formulated as follows:
[0041]
[0042] Where L T Denotes the triple contrast loss. By replacing L T To minimize, the encoder will make the Euclidean distance between the high-dimensional features encoded by samples of the same label as small as possible, and the Euclidean distance between the high-dimensional features encoded by samples of different categories as large as possible. By designing this triple form, a large number of sample triplets can be generated to train the encoder when there are fewer samples, thereby alleviating the overfitting problem and training an effective encoder.
[0043] In summary, when the sample size is small, by constructing a triple contrast learning deep neural network, we can explore the relationship between samples, improve the accuracy of the encoder, and combine the self-attention mechanism to consider the contribution of different samples to the encoder, further exploring the correlation between samples.
[0044] Step S130, using the trained encoder to encode the input sample data into high-dimensional features, and using the high-dimensional features to train a classifier, wherein the parameters of the encoder are frozen during the training process.
[0045] In one embodiment, step S130 includes the following sub-steps:
[0046] Step S131, constructing a classifier
[0047] After obtaining the appropriate encoder parameters, the encoder parameters are frozen, the input samples are encoded into high-dimensional features, and then the high-dimensional features are input into the subsequent classifier network to obtain the prediction results. The classifier network is as follows Figure 4 As shown, it as a whole includes a self-attention module and a multi-layer perceptron, and the multi-layer perceptron can be constructed using multiple layers of fully connected layers.
[0048] Specifically, the encoded high-dimensional features are input into the self-attention module, and then the processed features are input into a multi-layer perceptron composed of two fully connected layers. The number of input neurons in the first fully connected layer is 2048*N, where N represents the number of sensor sources of the input tactile data, the number of output neurons is 1024, followed by a Relu non-linear transformation activation layer; the number of input neurons in the second fully connected layer is 1024, the number of output neurons is 2, followed by a Sigmoid non-linear transformation activation layer, and finally the scores for predicting successful or failed grasps are output respectively.
[0049] Step S132, training the classifier
[0050] The above classifier neural network is trained using the training set with the Binary Cross-Entropy loss function. The Binary Cross-Entropy loss function L B is formally expressed as follows:
[0051] L B = -(y log(p(x)) + (1 - y) log(1 - p(x)))) (2)
[0052] where x represents the input sample data, y represents the category of whether the sample actually grasps successfully or fails. For example, y = 1 represents a successful grasp, y = 0 represents a failed grasp, and p(x) represents the probability that the classifier predicts the sample to grasp successfully under the current input sample data.
[0053] After the classifier training is completed, the test set can be used to verify the effectiveness of the model, and the evaluation index is the correct prediction accuracy rate.
[0054] Step S140, combining the trained encoder and classifier into a prediction model to predict the grasping result of the robot for the real-time collected tactile data.
[0055] After the training of the encoder and classifier is completed, the encoder and classifier can be combined into an overall prediction model, and this prediction model is used to predict whether the robotic arm's grasp is successful for the real-time collected tactile data.
[0056] In summary, compared with the prior art, the present invention has at least the following advantages:
[0057] 1) The present invention provides a deep neural network (including an encoder and a classifier, etc.) that adapts to any number of tactile sensors to predict the grasping stability through the tactile information before grasping. It can effectively predict the grasping stability in the scenarios of single or multiple tactile sensor information with basically the same framework, and its framework, core idea are independent of the types of tactile sensors and robotic hands, having good generalization.
[0058] 2) The present invention trains an encoder based on triplet contrastive learning. Through contrastive learning, the objective is to narrow the Euclidean distance of the deep high-dimensional features after encoding for similar samples and expand the Euclidean distance of the deep high-dimensional features after encoding for dissimilar samples, so as to train the feature encoder, discover the relationships between samples, and generate a large number of new training units by constructing triplets, thereby improving the network prediction accuracy performance in the few-shot scenario, alleviating the overfitting problem that is prone to occur during the training process when the number of samples is small, and obtaining better encoder parameters. By comparing samples, the mutual relationships between samples are mined, rather than considering samples separately for end-to-end training. Such an approach makes the training process make more full use of the potential information of the sample data and obtains better training results.
[0059] 3) The present invention introduces a self-attention mechanism into the network. The self-attention mechanism helps the network mine the internal correlations of features during the training process, highlight the important parts in the high-dimensional deep features that contribute to improving the classification accuracy, and further improve the prediction accuracy of the model.
[0060] 4) The network structure proposed by the present invention has been well verified on the Calandra public dataset, and the prediction performance is improved compared with the existing models.
[0061] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.
[0062] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the above. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0063] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0064] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.
[0065] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0066] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, the instructions of which implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0067] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0068] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.
[0069] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A robot grasping prediction method based on a triple contrast network, comprising the following steps: Construct a training set, where each sample data reflects the correspondence between tactile data at multiple moments during the process of a robotic arm grasping an object and a classification label, and this classification label is used to indicate whether the grasping is successful; Train an encoder based on a set loss function to obtain optimized parameters. During the training process, randomly select two samples with the same label from the training set as the anchor sample and the positive sample respectively, and randomly select a sample with the opposite label as the negative sample; input the sample data into the corresponding encoder respectively to encode the high-dimensional features of the anchor sample, the high-dimensional features of the positive sample, and the high-dimensional features of the negative sample; Freeze the optimized parameters of the encoder, and use the trained encoder to encode the input sample data into high-dimensional features, and then input these high-dimensional features into a classifier for training; Use the trained encoder and the trained classifier to predict the grasping result for real-time collected tactile data.
2. The method according to claim 1, characterized in that, the training set is constructed according to the following steps: Collect the tactile data of the robotic arm grasping an object, including the tactile data at three moments, namely, when the robotic arm is placed at the initial position, marked as moment T a , when the robotic arm grasps the target item but has not yet lifted it, marked as moment T b , and when the item is grasped and hovers in the air for a set time, marked as moment T c ; By randomly flipping the tactile data images at time T a and T b to enhance the data, and randomly cropping the tactile data to a set size to obtain a data set; Randomly select a set number of types of grasped items from the dataset, and randomly select a set number of data for each item to construct the training set.
3. The method according to claim 1, characterized in that, the encoder is trained using a triple contrast loss function as an evaluation metric, expressed as: where L T represents the triplet contrast loss, and f a , f p , f n represent the high-dimensional features of the anchor sample, positive sample, and negative sample, respectively.
4. The method according to claim 1, characterized in that, the classifier is trained using a binary cross-entropy loss function, expressed as: L B = -(y log(p(x) + (1 - y) log(1 - p(x)))) where x represents the input sample data, y represents the class label indicating whether the sample actually grasps successfully or fails, and p(x) represents the probability of predicting successful grasping under the current input sample data.
5. The method according to claim 2, characterized in that, The encoder includes a backbone network, which is used to extract high-dimensional features from the images stacked in the channel direction from G a and G b -G a where G a represents the tactile data collected at time T a and G b represents the tactile data collected at time T b .
6. The method according to claim 5, characterized in that, A plurality of backbone networks are set, and each backbone network is used to extract high-dimensional features from the corresponding tactile data, and after splicing the extracted high-dimensional features, input them into a first self-attention module to obtain fused high-dimensional features.
7. The method according to claim 6, characterized in that, The classifier includes a second self-attention module and a multi-layer perceptron. The multi-layer perceptron includes two fully connected layers, where a Relu non-linear transformation activation layer is connected after the first fully connected layer; the number of output neurons of the second fully connected layer is 2, and a Sigmoid non-linear transformation activation layer is connected after it, and the classifier outputs the score of predicting successful or failed grasping.
8. The method according to claim 5, characterized in that, The backbone network is constructed based on Resnet-50, VGG or AlexNet network.
9. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer device, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Physical measurement method and device for grabbing posture quality of manipulator
CN114083535A
Robotic grasping prediction using neural networks and geometry aware object representation
CN110691676A
Robot grabbing state distinguishing method based on tactile array
CN111459278A