A multi-modal tactile data joint perception method and related equipment
By building a convolutional cyclic autoencoder network and a bidirectional cyclic autoencoder network in the field of robotic medical care, and using joint loss functions combined with multimodal tactile data, the problem of difficulty in effectively utilizing tactile data in the existing technology is solved, and high-precision identification of tumor depth is achieved.
Patent Information
- Application Number
- CN202310245173.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-06
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-03-06
AI Technical Summary
The prior art is difficult to effectively utilize different modal haptic data, resulting in low recognition accuracy in the field of robotic medical care, especially in deep tumor recognition.
By obtaining the haptic array data set and the haptic color image data set, a convolutional cyclic autoencoder network and a bidirectional cyclic autoencoder network are built, and the two networks are jointly built using a joint loss function to form a tumor deep recognition model.
The effective combined utilization of multimodal tactile data with large characteristics is achieved, and the accuracy of tumor depth recognition is improved.
Smart Images

Figure CN116352735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot tactile technology, and in particular to a multimodal tactile data joint perception method, system, terminal and computer-readable storage medium. Background Art
[0002] Robotics plays an important role in the medical industry, such as robot-assisted minimally invasive surgery. Although it has many advantages, the surgeon can only indirectly contact the patient's soft tissue because the surgeon operates remotely on the main console and only uses visual feedback from the electronic screen to perform the operation. The tactile feedback of the surgical process is missing, such as the sense of touch of human tissue. This lack of tactile feedback limits the doctor's perception and recognition of lesions such as tumors to a certain extent.
[0003] How to sense and identify tumors is a common problem in the field of robotic medicine. At the same time, there has long been a problem in the field of robotic touch. There are many types of sensors, including piezoelectric sensors, piezoresistive sensors, capacitive sensors, and a series of sensors based on optical principles. The characteristics of data sets collected by different tactile sensors are very different, making it difficult to combine, utilize, and promote each other. Therefore, how to effectively mine and utilize different modal tactile data to improve the recognition accuracy of tumor depth has become a problem in the field of robotic touch.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a multimodal tactile data joint perception method, system, terminal and computer-readable storage medium, aiming to solve the problem in the prior art that it is difficult to effectively utilize tactile data with large feature differences in traditional classification tasks.
[0006] To achieve the above object, the present invention provides a multimodal tactile data joint perception method, the multimodal tactile data joint perception method comprising the following steps:
[0007] Acquire a tactile array dataset and a tactile color image dataset, wherein the tactile array dataset is acquired based on a capacitive sensor and obtained after preprocessing, and the tactile color image dataset is acquired based on a digit sensor and obtained after preprocessing, and divide the tactile array dataset and the tactile color image dataset into a training set, a validation set, and a test set according to a preset ratio;
[0008] Building a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and jointly building the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through joint loss to obtain a tumor depth recognition model;
[0009] The tumor depth recognition model is trained using the training set and the validation set, and the effectiveness of the trained tumor depth recognition model is verified using the test set, wherein the effectiveness evaluation index is the tumor depth classification accuracy.
[0010] Optionally, in the multimodal tactile data joint perception method, the capacitive sensor collecting tactile array data comprises:
[0011] Before touching, a BarrettHand robotic arm was placed at a height of 2.5 cm from the surface of the soft tissue mold, wherein the BarrettHand robotic arm includes three fingers, each of which is provided with a capacitive sensor;
[0012] Controlling the BarrettHand robotic arm to press downward the surface of the soft tissue mold at a speed of 0.028 m / s;
[0013] During the pressing process, the surface of the capacitive sensor provided on the finger of the BarrettHand mechanical arm is always parallel to the surface of the soft tissue, and the capacitive sensor acquires the original tactile array data and torque data;
[0014] Each soft tissue sample was pressed 60 times, and after each pressing, the soft tissue sample was rotated 7.5 degrees using a rotating stage to obtain 720 sets of original tactile array data and torque data.
[0015] Optionally, in the multimodal tactile data joint perception method, the capacitive sensor collects tactile array data for preprocessing, comprising:
[0016] For each set of original tactile array data and torque data, the torque data is only used to obtain the time point when the BarrettHand robot arm just touches the soft tissue mold surface, and this time point is selected as the starting point to intercept a fixed length of 72 time points of tactile array data;
[0017] Downsampling the intercepted tactile array data, taking the starting point as the first sampling point, setting the downsampling step length to 6, and obtaining tactile array data with a length of 12 time points;
[0018] The downsampled data is reshaped to convert the tactile array data at each time point into an 8×9 image size with 1 channel to the processed tactile array data set.
[0019] Optionally, in the multimodal tactile data joint perception method, the Digit sensor collecting tactile color image data comprises:
[0020] Before each compression, the UR5 robot arm was placed at a height of 2.5 cm from the surface of the soft tissue mold. The UR5 robot arm was equipped with a Digit sensor.
[0021] Control the UR5 robot arm to press the surface of the soft tissue mold downward at a speed of 0.03 m / s, and stop when the soft tissue mold is pressed to a depth of 0.5 cm;
[0022] During the pressing process, the surface of the Digit sensor is always pressed on the tumor position to obtain the original tactile color image data. The sampling frequency of the UR5 robot arm is set to 30fps, and the image resolution is 320×240;
[0023] Each soft tissue sample was pressed 80 times, and after each pressing, the soft tissue sample was rotated 7.5 degrees using a rotating stage to obtain 960 sets of original tactile color image data.
[0024] Optionally, in the multimodal tactile data joint perception method, the Digit sensor collects tactile color image data for preprocessing, comprising:
[0025] Treat each piece of original tactile color image data as a video of a pressing process, determine the starting frame in each piece of original tactile color image data where the Digit sensor just touches the soft tissue mold, and then select 10 frames of a fixed length as each piece of processed tactile color image data;
[0026] Whether the soft tissue is touched is determined based on the change in pixel values in the color image. First, 50 non-contact color images are selected, and the average and standard deviation are taken according to the channel direction to obtain the average matrix and standard deviation matrix;
[0027] For each current image in the original tactile color image data, the current image is first averaged in the channel direction, and then the average matrix is subtracted and compared with the 4 times standard deviation matrix. If more than 6% of the pixels are larger than the 4 times standard deviation corresponding to the current image in the standard deviation matrix, it means that the current image is in contact with the soft tissue mold. Then, the starting frame that just touched the soft tissue mold is calculated based on the time series. Finally, the captured image is normalized by a linear function to obtain the tactile color image dataset.
[0028] Optionally, in the multimodal tactile data joint perception method, the convolutional recurrent autoencoder network is constructed by a convolutional neural network and a long short-term memory network; the convolutional recurrent autoencoder network includes an encoder part and a decoder, the encoder includes a convolutional layer, a batch normalization layer, a maximum pooling layer and a long short-term memory network layer; the decoder includes a long short-term memory network layer and a deconvolution layer, which is used to decode the encoded features;
[0029] The bidirectional recurrent autoencoder network is constructed by a fully connected network and a bidirectional long short-term memory network; the bidirectional recurrent autoencoder network includes an encoder part and a decoder, the encoder includes a fully connected layer and two layers of bidirectional long short-term memory network layers; the decoder includes two layers of bidirectional long short-term memory network layers and a fully connected layer.
[0030] Optionally, in the multimodal tactile data joint perception method, the convolutional recurrent autoencoder network uses a cross entropy loss for the classification loss of the tactile array data set, and a mean square error is used for the reconstruction loss function, and the classification loss function and the reconstruction loss function are as follows:
[0031]
[0032]
[0033] Where N represents the number of tactile array datasets, C represents the number of categories, represents the indicator function, y i represents the label of the i-th sample, X i represents the i-th sample data, represents the output data of the i-th sample after the convolutional cyclic autoencoder network, Represents sample data X i Prediction vector after softmax function;
[0034] The bidirectional recurrent autoencoder network uses cross entropy loss for the classification loss of the tactile color image dataset, and the reconstruction loss function uses mean square error. The classification loss and reconstruction loss are as follows:
[0035]
[0036]
[0037] Among them, M represents the number of tactile color image datasets, C represents the number of categories, represents the indicator function, y j represents the label of the jth sample, X j represents the jth sample data, represents the output data of the jth sample after passing through the bidirectional cyclic autoencoder network, Represents sample data X j The prediction vector of
[0038] Based on the convolutional recurrent autoencoder network, the latent vector of the tactile array data set is obtained. Based on the bidirectional recurrent autoencoder network, the latent vector of the tactile color image data set is obtained. The method of combining the high-dimensional features of multimodal tactile data performs principal component analysis on the latent vectors with a dimension greater than a preset threshold, reduces the dimension to the same dimension as the latent vectors with a dimension less than the preset threshold, calculates the mean square error of the two to obtain the joint loss, and adds the joint loss to the total loss function of the network as follows:
[0039]
[0040] Among them, L represents the total loss, w r represents the weight of the reconstruction loss, w c Represents the weight of classification loss, w z represents the weight of the joint loss, z1 represents the latent vector of the tactile dataset whose dimension is greater than z2, z2 represents the latent vector of the tactile dataset whose dimension is less than z1, z1' represents the latent vector of z1 after principal component analysis, and the tactile dataset includes the tactile array dataset and the tactile color image dataset.
[0041] In addition, to achieve the above-mentioned purpose, the present invention further provides a multimodal tactile data joint perception system, wherein the multimodal tactile data joint perception system comprises:
[0042] a tactile data set acquisition module, used to acquire a tactile array data set and a tactile color image data set, wherein the tactile array data set is acquired based on a capacitive sensor and is obtained after preprocessing, and the tactile color image data set is acquired based on a digit sensor and is obtained after preprocessing, and the tactile array data set and the tactile color image data set are divided into a training set, a validation set, and a test set according to a preset ratio;
[0043] A model joint construction module, used to construct a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and to jointly construct the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through a joint loss to obtain a tumor depth recognition model;
[0044] A training and verification module is used to train the tumor depth recognition model using the training set and the verification set, and to verify the effectiveness of the trained tumor depth recognition model using the test set, wherein the effectiveness evaluation indicator is the tumor depth classification accuracy.
[0045] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal tactile data joint perception program stored in the memory and runnable on the processor, and when the multimodal tactile data joint perception program is executed by the processor, the steps of the multimodal tactile data joint perception method as described above are implemented.
[0046] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal tactile data joint perception program, and when the multimodal tactile data joint perception program is executed by a processor, the steps of the multimodal tactile data joint perception method as described above are implemented.
[0047] In the present invention, a tactile array data set and a tactile color image data set are obtained, the tactile array data set is obtained based on the acquisition of the capacitive sensor and after preprocessing, the tactile color image data set is obtained based on the acquisition of the digit sensor and after preprocessing, and the tactile array data set and the tactile color image data set are divided into a training set, a validation set and a test set according to a preset ratio; a convolutional cyclic autoencoder network for the tactile array data set and a bidirectional cyclic autoencoder network for the tactile color image data set are built, and the convolutional cyclic autoencoder network and the bidirectional cyclic autoencoder network are jointly built through a joint loss to obtain a tumor depth recognition model; the tumor depth recognition model is trained using the training set and the validation set, and the effectiveness of the trained tumor depth recognition model is verified using the test set, and the effectiveness evaluation index is the tumor depth classification accuracy. The present invention aims to explore and utilize the potential relevant information of multimodal tactile data sets with large feature differences, improve the problem of the difficulty in effectively utilizing tactile data with large feature differences in traditional classification tasks, and further enhance the mutual promotion effect of tactile data sets by entering a reasonable joint loss function, thereby improving the recognition accuracy of tumor depth. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flow chart of a preferred embodiment of the multimodal tactile data joint perception method of the present invention;
[0049] Figure 2 It is a flow chart of a tumor depth classification method in a preferred embodiment of the multimodal tactile data joint perception method of the present invention;
[0050] Figure 3 It is a schematic diagram of the principle of building a tumor depth recognition model in a preferred embodiment of the multimodal tactile data joint perception method of the present invention;
[0051] Figure 4It is a schematic diagram of the principle of a preferred embodiment of the multimodal tactile data joint perception system of the present invention;
[0052] Figure 5 Schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0054] Since the birth of the world's first da Vinci surgical robot, robotics has gradually become inseparable from the medical field. In traditional open surgery, experienced doctors can use palpation to diagnose deep tumors in human soft tissues that are difficult to observe with the naked eye, but traditional open surgery causes great trauma to patients. Robot-assisted minimally invasive surgery alleviates this problem, providing more precise control and better dexterity, but the lack of tactile feedback during surgery hinders doctors' judgment of lesion information. Therefore, robotic palpation tries to solve this problem. As a new technology that combines robotics and medical palpation, robotic palpation uses robot-assisted diagnosis to help doctors diagnose patients' pathological information, such as the depth of tumors. Generally speaking, the tactile sensor integrated in the medical robot arm touches the human soft tissue and senses and obtains tactile information, and then uses machine learning and deep learning algorithms to determine whether there is a tumor and the depth of the tumor from the surface. The tactile sensor itself also has the characteristic that data is difficult to combine with each other. Capacitive sensors (such as RoboTouch sensors) collect tactile array data, while sensors based on optical principles (such as Digit sensors and Gelsight sensors) collect color image data of the contact surface. The large modality differences between tactile datasets make it difficult for them to leverage each other and achieve a mutually reinforcing effect.
[0055] The existing technology has a tumor recognition method based on a multi-layer recognition convolutional neural network. This method mainly recognizes tumors by using an algorithm based on a multi-layer convolutional neural network. This method is divided into two stages: training and recognition. In the training stage, the tumor is first segmented using a full convolutional neural network, then the morphological recognition network is trained, and then the cost-sensitive recognition network is trained; in the recognition stage, the full convolutional neural network is used to segment the tumor in the image to be recognized, and then the segmented tumor image is fed into the network.
[0056] There is also a grasped object recognition method based on the fusion of tactile signals and visual images. The invention first uses a camera device to take a visual color photo of the grasped object. During the grasping process, the tactile sensor will also collect tactile vibration signals. Then, the tactile signal is converted into a tactile color photo according to the color photo value. Then, the tactile photo and the visual photo are superimposed in the channel direction, so that six-channel input image data is obtained. Finally, the six-channel image data is input into the convolutional neural network for training, and finally a network for identifying objects is obtained. Multimodal data is used to improve the robot's object recognition performance.
[0057] Although the above-mentioned methods all use machine learning and deep learning techniques, such as using convolutional neural networks to implement classification tasks, some of them realize the perception and recognition of tumors in medical scenarios, and some combine multimodal data by channel superposition, but they are different from the technical solutions proposed in the present invention. From the algorithm level, the deep learning network framework used in the above scheme is a convolutional neural network, which ignores the temporal information of the data and fails to further improve the performance of the algorithm. From the data level, the above scheme is too simple for the joint processing of multimodal data, and does not take into account the different problems of networks applicable to different modal data. The data features of different modalities are very different (such as tactile array signals and color images). The network suitable for a certain modality data may not be suitable for another modality. Forcibly converting the original tactile vibration signal into a color image may cause the original information to be lost to some information, so it should be considered to find a network suitable for different modalities, effectively utilize the data information of each modality, and then realize the combination of multimodal data in high-dimensional space through joint cost loss, so as to achieve better results. Therefore, based on this idea, the purpose of the present invention is to efficiently extract the spatial and temporal information of each modality data by using different autoencoder network structures, introduce joint loss rather than simple splicing, and ultimately achieve high-precision tumor depth recognition.
[0058] The multimodal tactile data joint perception method described in the preferred embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the multimodal tactile data joint perception method includes the following steps:
[0059] Step S10, obtaining a tactile array dataset and a tactile color image dataset, wherein the tactile array dataset is acquired based on a capacitive sensor and obtained after preprocessing, and the tactile color image dataset is acquired based on a digit sensor and obtained after preprocessing, and the tactile array dataset and the tactile color image dataset are divided into a training set, a validation set and a test set according to a preset ratio.
[0060] Specifically, Figure 2As shown, first prepare the materials required for the robot palpation experiment. First prepare the robot equipment. The present invention uses the BarrettHand robot arm and the UR5 robot arm. The sensors used are the capacitive sensor integrated in the BarrettHand and the Digit sensor installed on the UR5. Make the relevant molds. The present invention uses 12 soft tissue molds as simulated human soft tissues, including a soft tissue mold without a tumor inside and 11 soft tissue molds with tumors inside. The tumor depths are 1mm, 2mm, ..., 11mm respectively. Carry out the robot palpation experiment, use the capacitive sensor of the BarrettHand robot arm and the Digit sensor of the UR5 robot arm to press the soft tissue surface, and collect tactile data of different modes respectively. The capacitive sensor of the BarrettHand robot arm collects tactile array data, and the Digit sensor of the UR5 robot arm collects surface deformation color image data, both of which are time series data.
[0061] Among them, the tactile array dataset and the tactile color image dataset are called the tactile dataset (multimodal tactile dataset for network training). The multimodal tactile dataset for network training includes the tactile array signals generated during the contact process between the soft tissue mold and the capacitive sensor of the BarrettHand robotic arm, and the sensor surface deformation color image signals generated during the interaction with the Digit sensor of the UR5 robotic arm.
[0062] Tactile array data set acquisition: First, before each touch, the BarrettHand robotic arm is placed at a height of 2.5 cm from the soft tissue surface. Then, the BarrettHand robotic arm is controlled to press down on the surface of the soft tissue at a speed of 0.028 m / s. The BarrettHand robotic arm has three fingers, and each finger is equipped with a capacitive sensor. During the pressing process, the capacitive sensor surface of the BarrettHand robotic arm's fingers is always parallel to the soft tissue surface, and then the capacitive sensor obtains the original tactile array data and torque data (only the torque data is used for data preprocessing of the tactile array data, and the torque data does not participate in network training). Each soft tissue sample is pressed 60 times, and after each press, the sample is rotated 7.5 degrees using a rotating table to obtain more tactile data, so 720 sets of original tactile array data and torque data are obtained. The rotation operation is to increase the amount of data and improve the robustness of the algorithm.
[0063] Acquisition of tactile color image data set: First, before each press, the UR5 robotic arm (the UR5 robotic arm is equipped with a Digit sensor) is placed at a height of 2.5 cm from the surface of the soft tissue mold. Then, the UR5 robotic arm is controlled to press the surface of the soft tissue mold downward at a speed of 0.03 m / s, and stops when the soft tissue mold is pressed to a depth of 0.5 cm. During the pressing process, the surface of the Digit sensor is always pressed on the tumor position, and then the original tactile color image data is obtained. The sampling frequency of the Digit sensor is set to 30fps, and the image resolution is 320×240. Each soft tissue sample is pressed 80 times, and the soft tissue sample is rotated 7.5 degrees after each press, and finally 960 sets of original tactile color image data are obtained.
[0064] The tactile array dataset and tactile color image dataset are preprocessed respectively, including data truncation, downsampling, normalization and data reshaping.
[0065] Specifically, Figure 2 As shown in the figure, the preprocessing of the tactile array data set is as follows: for each piece of original tactile array data and torque data, as mentioned above, the torque data does not participate in the network training. The torque data is only used to obtain the time point when the BarrettHand robot arm just touches the surface of the soft tissue model. This time point is selected as the starting point, and then a fixed length of 72 time points of tactile array data (including the starting point) is intercepted. Since a small amount of data is sufficient to distinguish the depth of the tumor, the intercepted data is downsampled, the starting point is taken as the first sampling point, and the downsampling step is set to 6, and the tactile array data of 12 time points are obtained. Then the downsampled data is reshaped, and the tactile array data of each time point is converted into an 8×9 image size with a channel number of 1, so that the processed tactile array data set is obtained.
[0066] Specifically, Figure 2As shown in the figure, the preprocessing of the tactile color image data set is as follows: Similarly, each original tactile color image data can be regarded as a video of a pressing process (the video contains 240 frames). The key is to determine the starting frame of each original tactile color image data where the Digit sensor just touches the soft tissue, and then select 10 frames of fixed length as each processed tactile color image data (including the starting frame). How to determine the starting frame is the key. According to the change of pixel values in the color image, whether the soft tissue is touched is judged. First, 50 non-contact Digit color images are selected, and the average and standard deviation are taken according to the channel direction to obtain the average matrix and standard deviation matrix. Then, for the current image in each original tactile color image data, the current image is first averaged according to the channel direction, and then the average matrix is subtracted, and then compared with the 4 times standard deviation matrix. If more than 6% of the pixels are larger than the corresponding 4 times standard deviation in the standard deviation matrix, it means that the current image is in contact with the mold. Then, according to the time series, the starting frame that just touches the soft tissue can be inferred. Finally, the intercepted image is normalized by linear function to obtain the tactile color image data set.
[0067] Finally, the processed tactile array dataset is 720 in total, and the tactile color image dataset is 960 in total. For these two datasets, the training set, validation set, and test set are divided into 7:1:2 ratios. The tumor depths are 0mm, 1mm, ..., 11mm, and the labels are set as 1, 2, ..., 12 respectively.
[0068] Step S20, building a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and jointly building the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through a joint loss to obtain a tumor depth recognition model.
[0069] Specifically, Figure 3 As shown, for tactile array data, the convolutional cyclic autoencoder network used is built by a convolutional neural network (CNN) and a long short-term memory network (LSTM); CNN can effectively capture image information, and LSTM is good at processing time series data. The convolutional cyclic autoencoder network includes an encoder part and a decoder. The encoder includes a convolution layer (the convolution layer contains 32 convolution filters of size 3×3, with a step size of 1), a batch normalization layer, a maximum pooling layer, and a long short-term memory network layer; the decoder includes a long short-term memory network layer and a deconvolution layer, which is used to decode the encoded features and restore them to input data as much as possible.
[0070] like Figure 3As shown, for tactile color image data, the bidirectional recurrent autoencoder network used is built by a fully connected network and a bidirectional long short-term memory network (BiLSTM); the bidirectional recurrent autoencoder network includes an encoder part and a decoder, the encoder is first a fully connected layer, and then connected to two layers of bidirectional long short-term memory network layers; the decoder is first connected to two layers of bidirectional long short-term memory network layers, and then a fully connected layer.
[0071] The tactile array dataset and the tactile color image dataset are input into their respective suitable autoencoder networks, and the features of the middle hidden layer of the autoencoder are extracted for twelve classifications. Then, the joint loss is introduced so that the features of the hidden layers of the two autoencoder networks can promote each other, thereby improving the recognition accuracy of tumor depth.
[0072] like Figure 3 As shown, Figure 3 Conv2D in the figure represents a convolutional layer, ReLU represents a ReLU layer as the activation function, MP represents a maximum pooling layer, Reshape represents a data reshaping layer, Conv2D Transpose represents a deconvolutional layer, FC represents a fully connected layer, LSTM represents a long short-term memory network layer, and BiLSTM represents a bidirectional long short-term memory network layer; the red part marked L classification_1 and L classification_2 represents the classification loss, L reconstruction_1 and L reconstruction_2 represents the reconstruction loss, L joint Indicates joint loss.
[0073] like Figure 3 As shown, the convolutional cyclic autoencoder network uses cross entropy loss for the classification loss of the tactile array dataset, and its reconstruction loss function uses mean square error (MSE). Its classification loss function and reconstruction loss function are as follows:
[0074]
[0075]
[0076] Where N represents the number of tactile array datasets, C represents the number of categories, represents the indicator function, y i represents the label of the i-th sample, X i represents the i-th sample data, represents the output data of the i-th sample after the convolutional cyclic autoencoder network, Represents sample data X i The prediction vector after the softmax function.
[0077] The bidirectional recurrent autoencoder network uses cross entropy loss for the classification loss of the tactile color image dataset, and the reconstruction loss function uses mean square error. The classification loss and reconstruction loss are as follows:
[0078]
[0079]
[0080] Among them, M represents the number of tactile color image datasets, C represents the number of categories, represents the indicator function, y j represents the label of the jth sample, X j represents the jth sample data, represents the output data of the jth sample after passing through the bidirectional cyclic autoencoder network, Represents sample data X j The prediction vector of .
[0081] Through the network structure of the autoencoder, the latent vector that best expresses the characteristics of the tactile data set can be found. For tactile data sets with large feature differences, their latent vector dimensions may not be the same. Therefore, the method of combining the high-dimensional features of multimodal tactile data is to perform principal component analysis on the latent vector with larger dimension and reduce it to the same dimension as the latent vector with smaller dimension; that is, based on the convolutional cyclic autoencoder network, the latent vector of the tactile array data set is obtained, and based on the bidirectional cyclic autoencoder network, the latent vector of the tactile color image data set is obtained. The method of combining the high-dimensional features of multimodal tactile data performs principal component analysis on the latent vector with a dimension greater than a preset threshold, reduces it to the same dimension as the latent vector with a dimension lower than the preset threshold, calculates the mean square error of the two to obtain the joint loss, and adds the joint loss to the total loss function of the network, as follows:
[0082]
[0083] Among them, L represents the total loss, w r represents the weight of the reconstruction loss, w c Represents the weight of classification loss, w z represents the weight of the joint loss, z1 represents the latent vector of the tactile dataset whose dimension is greater than z2, z2 represents the latent vector of the tactile dataset whose dimension is less than z1, z1' represents the latent vector of z1 after principal component analysis, and the tactile dataset includes the tactile array dataset and the tactile color image dataset.
[0084] Finally, the network structure used was a dual autoencoder network, which combined the two autoencoders through the joint loss method to build a deep tumor recognition model.
[0085] Step S30: Use the training set and the validation set to train the tumor depth recognition model, and use the test set to verify the effectiveness of the trained tumor depth recognition model, wherein the effectiveness evaluation index is the tumor depth classification accuracy.
[0086] Specifically, Figure 2 As shown, the above neural network is trained using a training set and a validation set to obtain a tumor depth recognition model, wherein the tumor depth recognition model includes a convolutional training autoencoder network and a bidirectional recurrent autoencoder network; the effectiveness of the tumor depth recognition model is tested using a test set, and the evaluation index is the tumor depth classification accuracy.
[0087] The present invention aims to explore and utilize the potential relevant information of multimodal tactile data sets with large feature differences, improve the problem of difficult to effectively utilize tactile data with large feature differences in traditional classification tasks, and further enhance the mutual promotion effect of tactile data sets by entering a reasonable joint loss function, thereby improving the recognition accuracy of tumor depth. The purpose of the present invention is to efficiently extract the spatial and temporal information of each modality data by using different autoencoder network structures, introduce joint losses instead of simple splicing, and ultimately achieve high-precision tumor depth recognition.
[0088] Beneficial effects:
[0089] (1) The present invention proposes a multimodal tactile data joint perception method for tumor depth identification. A robot palpation experiment is carried out using different tactile sensors to collect a multimodal tactile dataset. The dataset is then intercepted and downsampled to finally obtain a complete and usable tactile dataset.
[0090] (2) The present invention uses convolutional neural networks and long short-term memory networks, which can not only effectively process spatial information, but also extract time series information features well. Compared with traditional data processing methods, the deep learning network used can better mine the spatiotemporal characteristics of data.
[0091] (3) The present invention uses an autoencoder network framework, which can combine multimodal data and efficiently extract the potential representation vector of the tactile data set to improve the algorithm performance.
[0092] (4) The present invention combines multimodal tactile data through a joint training method, which not only retains the effective representation of the latent vector, but also enables tactile data with large feature differences to utilize and promote each other, thereby improving the classification accuracy.
[0093] Compared with the existing tumor recognition, the present invention not only uses convolutional neural networks and long short-term memory networks to mine the spatiotemporal information of data, but also effectively extracts potential vector features by adopting an autoencoder network structure; using the idea of data union, a simple machine learning method is introduced to establish a joint loss of multimodal data, so that multimodal data with large feature differences can also promote each other, thereby improving the accuracy of deep tumor recognition. The deep tumor classification method based on the autoencoder network structure used by the present invention has been well verified on the BarrettHand tactile set and the Digit data set, and the classification accuracy has been effectively improved, proving that the method is feasible.
[0094] The multimodal data acquisition tool used in the present invention can be replaced by other tools, such as other types of tactile sensors, such as piezoelectric sensors, piezoresistive sensors, Gelsight and Gelslim sensors based on optical principles, and even visual cameras and other devices.
[0095] Furthermore, if Figure 4 As shown, based on the above-mentioned multimodal tactile data joint perception method, the present invention also provides a multimodal tactile data joint perception system accordingly, wherein the multimodal tactile data joint perception system includes:
[0096] A tactile data set acquisition module 51 is used to acquire a tactile array data set and a tactile color image data set, wherein the tactile array data set is acquired based on a capacitive sensor and is obtained after preprocessing, and the tactile color image data set is acquired based on a digit sensor and is obtained after preprocessing, and the tactile array data set and the tactile color image data set are divided into a training set, a verification set, and a test set according to a preset ratio;
[0097] A model joint construction module 52 is used to construct a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and to jointly construct the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through a joint loss to obtain a tumor depth recognition model;
[0098] The training verification module 53 is used to train the tumor depth recognition model using the training set and the verification set, and to verify the effectiveness of the trained tumor depth recognition model using the test set, wherein the effectiveness evaluation index is the tumor depth classification accuracy.
[0099] Furthermore, if Figure 3 As shown, based on the above-mentioned multimodal tactile data joint perception method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 3Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0100] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed in the terminal, such as the program code of the installation terminal, etc. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multimodal tactile data joint perception program 40 is stored on the memory 20, and the multimodal tactile data joint perception program 40 can be executed by the processor 10, thereby realizing the multimodal tactile data joint perception method in the present application.
[0101] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the multimodal tactile data joint perception method.
[0102] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.
[0103] In one embodiment, when the processor 10 executes the multimodal tactile data joint perception program 40 in the memory 20 , the steps of the multimodal tactile data joint perception method described above are implemented.
[0104] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal tactile data joint perception program, and when the multimodal tactile data joint perception program is executed by a processor, the steps of the multimodal tactile data joint perception method as described above are implemented.
[0105] In summary, the present invention provides a multimodal tactile data joint perception method and related equipment, the method comprising: acquiring a tactile array data set and a tactile color image data set, the tactile array data set is acquired based on a capacitive sensor and obtained after preprocessing, the tactile color image data set is acquired based on a digit sensor and obtained after preprocessing, and the tactile array data set and the tactile color image data set are divided into a training set, a validation set and a test set according to a preset ratio; building a convolutional recurrent autoencoder network for the tactile array data set and a bidirectional recurrent autoencoder network for the tactile color image data set, and jointly building the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through a joint loss to obtain a tumor depth recognition model; using the training set and the validation set to train the tumor depth recognition model, and using the test set to verify the effectiveness of the trained tumor depth recognition model, and the effectiveness evaluation index is the tumor depth classification accuracy. The present invention aims to explore and utilize the potential relevant information of multimodal tactile datasets with large feature differences, improve the problem that traditional classification tasks have difficulty in effectively utilizing tactile data with large feature differences, and further enhance the mutual promotion effect of tactile datasets by entering a reasonable joint loss function, thereby improving the recognition accuracy of tumor depth.
[0106] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.
[0107] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0108] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multimodal tactile data joint perception method, characterized in that: The multimodal tactile data joint perception method comprises: Acquire a tactile array dataset and a tactile color image dataset, wherein the tactile array dataset is acquired based on a capacitive sensor and obtained after preprocessing, and the tactile color image dataset is acquired based on a digit sensor and obtained after preprocessing, and divide the tactile array dataset and the tactile color image dataset into a training set, a validation set, and a test set according to a preset ratio; Building a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and jointly building the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through joint loss to obtain a tumor depth recognition model; The tumor depth recognition model is trained using the training set and the validation set, and the effectiveness of the trained tumor depth recognition model is verified using the test set, wherein the effectiveness evaluation index is the tumor depth classification accuracy; The convolutional cyclic autoencoder network uses cross entropy loss for the classification loss of the tactile array dataset, and the reconstruction loss function uses mean square error. The classification loss function and the reconstruction loss function are as follows: Where N represents the number of tactile array datasets, C represents the number of categories, represents the indicator function, y i represents the label of the i-th sample, x i represents the i-th sample data, represents the output data of the i-th sample after the convolutional cyclic autoencoder network, Represents sample data X i Prediction vector after softmax function; The bidirectional recurrent autoencoder network uses cross entropy loss for the classification loss of the tactile color image dataset, and the reconstruction loss function uses mean square error. The classification loss and reconstruction loss are as follows: Among them, M represents the number of tactile color image datasets, C represents the number of categories, represents the indicator function, y j represents the label of the jth sample, X j represents the jth sample data, represents the output data of the jth sample after passing through the bidirectional cyclic autoencoder network, Represents sample data X j The prediction vector of Based on the convolutional recurrent autoencoder network, the latent vector of the tactile array data set is obtained. Based on the bidirectional recurrent autoencoder network, the latent vector of the tactile color image data set is obtained. The method of combining the high-dimensional features of multimodal tactile data performs principal component analysis on the latent vectors with a dimension greater than a preset threshold, reduces the dimension to the same dimension as the latent vectors with a dimension less than the preset threshold, calculates the mean square error of the two to obtain the joint loss, and adds the joint loss to the total loss function of the network as follows: Among them, L represents the total loss, w r represents the weight of the reconstruction loss, w c Represents the weight of classification loss, w z represents the weight of the joint loss, z1 represents the latent vector of the tactile dataset with a dimension greater than z2, z2 represents the latent vector of the tactile dataset with a dimension less than z1, z1′ represents the latent vector of z1 after principal component analysis, and the tactile dataset includes the tactile array dataset and the tactile color image dataset.
2. The multimodal tactile data joint perception method according to claim 1, characterized in that: The capacitive sensor collects tactile array data including: Before touching, a BarrettHand robotic arm was placed at a height of 2.5 cm from the surface of the soft tissue mold, wherein the BarrettHand robotic arm includes three fingers, each of which is provided with a capacitive sensor; Controlling the BarrettHand robotic arm to press downward the surface of the soft tissue mold at a speed of 0.028 m / s; During the pressing process, the surface of the capacitive sensor provided on the finger of the BarrettHand mechanical arm is always parallel to the surface of the soft tissue, and the capacitive sensor acquires the original tactile array data and torque data; Each soft tissue sample was pressed 60 times, and after each pressing, the soft tissue sample was rotated 7.5 degrees using a rotating stage to obtain 720 sets of original tactile array data and torque data.
3. The multimodal tactile data joint perception method according to claim 2, characterized in that: The capacitive sensor collects tactile array data for preprocessing, including: For each set of original tactile array data and torque data, the torque data is only used to obtain the time point when the BarrettHand robot arm just touches the soft tissue mold surface, and this time point is selected as the starting point to intercept a fixed length of 72 time points of tactile array data; Downsampling the intercepted tactile array data, taking the starting point as the first sampling point, setting the downsampling step length to 6, and obtaining tactile array data with a length of 12 time points; The downsampled data is reshaped to convert the tactile array data at each time point into an 8×9 image size with 1 channel to the processed tactile array data set.
4. The multimodal tactile data joint perception method according to claim 1, characterized in that: The Digit sensor collects tactile color image data including: Before each compression, the UR5 robot arm was placed at a height of 2.5 cm from the surface of the soft tissue mold. The UR5 robot arm was equipped with a Digit sensor. Control the UR5 robot arm to press the surface of the soft tissue mold downward at a speed of 0.03 m / s, and stop when the soft tissue mold is pressed to a depth of 0.5 cm; During the pressing process, the surface of the Digit sensor is always pressed on the tumor position to obtain the original tactile color image data. The sampling frequency of the UR5 robot arm is set to 30fps, and the image resolution is 320×240; Each soft tissue sample was pressed 80 times, and after each pressing, the soft tissue sample was rotated 7.5 degrees using a rotating stage to obtain 960 sets of original tactile color image data.
5. The multimodal tactile data joint perception method according to claim 4, characterized in that: The Digit sensor collects tactile color image data and performs preprocessing, including: Treat each piece of original tactile color image data as a video of a pressing process, determine the starting frame in each piece of original tactile color image data where the Digit sensor just touches the soft tissue mold, and then select 10 frames of a fixed length as each piece of processed tactile color image data; Whether the soft tissue is touched is determined based on the change in pixel values in the color image. First, 50 non-contact color images are selected, and the average and standard deviation are taken according to the channel direction to obtain the average matrix and standard deviation matrix; For each current image in the original tactile color image data, the current image is first averaged in the channel direction, and then the average matrix is subtracted and compared with the 4 times standard deviation matrix. If more than 6% of the pixels are larger than the 4 times standard deviation corresponding to the current image in the standard deviation matrix, it means that the current image is in contact with the soft tissue mold. Then, the starting frame that just touched the soft tissue mold is calculated based on the time series. Finally, the captured image is normalized by a linear function to obtain the tactile color image dataset.
6. The multimodal tactile data joint perception method according to claim 1, characterized in that: The convolutional recurrent autoencoder network is constructed by a convolutional neural network and a long short-term memory network; the convolutional recurrent autoencoder network includes an encoder part and a decoder, the encoder includes a convolution layer, a batch normalization layer, a maximum pooling layer and a long short-term memory network layer; the decoder includes a long short-term memory network layer and a deconvolution layer, which is used to decode the encoded features; The bidirectional recurrent autoencoder network is constructed by a fully connected network and a bidirectional long short-term memory network; the bidirectional recurrent autoencoder network includes an encoder part and a decoder, the encoder includes a fully connected layer and two layers of bidirectional long short-term memory network layers; the decoder includes two layers of bidirectional long short-term memory network layers and a fully connected layer.
7. A multimodal tactile data joint perception system, characterized in that: The multimodal tactile data joint perception system is applied to the multimodal tactile data joint perception method according to any one of claims 1 to 6, and the multimodal tactile data joint perception system includes: a tactile data set acquisition module, used to acquire a tactile array data set and a tactile color image data set, wherein the tactile array data set is acquired based on a capacitive sensor and is obtained after preprocessing, and the tactile color image data set is acquired based on a digit sensor and is obtained after preprocessing, and the tactile array data set and the tactile color image data set are divided into a training set, a validation set, and a test set according to a preset ratio; A model joint construction module, used to construct a convolutional recurrent autoencoder network for the tactile array dataset and a bidirectional recurrent autoencoder network for the tactile color image dataset, and to jointly construct the convolutional recurrent autoencoder network and the bidirectional recurrent autoencoder network through a joint loss to obtain a tumor depth recognition model; A training and verification module is used to train the tumor depth recognition model using the training set and the verification set, and to verify the effectiveness of the trained tumor depth recognition model using the test set, wherein the effectiveness evaluation index is the tumor depth classification accuracy.
8. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multimodal tactile data joint perception program stored in the memory and executable on the processor. When the multimodal tactile data joint perception program is executed by the processor, the steps of the multimodal tactile data joint perception method as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multimodal tactile data joint perception program, and when the multimodal tactile data joint perception program is executed by a processor, the steps of the multimodal tactile data joint perception method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Remote sensing image semantic generation method based on fast region convolutional neural network
CN108960330A
Grabbed object recognition method based on tactile vibration signal and visual image fusion
CN112388655A