Electronic skin tactile perception feature extraction method based on hierarchical semantic coding
Through a hierarchical semantic coding method, the tactile information collected by MEMS electronic skin is mapped to a high-dimensional space and uniformly embedded with visual-language information, solving the problems of accurate perception and multimodal fusion, and realizing efficient abstraction of tactile information and real-time decision-making.
Patent Information
- Application Number
- CN202510766472.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies make it difficult to achieve accurate perception of 'force', 'material', 'micro-surface structure', etc. in the environment, and lack unified coding standards and feature extraction methods for alignment and fusion of multimodal information.
A hierarchical semantic coding method is used to collect tactile information through MEMS electronic skin. Combined with signal preprocessing, semantic coding network and cross-modal alignment module, the tactile information is mapped to a high-dimensional space and uniformly embedded with visual-linguistic information.
It realizes the high-dimensional semantic vector abstraction of tactile information, supports the alignment and fusion of multimodal information, has strong generalization and adaptability, and can collaborate with large models to achieve real-time decision-making.
Smart Images

Figure CN120671077A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot tactile perception, and in particular to a method for extracting tactile perception features of electronic skin based on hierarchical semantic coding. Background Art
[0002] In recent years, skin-based embodied intelligent robots have become a key frontier in robotics research. With their highly anthropomorphic appearance, natural interaction capabilities, and realistic tactile experience, these robots have found widespread application in various fields, including the tourism and entertainment industries, public government services, and healthcare and education.
[0003] The research and development of skin-based embodied intelligent robots is highly multidisciplinary, requiring the integration of expertise from multiple fields, including materials science, artificial intelligence, biomimetics, and psychology. Technological breakthroughs in this field will play a crucial role in guiding and driving the development of the entire robotics industry.
[0004] Currently, skin-based embodied intelligent robotic systems widely use vision and language as their primary channels for perception and interaction. However, accurate perception of environmental "force," "material," and "microsurface structure" remains difficult with vision and language alone. MEMS electronic skin has demonstrated high-density, high-resolution tactile acquisition capabilities in multiple studies, sensing multidimensional tactile signals such as pressure, shear, vibration, and temperature, and possessing rich physical semantic information.
[0005] Existing technologies mainly focus on inputting tactile signals into shallow CNN / LSTM networks for target recognition or material classification. A unified encoding standard has not yet been formed, and there is a lack of features that can effectively align and fuse information with large multimodal models (such as PaLM-E and Octopi).
[0006] Therefore, the present invention proposes a universal tactile semantic encoding mechanism, which processes the MEMS dot matrix data collected by the electronic skin of the robot hand into abstract features and performs semantic mapping, so that it can serve as the "third modality" input of the visual-language large model, realizing the joint perception, understanding and reasoning of scenes, objects and actions. Summary of the Invention
[0007] The technical problem to be solved by the present invention is: for the tactile information collected by the MEMS electronic skin tactile sensor of the skin-type robot, the signal preprocessing module, the hierarchical semantic encoding module and the cross-modal alignment module are combined to extract the semantic features of the tactile information, and the processed tactile semantic features are projected into a high-dimensional space unified with vision / language, so as to facilitate the subsequent fusion of vision-language-tactile multimodal information.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is: a method for extracting tactile perception features of electronic skin based on hierarchical semantic coding, comprising the following steps:
[0009] Step S1: electronic skin signal acquisition;
[0010] A high-density flexible electronic skin (16x16 MEMS array) is attached to the robot's hand or fingertips. All parameters in the data stream are configured to collect pressure, temperature, and other signals during the robot's interaction with the environment (such as grasping and sliding). The data output from each sensor element is packaged into a four-dimensional tensor, forming a sequence of basic tactile snapshots.
[0011] Step S2: tactile signal preprocessing;
[0012] First, the tactile signal is temporally aligned with the image signal and voice signal through a multimodal data synchronization mechanism. Then, all tactile signal data is debiased and normalized. Adaptive filtering (such as wavelet reconstruction) is used to remove high-frequency interference to retain the tactile signal after noise removal. Then, an event triggering mechanism (such as contact / release boundaries) is used to annotate the action cycle and output a standard data structure: [Batch, Time, X, Y, Channel].
[0013] Step S3: multi-level semantic encoding of tactile signals;
[0014] The semantic encoding network consists of a CNN module, a BiLSTM module, and a Transformer Encoder module. Tactile information passes through each module in sequence to extract different features, ultimately generating a high-dimensional tactile semantic vector. Specifically, the tactile signal with a standard data structure is first input into a local CNN module to extract the local contact structure and force distribution. This module consists of three convolutional layers and two pooling layers with a convolution kernel size of 3x3. The output after convolution is fed into the BiLSTM module for analysis of its temporal evolution. The Transformer Encoder module uses an attention mechanism to encode the output of the BiLSTM module for the entire contact segment, resulting in a final 512-dimensional tactile semantic vector.
[0015] During the training phase, three independent sub-networks were first built, each with a CNN, BiLSTM, and Transformer backbone. Each sub-network was trained independently to achieve alignment with semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot"). The CNN-based network consisted of three convolutional layers, two pooling layers, two fully connected layers, and a softmax classification layer. The BiLSTM-based network consisted of two bidirectional LSTM layers, a feature fusion layer, a fully connected layer, and a softmax classification layer. The Transformer-based network consisted of one convolutional layer, one embedding layer, two Transformer Encoder modules, and an MLP head. Then, based on these three sub-networks, an end-to-end network architecture was constructed, comprising the CNN, BiLSTM, and Transformer backbones, to fine-tune the semantic information. Finally, alignment with the semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot") was achieved by constraining the network's sigmoid-gated outputs.
[0016] Step S4: cross-modal mapping;
[0017] The tactile semantic vector shares a contrastive learning objective with the image and speech semantic vectors (e.g., CLIP-ViT models). Language labels (e.g., "cloth" and "glass") are introduced for paired triplet training. This ultimately forms a shared high-dimensional semantic space for a unified "vision+language+tactile" embedding. Specifically, a paired hard sample mining strategy is employed during the training phase. Each training sample in a paired triplet consists of an anchor (the tactile semantic vector), a positive sample (the matching image and speech semantic vectors), and a negative sample (the mismatching image and speech semantic vectors). A triplet consists of an anchor, a hard positive sample (the lower-scoring matching image and speech semantic vectors), and a hard negative sample (the higher-scoring mismatching image and speech semantic vectors). The training phase achieves consistency across modalities by minimizing the distance between the tactile semantic vector and the image and speech semantic vectors in the positive sample, and maximizing the distance between the tactile semantic vector and the image and speech semantic vectors in the negative sample. The triplet loss is expressed as:
[0018]
[0019] Among them, h t 、 and denote the tactile semantic vector, the image semantic vector in the hard positive sample, the image semantic vector in the hard negative sample, the speech semantic vector in the hard positive sample, and the speech semantic vector in the hard negative sample, respectively. α and β denote the margin hyperparameters for the visual and language modalities, respectively. λ is the language supervision weight.
[0020] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0021] 1. This invention proposes a hierarchical semantic encoder for the first time, which abstracts tactile signals step by step into interpretable high-dimensional semantic vectors to achieve human-like tactile understanding.
[0022] 2. The tactile features extracted by the present invention support multimodal alignment and embedding in a shared space, and can natively collaborate with CLIP / VLM-type models to break down modal barriers.
[0023] 3. The tactile information feature extraction method proposed in this invention has strong generalization and is adaptable to various electronic skin architectures, signal dimensions, and interaction tasks;
[0024] 4. The tactile information features extracted by the present invention support end-to-end deployment and can be directly applied to large models of vision-language-tactile multimodal information fusion to achieve real-time training and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0026] Figure 1 This is the overall block diagram of the electronic skin tactile perception feature extraction based on hierarchical semantic coding of the present invention.
[0027] Figure 2 This is a multi-level semantic encoding flow chart of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions of the present invention with reference to the embodiments. Obviously, the embodiments described are only some of the embodiments of the present invention, rather than all of them. The embodiments of the present invention and all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] Example 1
[0030] The present invention provides a method for extracting tactile perception features of electronic skin based on hierarchical semantic coding, which includes the following steps:
[0031] Step S1: electronic skin signal acquisition;
[0032] A high-density flexible electronic skin (16x16 MEMS array) is attached to the robot's hand or fingertips. All parameters in the data stream are configured to collect pressure, temperature, and other signals during the robot's interaction with the environment (such as grasping and sliding). The data output from each sensor element is packaged into a four-dimensional tensor, forming a sequence of basic tactile snapshots.
[0033] Step S2: tactile signal preprocessing;
[0034] First, the tactile signal is temporally aligned with the image signal and voice signal through a multimodal data synchronization mechanism. Then, all tactile signal data is debiased and normalized. Adaptive filtering (such as wavelet reconstruction) is used to remove high-frequency interference to retain the tactile signal after noise removal. Then, an event triggering mechanism (such as contact / release boundaries) is used to annotate the action cycle and output a standard data structure: [Batch, Time, X, Y, Channel].
[0035] Step S3: multi-level semantic encoding of tactile signals;
[0036] The semantic encoding network consists of a CNN module, a BiLSTM module, and a Transformer Encoder module. Tactile information passes through each module in sequence to extract different features, ultimately generating a high-dimensional tactile semantic vector. Specifically, the tactile signal with a standard data structure is first input into a local CNN module to extract the local contact structure and force distribution. This module consists of three convolutional layers and two pooling layers with a convolution kernel size of 3x3. The output after convolution is fed into the BiLSTM module for analysis of its temporal evolution. The Transformer Encoder module uses an attention mechanism to encode the output of the BiLSTM module for the entire contact segment, resulting in a final 512-dimensional tactile semantic vector.
[0037] During the training phase, three independent sub-networks were first built, each with a CNN, BiLSTM, and Transformer backbone. Each sub-network was trained independently to achieve alignment with semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot"). The CNN-based network consisted of three convolutional layers, two pooling layers, two fully connected layers, and a softmax classification layer. The BiLSTM-based network consisted of two bidirectional LSTM layers, a feature fusion layer, a fully connected layer, and a softmax classification layer. The Transformer-based network consisted of one convolutional layer, one embedding layer, two Transformer Encoder modules, and an MLP head. Then, based on these three sub-networks, an end-to-end network architecture was constructed, comprising the CNN, BiLSTM, and Transformer backbones, to fine-tune the semantic information. Finally, alignment with the semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot") was achieved by constraining the network's sigmoid-gated outputs.
[0038] Step S4: cross-modal mapping;
[0039] The tactile semantic vector shares a contrastive learning objective with the image and speech semantic vectors (e.g., CLIP-ViT models). Language labels (e.g., "cloth" and "glass") are introduced for paired triplet training. This ultimately forms a shared high-dimensional semantic space for a unified "vision+language+tactile" embedding. Specifically, a paired hard sample mining strategy is employed during the training phase. Each training sample in a paired triplet consists of an anchor (the tactile semantic vector), a positive sample (the matching image and speech semantic vectors), and a negative sample (the mismatching image and speech semantic vectors). A triplet consists of an anchor, a hard positive sample (the lower-scoring matching image and speech semantic vectors), and a hard negative sample (the higher-scoring mismatching image and speech semantic vectors). The training phase achieves consistency across modalities by minimizing the distance between the tactile semantic vector and the image and speech semantic vectors in the positive sample, and maximizing the distance between the tactile semantic vector and the image and speech semantic vectors in the negative sample. The triplet loss is expressed as:
[0040]
[0041] Among them, h t 、 and denote the tactile semantic vector, the image semantic vector in the hard positive sample, the image semantic vector in the hard negative sample, the speech semantic vector in the hard positive sample, and the speech semantic vector in the hard negative sample, respectively. α and β denote the margin hyperparameters for the visual and language modalities, respectively. λ is the language supervision weight.
[0042] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for extracting tactile perception features of electronic skin based on hierarchical semantic coding, characterized in that: The following steps are involved: Step S1: electronic skin signal acquisition; A high-density flexible electronic skin (16x16 MEMS array) is attached to the robot's hand or fingertips. All parameters in the data stream are configured to collect pressure, temperature, and other signals during the robot's interaction with the environment (such as grasping and sliding). The data output from each sensor element is packaged into a four-dimensional tensor, forming a sequence of basic tactile snapshots. Step S2: tactile signal preprocessing; First, the tactile signal is temporally aligned with the image signal and voice signal through a multimodal data synchronization mechanism. Then, all tactile signal data is debiased and normalized. Adaptive filtering (such as wavelet reconstruction) is used to remove high-frequency interference to retain the tactile signal after noise removal. Then, an event triggering mechanism (such as contact / release boundaries) is used to annotate the action cycle and output a standard data structure: [Batch, Time, X, Y, Channel]. Step S3: multi-level semantic encoding of tactile signals; The semantic encoding network consists of a CNN module, a BiLSTM module, and a Transformer Encoder module. Tactile information passes through each module in sequence to extract different features, ultimately generating a high-dimensional tactile semantic vector. Specifically, the tactile signal with a standard data structure is first input into a local CNN module to extract the local contact structure and force distribution. This module consists of three convolutional layers and two pooling layers with a convolution kernel size of 3x3. The output after convolution is fed into the BiLSTM module for analysis of its temporal evolution. The Transformer Encoder module uses an attention mechanism to encode the output of the BiLSTM module for the entire contact segment, resulting in a final 512-dimensional tactile semantic vector. During the training phase, three independent sub-networks were first built, each with a CNN, BiLSTM, and Transformer backbone. Each sub-network was trained independently to achieve alignment with semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot"). The CNN-based network consisted of three convolutional layers, two pooling layers, two fully connected layers, and a softmax classification layer. The BiLSTM-based network consisted of two bidirectional LSTM layers, a feature fusion layer, a fully connected layer, and a softmax classification layer. The Transformer-based network consisted of one convolutional layer, one embedding layer, two Transformer Encoder modules, and an MLP head. Then, based on these three sub-networks, an end-to-end network architecture was constructed, comprising the CNN, BiLSTM, and Transformer backbones, to fine-tune the semantic information. Finally, alignment with the semantic labels (e.g., 30 tactile features such as "slippery," "thorny," and "hot") was achieved by constraining the network's sigmoid-gated outputs. Step S4: cross-modal mapping; The tactile semantic vector shares a contrastive learning objective with the image and speech semantic vectors (e.g., CLIP-ViT models). By introducing language labels (e.g., "cloth" and "glass"), paired triplet loss training is performed. Ultimately, a shared high-dimensional semantic space is formed, representing a unified "vision+language+tactile" embedding. Specifically, a paired hard sample mining strategy is employed during the training phase. Each training sample in a paired triplet consists of an anchor (the tactile semantic vector), a positive sample (the matching image and speech semantic vectors), and a negative sample (the mismatching image and speech semantic vectors). A triplet consists of an anchor, a hard positive sample (the lower-scoring matching image and speech semantic vectors), and a hard negative sample (the higher-scoring mismatching image and speech semantic vectors). The training phase achieves consistency across modalities by minimizing the distance between the tactile semantic vector and the image and speech semantic vectors in the positive sample, and maximizing the distance between the tactile semantic vector and the image and speech semantic vectors in the negative sample. The triplet loss is expressed as: Among them, h t 、 and denote the tactile semantic vector, the image semantic vector in the hard positive sample, the image semantic vector in the hard negative sample, the speech semantic vector in the hard positive sample, and the speech semantic vector in the hard negative sample, respectively. α and β denote the margin hyperparameters for the visual and language modalities, respectively. λ is the language supervision weight.
Citation Information
Cited By
Data compression and transmission method for electronic skin of humanoid robot
CN121814856A