A tactile sensor, electronic skin, and robot
By covering the visual-tactile sensor with a magnetic thin film and combining it with a magnetic graphics card and a light source, magnetic sensing and imaging data are collected simultaneously. By using a multimodal pressure detection model for analysis, the problems of detection sensitivity and comprehensiveness in existing technologies are solved, and high-precision detection of forces on the top and side surfaces is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-10
AI Technical Summary
Existing visual-tactile sensor structures increase the overall thickness of the magnetic film by setting a surface coating on the transparent silicone surface, thereby increasing the sensitivity of influence detection. Furthermore, planar structures cannot detect lateral forces.
A magnetic thin film is used to cover the top and side areas. Combined with a magnetic graphics card and a light source, magnetic sensing data and imaging data are collected simultaneously by a visual-tactile sensor. The data are then analyzed using a multimodal pressure detection model to achieve high-precision detection of the stress on the magnetic thin film.
It improves the sensitivity and comprehensiveness of magnetic thin film deformation detection, enabling simultaneous detection of forces on the top and side surfaces, and achieving high-precision multimodal information acquisition.
Smart Images

Figure CN121347008B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of sensors, in particular a visual tactile sensor, electronic skin and robot. BACKGROUND
[0002] Electronic skin is a kind of flexible electronic device that simulates the sensing function of human skin. Its core working principle is to capture external stimulus signals through sensors, convert and optimize them through signal processing modules, and finally output recognizable signals to achieve accurate detection of physical quantities such as pressure, temperature, humidity, and strain. Some electronic skins even have stimulation feedback function.
[0003] In recent years, with the development of artificial intelligence and robot technology, electronic skin that can accurately detect physical quantities such as pressure, temperature, humidity, and strain has been increasingly valued and shown broad application prospects. In particular, the sensing method based on visual tactile sensor combined with magnetic induction solves the problems of missing 3D position information reception of marker points, limited camera frequency and response rate, and lack of state information sensing ability of non-contact objects for visual tactile sensor in the case of monocular camera, realizing high-precision force measurement, high response, and proximity sensing function of visual tactile sensor.
[0004] However, such sensor structure sets a surface plating layer on the surface of transparent silicone to isolate external light sources, and reconstructs the surface texture of the sensed object, and receives the deformation information of the surface plating layer through the camera to realize information collection. The presence of the surface plating layer increases the overall thickness of the magnetic film and affects the sensitivity of force detection.
[0005] In addition, the magnetic film of the known electronic skin structure is usually a planar structure, and when contact occurs on the side of the electronic skin, the force acting thereon cannot be detected. SUMMARY
[0006] The purpose of the present application is to provide a visual tactile sensor, electronic skin and robot to improve the comprehensiveness and sensitivity of force detection in response to all or part of the above problems.
[0007] The technical solution adopted by the present application is as follows:
[0008] A visual tactile sensor comprises:
[0009] a magnetic film covering the top surface area and the side surface area;
[0010] a support supporting the magnetic film;
[0011] a magnetic display card installed on the support and responding to the force of the magnetic film to develop;
[0012] a light source installed in the bracket;
[0013] a visual-tactile collector collecting magnetic sensing data of the magnetic film under stress and developing data of the magnetic card.
[0014] Further, the visual-tactile collector comprises a mounting plate, and an image collecting unit and a magnetic flux collecting unit arranged on the mounting plate, the image collecting unit collects the developing data, and the magnetic flux collecting unit collects the magnetic sensing data.
[0015] Further, the magnetic film is made of magnetic powder and silicone rubber.
[0016] Further, the mixing mass ratio of the magnetic powder and the silicone rubber is 0.8:1-1.2:1.
[0017] Further, the stress detection and analysis method of the visual-tactile sensor comprises:
[0018] inputting the magnetic sensing data and the developing data into a trained multi-modal pressure detection model, and outputting a stress position and / or a stress size from the multi-modal pressure detection model;
[0019] The multi-modal pressure detection model is configured to:
[0020] respectively pre-process the input magnetic sensing data and developing data;
[0021] extracting magnetic sensing features from the pre-processed magnetic sensing data and extracting developing features from the pre-processed developing data;
[0022] fusing the magnetic sensing features and the developing features to obtain fused features;
[0023] reasoning the stress position and / or the stress size according to the fused features.
[0024] Further, extracting the magnetic sensing features from the pre-processed magnetic sensing data comprises:
[0025] using a plurality of full connection layers to perform feature extraction on the pre-processed magnetic sensing data layer by layer, wherein each full connection layer respectively performs dimension expansion, activation operation and regularization processing on the input data.
[0026] Further, extracting the developing features from the pre-processed developing data comprises:
[0027] using a plurality of convolution layers to perform feature extraction on the pre-processed developing data layer by layer, wherein each convolution layer respectively performs dimension expansion on the input data.
[0028] The features after convolution processing are subjected to activation processing and pooling processing to obtain the developed features.
[0029] Further, the magnetic sensing features and the developed features are fused, including:
[0030] The magnetic sensing features and the developed features are spliced.
[0031] The spliced features are input into an SE module for attention fusion.
[0032] The features output by the SE module are subjected to dimension compression by a fully connected layer to obtain the fused features.
[0033] The application further provides an electronic skin, which comprises a skin body and at least one above-mentioned visual-tactile sensor arranged on the skin body.
[0034] The application further provides a robot, which is provided with the above-mentioned electronic skin on the surface.
[0035] As described above, due to the adoption of the above technical solutions, the application has the following beneficial effects:
[0036] The electronic skin structure provided in the application can effectively improve the deformation detection sensitivity of the magnetic thin film by developing the deformation information of the magnetic thin film under stress through the magnetic development card. The high-precision and multi-modal information acquisition of the stress condition of the electronic skin can be realized by synchronously collecting the magnetic sensing data (magnetic flux change information) and the visual development data (development image) of the magnetic thin film under stress. The magnetic thin film covers the top surface area and the side surface area of the sensor. Whether the stress is on the top surface or the side surface, the magnetic sensing data and the development data will change accordingly, so that the forward stress and the lateral stress of the electronic skin can be comprehensively detected. Through the analysis of the multi-modal pressure detection model, the high-precision inference of the stress position and the stress size can be realized at the same time. BRIEF DESCRIPTION OF DRAWINGS
[0037] The application will be described by way of example and with reference to the accompanying drawings, in which:
[0038] Figure 1 is a forward cross-sectional view of the visual-tactile sensor in one embodiment.
[0039] Figure 2 is an exploded view of the visual-tactile sensor in one embodiment.
[0040] Figure 3 is a diagonal cross-sectional view of the tactile sensor in one embodiment.
[0041] Figure 4 is a data flow diagram of the stress detection and analysis method of the visual-tactile sensor in one embodiment.
[0042] Figure 5 is a hierarchical structure diagram of a multi-modal pressure detection model in one embodiment.
[0043] Figure 6 is a network architecture diagram of a multi-modal pressure detection model in one embodiment.
[0044] Figure 7 is a flowchart of a force detection analysis method of a visual-tactile sensor in one embodiment.
[0045] In the figure, 1 is a magnetic film, 2 is a support, 3 is a magnetic display card, 4 is a light source, 5 is a visual-tactile collector, 6 is a base, 11 is a top surface area, 12 is a side surface area, 13 is a edge, 21 is a horizontal support, 22 is a vertical support, 51 is a mounting plate, 52 is an image collection unit, 53 is a magnetic flux collection unit, 61 is a base, and 62 is a connecting piece. DETAILED DESCRIPTION
[0046] All features disclosed in this specification, and / or all steps of any methods or processes disclosed in this specification, can be combined in any manner, except where features and / or steps are mutually exclusive.
[0047] Any feature disclosed in this specification, unless stated otherwise, can be replaced by alternative features or equivalents having the same or a similar effect. That is, unless stated otherwise, each feature is one example only of a number of alternative or similar features.
[0048] The existing visual-tactile sensor structure achieves information collection by setting a surface plating layer for isolating external light sources and reconstructing the surface texture of the perceived object on the transparent silicone surface, and receiving deformation information of the surface plating layer through a camera. However, it has the problems of increasing the overall thickness of the magnetic film and low sensitivity of force detection. In addition, the existing electronic skin structure is generally a planar structure, and when the contact position occurs on the side surface, the side surface force cannot be detected.
[0049] In view of the above problems, the embodiments of the present application provide a visual-tactile sensor, an electronic skin and a robot, aiming to improve the sensitivity and comprehensiveness of pressure (contact force) detection.
[0050] In a feasible embodiment, the present application provides a visual-tactile sensor, which includes a magnetic film 1, a support 2, a magnetic display card 3, a light source 4, and a visual-tactile collector 5.
[0051] The magnetic film 1 is located at the outermost layer of the visual tactile sensor, and serves as a component for sensing external force. The magnetic film 1 covers at least the top surface area 11 and the side surface area 12 of the visual tactile sensor to receive the force applied to the top surface and the side surface. The magnetic film 1 has magnetism, and deformation of the magnetic film 1 will cause a change in magnetic flux.
[0052] As an optional embodiment, the magnetic film 1 is made of magnetic powder mixed with silicone rubber. That is, the magnetic powder is mixed in the silicone rubber film instead of being coated on the surface of the silicone rubber film, which can effectively reduce the thickness of the film layer and improve the sensitivity of force detection. For example, by designing the magnetic film 1 in this way, the average thickness of the magnetic film can be configured to be 2mm-5mm, which has a higher deformation response sensitivity.
[0053] In a feasible embodiment, the mass ratio of the magnetic powder and the silicone rubber is 0.8:1-1.2:1. Specifically, the magnetic film 1 with magnetism is obtained after mixing, stirring, molding, and magnetizing the magnetic powder and the silicone rubber according to a mass ratio of 0.8:1-1.2:1. The silicone rubber is obtained by mixing platinum A reagent and platinum B reagent according to a mass ratio of 1:1.
[0054] In a specific implementation, the average surface magnetic field of the magnetic film 1 is configured to be 10mT-30mT, which can realize accurate detection of the force condition.
[0055] The magnetic display card 3 can adopt a magneto-optical display film, which has the characteristics of high sensitivity (up to 10 -8 Wb order of magnitude) and strong anti-interference, can perform magnetic flux distribution imaging, can detect weak magnetic fields, and thus can detect weak force conditions.
[0056] The support 2 is used to support the magnetic film 1, so that the magnetic film 1 can have a stable shape / state when not subjected to force, and can return to the shape / state after the force is removed.
[0057] The magnetic display card 3 is installed on the support 2, such as in the support 2, and is adjacent to the magnetic film 1. If the magnetic film 1 deforms, the image of the magnetic display card 3 will change accordingly. Thus, the magnetic display card 3 will develop in response to the force of the magnetic film 1, reflecting the deformation information of the magnetic film 1.
[0058] As a preferred embodiment, the support 2 includes a horizontal support 21 supporting the top surface area 11 of the magnetic film 1, and a vertical support 22 supporting the side surface area 12 of the magnetic film 1. As shown in Figure 2 If the magnetic film 1 covers the top surface area 11 and four side surface areas 12, the support 2 includes a horizontal support 21 supporting the top surface and four vertical supports 22 supporting the four side surface areas 12, Figure 2The image shows a visible horizontal bracket 21 and two vertical brackets 22. The magnetic graphics card 3 is mounted parallel to the horizontal bracket 21, for example, mounted on the bottom surface of the horizontal bracket 21.
[0059] The light source 4 is installed inside the bracket 2. The light emitted by the light source 4 can illuminate the magnetic graphics card 3, thereby facilitating the acquisition of the image of the magnetic graphics card 3, i.e., the development data.
[0060] The visual-tactile sensor 5 collects magnetic flux data (i.e., magnetic flux change information) of the magnetic thin film and development data of the magnetic display 3, respectively. This can be achieved through two separate acquisition units; for example, magnetic flux acquisition unit 53 can collect magnetic flux data, while image acquisition unit 52 can collect development data. The image acquisition unit 52, which collects development data, faces the magnetic display 3. There can be one or more magnetic flux acquisition units 53, located within the magnetic field range of the magnetic thin film 1, such as within the space enclosed by the magnetic thin film 1.
[0061] In a preferred embodiment, the light source 4 is located behind the photosensitive element of the image acquisition unit 52. In this way, the image acquisition unit 52 only acquires the image of the magnetic graphics card 3 and does not acquire the image of the light source 4, thereby effectively reducing the influence of light on the accuracy of the acquisition of the color, outline, etc. of the image of the magnetic graphics card 3.
[0062] The visual-tactile sensor can be mounted and fixed using the base 6. That is, other components are set on the base 6. Specifically, the visual-tactile sensor 5 and the bracket 2 are mounted on the base 6, and other components can be supported by the visual-tactile sensor 5 or the bracket 2.
[0063] like Figures 1-3 As shown, in one feasible embodiment, the base 6 includes two parts: a base 61 and a connector 62. The base 61 serves as the foundation for mounting the visual-tactile sensor, and the connector 62 is connected to the base 61, for example, via plastic bolts. The connector 62 is used to mount the visual-tactile sensor 5 and the bracket 2, and also to lock the magnetic film 1, keeping it taut. Wherein, as... Figure 3 As shown, the connector 62, the visual-tactile sensor 5, and the bracket 2 are connected by bolts passing through them in sequence. The bottom edge 13 of the magnetic film 1 is sandwiched between the connector 62 and the base 61. If the base 6 is omitted, the magnetic film 1 can be locked to the bracket 2, for example, by locking the edge 13 of the magnetic film 1 to the bottom of the bracket 2 with plastic bolts.
[0064] In one optional embodiment, the visual-tactile sensor 5 includes a mounting plate 51, and an image acquisition unit 52 and a magnetic flux acquisition unit 53 disposed on the mounting plate 51. Figures 1-3As shown, the mounting plate 51 is, for example, a PCB board on which a plurality of (4 in the figure, but only an example, which can be adjusted according to actual conditions) Hall sensors are mounted as the magnetic flux collection unit 53, and the Hall sensor is, for example, a three-axis force Hall sensor that collects magnetic flux change data in three dimensions. In the middle of the plurality of Hall sensors, one (one is an example, which can be adjusted according to actual conditions) depth camera is installed as the image collection unit 52, and the depth camera is, for example, a monocular camera.
[0065] The light strip is used as the light source 4, and the light strip is arranged around the image collection unit 52. For example, in the above embodiment in which the depth camera is surrounded by a plurality of Hall sensors, the light strip is arranged to surround each Hall sensor to improve the uniformity of the light illumination and enhance the imaging effect. In this way, the light source 4 can be arranged on the mounting plate 51.
[0066] The visual-tactile sensor with the above structure can synchronously collect the magnetic sensing data and the imaging data generated by the deformation of the magnetic film 1 under the action of an external force. The magnetic film does not need to be coated with a film layer and has a relatively thin thickness, and has higher response sensitivity to the force. In this way, the accuracy of the visual-tactile sensor in detecting the force can be improved. The magnetic film 1 covers the top surface area 11 and the side surface area 12, the magnetic imaging card 3 can respond to the normal force and the lateral force, and the visual-tactile collector 5 can synchronously collect the magnetic sensing data and the imaging data when the normal force or the lateral force is applied, so as to realize comprehensive, sensitive, multi-modal and high-precision sensing data collection of the normal force or the lateral force.
[0067] On the basis of the visual-tactile sensor in the above embodiment, the stress detection and analysis method of the embodiment of the present application is also introduced, and the method comprises:
[0068] The magnetic sensing data and the imaging data are input into the trained multi-modal pressure detection model, and the stress position and / or stress size are output by the multi-modal pressure detection model.
[0069] Referring to Figures 4-7 , the multi-modal pressure detection model is configured to:
[0070] (1) The input magnetic sensing data and imaging data are preprocessed respectively.
[0071] The magnetic sensing data and the imaging data belong to different modal sensing data, and different preprocessing methods need to be used. The preprocessing of the input data is performed in the input layer of the multi-modal pressure detection model.
[0072] As shown in Figure 4 , Figure 5 , in the input layer, for the magnetic sensing data, the magnetic sensing data is standardized to convert the magnetic sensing data into a range that can be processed by the model. For example, the numerical value of the magnetic sensing data is normalized to the range of [0, 1].
[0073] For the developed data, its resolution is controlled to a preset size, such as 64 pixels × 64 pixels, through scaling or cropping; then, the developed data is converted to grayscale; finally, the pixel values of the developed data are normalized to the range [0,1] to facilitate model processing. Additionally, in an optional implementation, sample data augmentation can be performed on the developed data samples to expand the training samples. For example, the developed data samples can be processed through rotation, flipping, masking operations, etc., to obtain new samples, which are then preprocessed.
[0074] (2) Extract magnetic features from the preprocessed magnetic data and extract development features from the preprocessed development data.
[0075] like Figure 4 , Figure 5 As shown, the multimodal pressure detection model uses a feature extraction layer to extract magnetic and imaging features respectively.
[0076] In one alternative implementation, extracting magnetic features from the preprocessed magnetic field data includes:
[0077] A multi-layer fully connected approach is used to extract features from the preprocessed magnetic field data, obtaining magnetic field features. Each fully connected layer performs dimensionality expansion, activation operations, and regularization on the input data.
[0078] For example, assuming a visual-haptic sensor contains 5 Hall sensors, the magnetic data it collects is 1×5×3=1×15 dimensional, therefore the input dimension of the first fully connected layer is 15. Figure 4 As shown, the multi-layer first fully connected network of the multimodal pressure detection model uses four fully connected layers to extract features from the preprocessed magnetic field data. The output dimensions of the four fully connected layers are 15-64-128-256, respectively. The ReLU function is used as the activation function for the fully connected layers, and Dropout is used for regularization after each fully connected layer. Thus, after feature extraction by the four fully connected layers, the multi-layer first fully connected network outputs the corresponding magnetic field features. These magnetic field features reflect the mapping relationship between the change in magnetic field and the location / magnitude of the force when force is applied.
[0079] like Figure 4 , Figure 5 As shown, in one optional implementation, extracting development features from the preprocessed development data includes:
[0080] Feature extraction is performed on the preprocessed development data layer by layer using multiple convolutional layers, where each convolutional layer expands the dimensions of the input data.
[0081] like Figure 5As shown, a lightweight CNN (Convolutional Neural Network) (CNN Lite), such as MobileNet V2, is used to extract imaging features in a multimodal stress detection model. This lightweight CNN consists of three convolutional layers, each with a 3×3 kernel. The imaging data acquired by the image acquisition unit 52 has depth information, therefore the imaging data has three dimensions: length, width, and height. Thus, the input dimension of the first convolutional layer is 3. The output dimensions of the three convolutional layers are 16-32-64, respectively.
[0082] The convolutional features are then activated and pooled to obtain the development features. In the development feature extraction branch, the convolutional features are activated using the ReLU function, followed by max pooling and global average pooling to obtain 256-dimensional development features. These development features reflect the mapping relationship between the changes in the development texture of the magnetic graphics card 3 under stress and the location / magnitude of the stress.
[0083] (3) Combine magnetic induction features and imaging features to obtain fused features.
[0084] As an alternative implementation, the multimodal pressure detection model employs a cross-modal fusion layer to fuse magnetic and imaging features.
[0085] like Figure 4 , Figure 5 As shown, the cross-modal fusion layer first splices the magnetic sensing features and the development features. Taking the previous implementation as an example, the 256-dimensional magnetic sensing features and the 256-dimensional development features are spliced into a 512-dimensional spliced feature.
[0086] The cross-modal fusion layer then uses the SE module to perform attention fusion on the spliced features. The SE module dynamically assigns fusion weights to the magnetic and imaging features, and the weighted fusion yields a 512-dimensional vector.
[0087] The cross-modal fusion layer also employs a fully connected layer to compress the dimensionality of the features output by the SE module, resulting in fused features. Specifically, based on the 512-dimensional vector obtained by the weighted fusion above, a fully connected layer is used to compress it into a 256-dimensional fused feature, which is then used as the output of the cross-modal fusion layer.
[0088] (4) Based on the fusion characteristics, infer the location and / or magnitude of the force.
[0089] As an optional implementation method, such as Figure 4 , Figure 5 As shown, the multimodal stress detection model uses the output layer to infer the fused features.
[0090] According to different inference purposes, the number of channels of the output layer is different. For example, if only the force size needs to be inferred, the output channel is 1; if the force position and the force size need to be inferred simultaneously, and the force position is represented by two-dimensional coordinates, the output channel is 3, which outputs the x coordinate, the y coordinate and the force size, respectively. The output of the three channels can be simplified as (x, y, N), where x and y represent the x coordinate and the y coordinate, respectively, and N represents the value of the force size.
[0091] As shown in Figure 5 , specifically, the output layer first uses two fully connected layers (second fully connected network) for dimension compression, compressing the 256-dimensional fusion features to 128 dimensions, and then compressing the 128-dimensional features to a 3 (output channel number)-dimensional feature vector for output. The output result of the output layer is a continuous value, where the range of x is [0, W], the range of y is [0, H], and the range of N is [0, F_max], W and H are the length and width dimensions of the tactile sensor, and F_max is the maximum pressure.
[0092] To more clearly illustrate the working process of the multi-modal pressure detection model proposed in the present application, as shown in Figure 6 , the construction of the multi-modal pressure detection model is exemplarily illustrated in the embodiments of the present application.
[0093] As shown in Figure 6 , the multi-modal pressure detection model as a whole includes four layers of structures, namely, an input layer, a feature extraction layer, a cross-modal feature fusion layer and an output layer. The input layer is responsible for pre-processing the input data; the feature extraction layer is responsible for feature extraction of the multi-modal data; the cross-modal feature fusion layer is responsible for fusion of the multi-modal features; and the output layer is responsible for task inference based on the fusion features.
[0094] The input layer receives two modal data, i.e., magnetic induction data and imaging data; and pre-processes the magnetic induction data and the imaging data, respectively. The magnetic induction data is controlled within a required range by subtracting the mean value or dividing by the standard deviation. For example, the magnetic induction data is normalized within [0, 1]. For the imaging data, scaling and / or cropping can be performed to control the resolution to a preset size, for example, 64 pixels x 64 pixels; the imaging data is then subjected to grayscale processing, and finally the pixel value is also normalized within [0, 1]. In addition, for the imaging data, rotation, flipping, mosaic addition and the like can be performed to enhance the training sample set and expand the training samples.
[0095] The preprocessed magnetic sensing data and the developed data are input into a feature extraction layer for multi-modal feature extraction. The feature extraction layer extracts magnetic sensing features by using a first fully connected network on the preprocessed magnetic sensing data, and extracts developed features by using a lightweight CNN on the preprocessed developed data. The structure of the first fully connected network can refer to the previous embodiment, which will not be described here. The lightweight CNN performs multi-layer convolution processing on the input developed data through multiple convolution layers, then performs batch normalization processing through a BN (Batch Normalization) layer, performs activation processing through a ReLu function, and finally performs pooling processing through a pooling layer. The pooling processing can be maximum pooling and global average pooling.
[0096] The magnetic sensing features and the developed features extracted by the feature extraction layer are input into a cross-modal fusion layer, which fuses the magnetic sensing features and the developed features. Specifically, the cross-modal fusion layer first concatenates the magnetic sensing features and the developed features to form 512-dimensional concatenated features from 256-dimensional magnetic sensing features and 256-dimensional developed features. Then, the SE module is used to perform attention weighting on the magnetic sensing features and the developed features in the concatenated features. Finally, a fully connected layer is used to compress the dimension of the weighted concatenated features from 512 dimensions to 256 dimensions to obtain the fused features.
[0097] The fused features are finally input into an output layer for task reasoning. The output layer uses a second fully connected network to reduce the dimension of the fused features to a dimension corresponding to the number of task channels. Taking reasoning of the force position and the force size as an example, the second fully connected network uses two fully connected layers to complete the reasoning of the fused features to the output task according to the dimension compression rule of 256-128-3 dimensions. The three output channels are the x-coordinate of the force position, the y-coordinate of the force position, and the force size.
[0098] The multi-modal pressure detection model needs to be trained before it can accurately reason the force position and / or the force size. In an optional implementation, the training method of the multi-modal pressure detection model includes:
[0099] (1) Obtain a training sample set.
[0100] The facility capable of accurately applying pressure such as a controllable mechanical arm continuously presses on the magnetic film 1 of the visual tactile sensor, and the mechanical arm is taken as an example. The magnetic induction data and the developed data when pressing are obtained respectively, as well as the force position and / or force size of the mechanical arm pressing (determined according to the final detection purpose), and the magnetic induction data, the developed data, and the force position and / or force size are associated as a set of training sample data. By controlling the mechanical arm to press with different forces at different positions, a plurality of sets of training sample data are obtained. In addition, the developed data can also be subjected to sample enhancement processing such as flipping, rotating, mask processing (i.e. increasing mosaic), and the like, and as the corresponding new developed data under the current conditions, together as training sample data. All the training sample data are combined to obtain a training sample set. For example, assuming that the force position and the force size need to be detected at the same time, the training sample set at least contains (magnetic induction data, developed data, force position, force size), wherein the force position and the force size are label values corresponding to the magnetic induction data and the developed data.
[0101] (2) Model training.
[0102] The training sample set is randomly divided into a training set and a validation set according to a certain proportion, the training set is sequentially input to the initialized multi-modal pressure detection model for iterative training, and the corresponding predicted force position and force size are obtained every iteration. And based on the label value, the corresponding model loss is calculated by the weighted MSE loss function. Wherein, the weighted MSE loss function is designed as:
[0103] ;
[0104] Wherein, Loss represents the model loss, represents the weight of the i-th channel, and respectively represent the predicted value and the label value of the j-th training sample data in the i-th channel, and m represents the number of training sample data.
[0105] The weighted MSE loss function assigns weights to the three output channels x, y, and N respectively, and the weights are set in advance according to the priority of the detection target / task.
[0106] The stop condition of model training is:
[0107] 1) The model loss on the validation set decreases by less than a preset value for K consecutive rounds (K is a positive integer, set in advance). The preset value is, for example, .
[0108] 2) The model loss on the validation set is reduced to below a preset threshold. For example, the pixel-level loss corresponding to the force position is <0.1, and the force size loss is <0.05.
[0109] When the stop condition is met, the model parameters are saved, and the trained multi-modal pressure detection model can be used to detect the force position and force size according to the magnetic sensing data and the imaging data.
[0110] The visual tactile sensor provided by the above embodiment can perceive force through the magnetic film 1, does not need to be coated with a film layer, and can effectively improve the sensitivity of force perception. In addition, the magnetic film 1 covers the top surface area 11 and the side surface area 12, and integrates the magnetic flux collecting unit 53 and the image collecting unit 52 into the visual tactile collector 5 to simultaneously collect magnetic sensing data and imaging data, and the two are combined for force detection analysis, and the visual tactile collector 5 has extremely high detection accuracy for top surface force and side surface force. The designed multi-modal pressure detection model extracts magnetic sensing features and imaging features through two branches, and then performs attention weighted fusion, and performs multi-task reasoning according to the fused features, which can effectively improve the accuracy, richness and comprehensiveness of task detection.
[0111] According to the idea of the present application, the electronic skin includes a skin body and at least one visual tactile sensor arranged on the skin body. By dispersing / centralizing the visual tactile sensor on the skin body, the force position and / or force size of the force acting on each visual tactile sensor can be detected to accurately detect the force position and / or force size on the electronic skin.
[0112] In addition, the present application also provides a robot, and the robot is provided with the above-mentioned electronic skin. The robot can be a mechanical arm or a humanoid robot, and the electronic skin can be arranged on the position where force detection is needed.
[0113] The present application is not limited to the foregoing specific embodiments. The present application extends to any new feature or any new combination disclosed in the specification, and any new method or process steps or any new combination disclosed.
Claims
1. A visual-tactile sensor characterized by, The application relates to a visual-tactile sensor, which comprises the following parts: a magnetic film covering a top surface area and a side surface area; a support supporting the magnetic film; a magnetic display card installed on the support and displaying in response to force of the magnetic film; a light source installed in the support; a visual-tactile collector collecting magnetic induction data of the force of the magnetic film and display data of the magnetic display card; the visual-tactile collector comprises a mounting plate, and an image collecting unit and a magnetic flux collecting unit arranged on the mounting plate, wherein the image collecting unit collects the display data, and the magnetic flux collecting unit collects the magnetic induction data.
2. The visuo-tactile sensor of claim 1, wherein, The magnetic film is made of magnetic powder and silicone rubber.
3. The visuo-tactile sensor of claim 2, wherein, The mass ratio of the magnetic powder and the silicone rubber is 0.8:1-1.2:
1.
4. The tactile sensor according to any one of claims 1 to 3, wherein The force detection and analysis method of the visual-tactile sensor comprises the following steps: inputting the magnetic induction data and the display data into a trained multi-modal pressure detection model, and outputting force position and / or force size from the multi-modal pressure detection model; the multi-modal pressure detection model is configured to: respectively pre-process the input magnetic induction data and display data; extract magnetic induction features from the pre-processed magnetic induction data and extract display features from the pre-processed display data; fuse the magnetic induction features and the display features to obtain fused features; infer the force position and / or the force size according to the fused features.
5. The visuo-tactile sensor of claim 4, wherein, extracting the magnetic induction features from the pre-processed magnetic induction data comprises the following steps: adopting a plurality of full connection layers to perform feature extraction on the pre-processed magnetic induction data layer by layer, so as to obtain the magnetic induction features; wherein each full connection layer respectively performs dimension expansion, activation operation and regularization processing on the input data.
6. The visuo-tactile sensor of claim 4, wherein, extracting the display features from the pre-processed display data comprises the following steps: adopting a plurality of convolution layers to perform feature extraction on the pre-processed display data layer by layer, wherein each convolution layer respectively performs dimension expansion on the input data; performing activation processing and pooling processing on the features after convolution processing to obtain the display features.
7. The visuo-tactile sensor of claim 4, wherein, fusing the magnetic induction features and the display features comprises the following steps: splicing the magnetic induction features and the display features; inputting the spliced features into an SE module for attention fusion; adopting a full connection layer to perform dimension compression on the features output by the SE module to obtain the fused features.
8. An e-skin, characterized by, The application further relates to an electronic skin, which comprises a skin body and at least one visual-tactile sensor as described in any one of claims 1-7 arranged on the skin body.
9. A robot, characterized in that The robot surface is provided with the electronic skin as described in claim 8.
Citation Information
Patent Citations
Visual tactile sensor based on magnetic imaging principle
CN121207402A