Visual tactile sensor stress detection and analysis method
By employing a magnetic thin film and magnetic graphics card design in the visual-tactile sensor, combined with a multimodal pressure detection model, the problem of low sensitivity in detecting top and side forces in the visual-tactile sensor is solved, achieving high-precision detection of top and side forces.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENGDU HUMANOID ROBOT INNOVATION CENT CO LTD
- Filing Date
- 2025-12-17
- Publication Date
- 2026-05-12
AI Technical Summary
Existing visual-tactile sensors have low sensitivity in detecting forces on the top and sides, and planar structures cannot detect lateral forces.
A magnetic thin film is used to cover the top and side areas. Combined with a magnetic graphics card and a light source, magnetic and imaging data are collected simultaneously by a visual-tactile sensor. The data is then analyzed using a multimodal pressure detection model to achieve high-precision detection of the forces acting on the top and side surfaces.
It enables comprehensive and high-precision detection of forces on the top and sides of the visual-tactile sensor, improving the sensitivity and comprehensiveness of force detection.
Smart Images

Figure CN122016093A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application No. 202511906273.7, filed on December 17, 2025, entitled "A visual-tactile sensor, electronic skin and robot". Technical Field
[0002] This invention relates to the field of sensors, and in particular to a method for force detection and analysis using a visual-tactile sensor. Background Technology
[0003] Electronic skin is a flexible electronic device that mimics the sensory function of human skin. Its core working principle is to capture external stimulus signals through sensors, convert and optimize them through signal processing modules, and finally output identifiable signals to achieve accurate detection of physical quantities such as pressure, temperature, humidity, and strain. Some electronic skins even have stimulus feedback functions.
[0004] In recent years, with the development of artificial intelligence and robotics, electronic skin capable of accurately detecting physical quantities such as pressure, temperature, humidity, and strain has received increasing attention and shown broad application prospects. In particular, the perception method based on visual-tactile sensors combined with magnetic induction has solved the problems of current visual-tactile sensors lacking the ability to receive 3D position information of marked points, limited camera frequency and response rate, and lack of perception of the state information of the object not being touched when using a monocular camera. This has enabled visual-tactile sensors to achieve high-precision force measurement, high response, and proximity sensing functions.
[0005] However, this type of sensor structure uses a surface coating on a transparent silicone surface to isolate external light sources and reconstruct the surface texture of the object being sensed. Information is then acquired by receiving deformation information from the surface coating via a camera. The presence of this surface coating increases the overall thickness of the magnetic thin film and affects the sensitivity of the sensor.
[0006] Furthermore, the magnetic thin films of known electronic skin structures are typically planar, and the forces acting on them cannot be detected when contact occurs on the side of the electronic skin. Summary of the Invention
[0007] The purpose of this invention is to provide a method for force detection and analysis of visual-tactile sensors, addressing all or part of the problems mentioned above, so as to achieve comprehensive and high-precision force detection and analysis of the top and side surfaces of the visual-tactile sensors.
[0008] The technical solution adopted in this invention is as follows: A method for force detection and analysis using a visual-tactile sensor, wherein the visual-tactile sensor comprises: A magnetic thin film covering the top and side areas; A support bracket is provided to support the magnetic film. A magnetic graphics card, mounted on the bracket, develops in response to the force applied to the magnetic thin film; The light source is installed inside the bracket; A visual-tactile sensor collects magnetic data of the force applied to the magnetic thin film and development data of the magnetic graphics card. Stress testing and analysis methods include: The magnetic field data and imaging data are input into a trained multimodal pressure detection model, which outputs the location and / or magnitude of the force.
[0009] Optionally, the multimodal pressure detection model is configured as follows: The input magnetic sensing data and development data are preprocessed separately; Magnetic features are extracted from the preprocessed magnetic data, and development features are extracted from the preprocessed development data. The magnetic field feature and the development feature are fused to obtain the fused feature; The location and / or magnitude of the force are inferred based on the fusion features.
[0010] Optionally, the preprocessing of the input magnetic sensing data and development data includes: The magnetic sensing data is standardized to convert it into a range that the multimodal pressure detection model can process. For the developed data, its resolution is controlled to a preset size, then the developed data is grayscale processed, and finally the pixel values of the developed data are normalized to the range of [0,1].
[0011] Optionally, the standardization process includes: The magnetic field data can be controlled within the desired range by subtracting the mean or dividing by the standard deviation.
[0012] Optionally, magnetic features are extracted from the preprocessed magnetic data, including: The magnetic features are obtained by extracting features from the preprocessed magnetic data layer by layer using a multi-layer fully connected layer. Extract development features from the preprocessed development data, including: Feature extraction is performed on the preprocessed development data layer by layer using multiple convolutional layers, where each convolutional layer expands the dimensions of the input data. The convolutional features are then subjected to activation and pooling processes to obtain the development features.
[0013] Optionally, the magnetic features and the imaging features are fused to obtain fused features, including: A cross-modal fusion layer is used to fuse the magnetic features and the imaging features, wherein: The cross-modal fusion layer first splices the magnetic features and the imaging features; then, it uses the SE module to perform attention fusion on the spliced features; finally, it uses a fully connected layer to compress the dimensions of the features output by the SE module to obtain the fused features.
[0014] Optionally, inferring the location and / or magnitude of the force based on the fused features includes: The fused features are inferred using an output layer. The output layer uses two fully connected layers to compress the dimensionality of the fused features, reducing the feature dimension to the number of output channels. When only the magnitude of the force needs to be inferred, the number of output channels is 1, outputting the magnitude of the force. When both the location and magnitude of the force need to be inferred simultaneously, and the location of the force is represented by two-dimensional coordinates, the number of output channels is 3, outputting the x-coordinate, y-coordinate, and magnitude of the force, respectively.
[0015] Optionally, the training method for the multimodal stress detection model includes: Obtain the training sample set; The training sample set is divided into a training set and a validation set according to a predetermined ratio; the training set is then sequentially input into the initialized multimodal stress detection model for iterative training. Once the model training stops, save the model parameters.
[0016] Optionally, obtain the training sample set, including: Pressing is performed on the magnetic film of the visual-touch sensor, and magnetic sensing data and imaging data are acquired at the time of pressing, as well as the force location and / or force magnitude. The magnetic sensing data, imaging data, and force location and / or force magnitude are associated with a set of training sample data, where the force location and / or force magnitude are the label values of the magnetic sensing data and imaging data. Multiple sets of training sample data are obtained by pressing with different forces at different locations. All training sample data are merged to obtain the training sample set.
[0017] Optionally, the loss function used to train the multimodal stress detection model is a weighted MSE loss function. When the final detection objective is to simultaneously detect the force location and force magnitude, the weighted MSE loss function is designed as follows: ; Where Loss represents the model loss. This represents the weight of the i-th channel, which has a total of three channels: the x-coordinate of the force application location, the y-coordinate of the force application location, and the magnitude of the force. and Let represent the predicted value and label value of the j-th training sample in the i-th channel, respectively, and m represent the number of training samples.
[0018] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: The electronic skin structure provided in this application effectively improves the sensitivity of magnetic film deformation detection by developing the deformation information of the magnetic thin film under stress using a magnetic graphics card. By simultaneously acquiring magnetic flux density data (magnetic flux change information) and visual development data (developed images) of the magnetic thin film under stress, high-precision, multi-modal information acquisition of the electronic skin's stress conditions is achieved. The magnetic thin film covers the top and side areas of the sensor; whether the force is applied to the top or side, it causes corresponding changes in the magnetic flux density and development data, enabling comprehensive detection of both forward and lateral forces on the electronic skin. Analysis using a multi-modal pressure detection model allows for high-precision inference of the force location and magnitude simultaneously. Attached Figure Description
[0019] The present invention will be described by way of example and with reference to the accompanying drawings, wherein: Figure 1 This is a front cross-sectional view of a visual-tactile sensor in one embodiment.
[0020] Figure 2 This is an exploded view of a visual-tactile sensor in one embodiment.
[0021] Figure 3 This is a diagonal cross-sectional view of a tactile sensor in one embodiment.
[0022] Figure 4 This is a data flow diagram of a force detection and analysis method for visual-tactile sensors in one embodiment.
[0023] Figure 5 This is a hierarchical structure diagram of a multimodal pressure detection model in one embodiment.
[0024] Figure 6 This is a network architecture diagram of a multimodal stress detection model in one embodiment.
[0025] Figure 7 This is a flowchart of one embodiment of the force detection and analysis method for visual-tactile sensors.
[0026] In the diagram, 1 is the magnetic film, 2 is the bracket, 3 is the magnetic graphics card, 4 is the light source, 5 is the visual and tactile sensor, 6 is the base, 11 is the top area, 12 is the side area, 13 is the edging, 21 is the horizontal bracket, 22 is the vertical bracket, 51 is the mounting plate, 52 is the image acquisition unit, 53 is the magnetic flux acquisition unit, 61 is the base, and 62 is the connector. Detailed Implementation
[0027] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.
[0028] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.
[0029] Existing visual-tactile sensor structures acquire information by depositing a surface coating on a transparent silicone surface to isolate external light sources and reconstruct the surface texture of the object being sensed, and then receiving deformation information of the surface coating through a camera. However, this approach suffers from increased overall thickness of the magnetic film and low sensitivity in force detection. Furthermore, existing electronic skin structures are generally planar, making it impossible to detect lateral forces when contact occurs on the side.
[0030] To address the aforementioned problems, this application provides a visual-tactile sensor, electronic skin, and robot, aiming to improve the sensitivity and comprehensiveness of pressure (contact force) detection.
[0031] In one feasible embodiment, this application provides a visual-tactile sensor, which includes a magnetic thin film 1, a bracket 2, a magnetic graphics card 3, a light source 4, and a visual-tactile sensor 5.
[0032] The magnetic thin film 1 is located on the outermost layer of the visual-tactile sensor. As a component that senses external forces, it covers at least the top surface region 11 and the side surface region 12 of the visual-tactile sensor to receive forces occurring on the top and sides. The magnetic thin film 1 is magnetic, and its deformation will cause a change in magnetic flux.
[0033] As an optional implementation, the magnetic thin film 1 is made of a mixture of magnetic powder and silicone rubber. That is, the magnetic powder is mixed into the silicone film rather than coated onto the surface of the silicone film. This effectively reduces the film thickness and improves the sensitivity of force detection. For example, by designing the magnetic thin film 1 in this way, the average thickness of the magnetic thin film can be configured to be 2mm to 5mm, resulting in high deformation response sensitivity.
[0034] In one feasible embodiment, the mass ratio of magnetic powder to silicone rubber is 0.8:1 to 1.2:1. Specifically, the magnetic powder and silicone rubber are mixed, stirred, molded, and magnetized at a mass ratio of 0.8:1 to 1.2:1 to obtain a magnetic thin film 1 with magnetic properties. The silicone rubber is obtained by mixing platinum reagent A and platinum reagent B at a mass ratio of 1:1.
[0035] In practical implementation, the average surface magnetic field of the magnetic thin film 1 is configured to be 10mT to 30mT, which can achieve accurate detection of the stress condition.
[0036] The Magneton Graphics Card 3 can use magneto-optical display for video, and it features high sensitivity (up to 10). -8 With its characteristics of being on the order of Wb (in terms of magnitude), strong anti-interference ability, and ability to perform magnetic flux distribution imaging, it can detect weak magnetic fields and thus detect subtle force conditions.
[0037] The support 2 is used to support the magnetic film 1 so that it has a stable shape / state when it is not under force and can return to that shape / state after the force is removed.
[0038] The magnetic graphics card 3 is mounted on the bracket 2, and can be close to the magnetic film 1 inside the bracket 2. If the magnetic film 1 deforms, the image of the magnetic graphics card 3 will change accordingly. Thus, the magnetic graphics card 3 will respond to the force on the magnetic film 1 and develop, reflecting the deformation information of the magnetic film 1.
[0039] In a preferred embodiment, the support 2 includes a horizontal support 21 supporting the top surface region 11 of the magnetic film 1, and a vertical support 22 supporting the side surface region 12 of the magnetic film 1. Figure 2 As shown, assuming the magnetic film 1 covers the top surface region 11 and the four side surface regions 12, the support 2 includes a horizontal support 21 supporting the top surface and four vertical supports 22 supporting the four side surface regions 12. Figure 2 The image shows a visible horizontal bracket 21 and two vertical brackets 22. The magnetic graphics card 3 is mounted parallel to the horizontal bracket 21, for example, mounted on the bottom surface of the horizontal bracket 21.
[0040] The light source 4 is installed inside the bracket 2. The light emitted by the light source 4 can illuminate the magnetic graphics card 3, thereby facilitating the acquisition of the image of the magnetic graphics card 3, i.e., the development data.
[0041] The visual-tactile sensor 5 collects magnetic flux data (i.e., magnetic flux change information) of the magnetic thin film and development data of the magnetic display 3, respectively. This can be achieved through two separate acquisition units; for example, magnetic flux acquisition unit 53 can collect magnetic flux data, while image acquisition unit 52 can collect development data. The image acquisition unit 52, which collects development data, faces the magnetic display 3. There can be one or more magnetic flux acquisition units 53, located within the magnetic field range of the magnetic thin film 1, such as within the space enclosed by the magnetic thin film 1.
[0042] In a preferred embodiment, the light source 4 is located behind the photosensitive element of the image acquisition unit 52. In this way, the image acquisition unit 52 only acquires the image of the magnetic graphics card 3 and does not acquire the image of the light source 4, thereby effectively reducing the influence of light on the accuracy of the acquisition of the color, outline, etc. of the image of the magnetic graphics card 3.
[0043] The visual-tactile sensor can be mounted and fixed using the base 6. That is, other components are set on the base 6. Specifically, the visual-tactile sensor 5 and the bracket 2 are mounted on the base 6, and other components can be supported by the visual-tactile sensor 5 or the bracket 2.
[0044] like Figures 1-3 As shown, in one feasible embodiment, the base 6 includes two parts: a base 61 and a connector 62. The base 61 serves as the foundation for mounting the visual-tactile sensor, and the connector 62 is connected to the base 61, for example, via plastic bolts. The connector 62 is used to mount the visual-tactile sensor 5 and the bracket 2, and also to lock the magnetic film 1, keeping it taut. Wherein, as... Figure 3 As shown, the connector 62, the visual-tactile sensor 5, and the bracket 2 are connected by bolts passing through them in sequence. The bottom edge 13 of the magnetic film 1 is sandwiched between the connector 62 and the base 61. If the base 6 is omitted, the magnetic film 1 can be locked to the bracket 2, for example, by locking the edge 13 of the magnetic film 1 to the bottom of the bracket 2 with plastic bolts.
[0045] In one optional embodiment, the visual-tactile sensor 5 includes a mounting plate 51, and an image acquisition unit 52 and a magnetic flux acquisition unit 53 disposed on the mounting plate 51. Figures 1-3 As shown, the mounting plate 51 is, for example, a PCB board. On the PCB board, multiple Hall sensors (four in the figure, but only for example, and can be adjusted according to the actual situation) are installed as magnetic flux acquisition units 53. The Hall sensors are, for example, triaxial force Hall sensors, which acquire magnetic flux change data in three dimensions. In the middle of the multiple Hall sensors, one depth camera (one for example, and can be adjusted according to the actual situation) is installed as an image acquisition unit 52. The depth camera is, for example, a monocular camera.
[0046] An LED strip is used as the light source 4, and the LED strip is arranged around the image acquisition unit 52. For example, in the embodiment described above where multiple Hall sensors surround the depth camera, the LED strip is arranged to surround each Hall sensor to improve the uniformity of light illumination and enhance the imaging effect. In this way, the light source 4 can be mounted on the mounting plate 51.
[0047] The aforementioned visual-tactile sensor can simultaneously acquire magnetic and imaging data generated by the deformation of the magnetic thin film 1 under external force through the visual-tactile collector 5. The magnetic thin film requires no coating layer, is relatively thin, and has higher sensitivity to applied forces, thus improving the accuracy of force detection by the visual-tactile sensor. The magnetic thin film 1 covers the top surface region 11 and the side surface region 12. The magnetic image sensor 3 can respond to normal and lateral forces. The visual-tactile collector 5 can simultaneously acquire magnetic and imaging data under normal or lateral forces, thereby achieving comprehensive, sensitive, multimodal, and high-precision sensing data acquisition for normal or lateral forces.
[0048] Based on the visual-tactile sensor described in the above embodiments, this application also introduces a force detection and analysis method, which includes: The magnetic field data and imaging data are input into the trained multimodal pressure detection model, which outputs the location and / or magnitude of the force.
[0049] See Figures 4-7 The multimodal pressure detection model is configured as follows: (1) Preprocess the input magnetic induction data and development data respectively.
[0050] Magnetic sensing data and imaging data belong to different modes of sensing data and require different preprocessing methods. Preprocessing of the input data is performed at the input layer of the multimodal pressure detection model.
[0051] like Figure 4 , Figure 5 As shown, in the input layer, the magnetic data is standardized to convert it into a range that the model can handle. For example, the values of the magnetic data are normalized to the range [0,1].
[0052] For the developed data, its resolution is controlled to a preset size, such as 64 pixels × 64 pixels, through scaling or cropping; then, the developed data is converted to grayscale; finally, the pixel values of the developed data are normalized to the range [0,1] to facilitate model processing. Additionally, in an optional implementation, sample data augmentation can be performed on the developed data samples to expand the training samples. For example, the developed data samples can be processed through rotation, flipping, masking operations, etc., to obtain new samples, which are then preprocessed.
[0053] (2) Extract magnetic features from the preprocessed magnetic data and extract development features from the preprocessed development data.
[0054] like Figure 4 , Figure 5 As shown, the multimodal pressure detection model uses a feature extraction layer to extract magnetic and imaging features respectively.
[0055] In one alternative implementation, extracting magnetic features from the preprocessed magnetic field data includes: A multi-layer fully connected approach is used to extract features from the preprocessed magnetic field data, obtaining magnetic field features. Each fully connected layer performs dimensionality expansion, activation operations, and regularization on the input data.
[0056] For example, assuming a visual-haptic sensor contains 5 Hall sensors, the magnetic data it collects is 1×5×3=1×15 dimensional, therefore the input dimension of the first fully connected layer is 15. Figure 4 As shown, the multi-layer first fully connected network of the multimodal pressure detection model uses four fully connected layers to extract features from the preprocessed magnetic field data. The output dimensions of the four fully connected layers are 15-64-128-256, respectively. The ReLU function is used as the activation function for the fully connected layers, and Dropout is used for regularization after each fully connected layer. Thus, after feature extraction by the four fully connected layers, the multi-layer first fully connected network outputs the corresponding magnetic field features. These magnetic field features reflect the mapping relationship between the change in magnetic field and the location / magnitude of the force when force is applied.
[0057] like Figure 4 , Figure 5 As shown, in one optional implementation, extracting development features from the preprocessed development data includes: Feature extraction is performed on the preprocessed development data layer by layer using multiple convolutional layers, where each convolutional layer expands the dimensions of the input data.
[0058] like Figure 5 As shown, a lightweight CNN (Convolutional Neural Network) (CNN Lite), such as MobileNet V2, is used to extract imaging features in a multimodal stress detection model. This lightweight CNN consists of three convolutional layers, each with a 3×3 kernel. The imaging data acquired by the image acquisition unit 52 has depth information, therefore the imaging data has three dimensions: length, width, and height. Thus, the input dimension of the first convolutional layer is 3. The output dimensions of the three convolutional layers are 16-32-64, respectively.
[0059] The convolutional features are then activated and pooled to obtain the development features. In the development feature extraction branch, the convolutional features are activated using the ReLU function, followed by max pooling and global average pooling to obtain 256-dimensional development features. These development features reflect the mapping relationship between the changes in the development texture of the magnetic graphics card 3 under stress and the location / magnitude of the stress.
[0060] (3) Combine magnetic induction features and imaging features to obtain fused features.
[0061] As an alternative implementation, the multimodal pressure detection model employs a cross-modal fusion layer to fuse magnetic and imaging features.
[0062] like Figure 4 , Figure 5 As shown, the cross-modal fusion layer first splices the magnetic sensing features and the development features. Taking the previous implementation as an example, the 256-dimensional magnetic sensing features and the 256-dimensional development features are spliced into a 512-dimensional spliced feature.
[0063] The cross-modal fusion layer then uses the SE module to perform attention fusion on the spliced features. The SE module dynamically assigns fusion weights to the magnetic and imaging features, and the weighted fusion yields a 512-dimensional vector.
[0064] The cross-modal fusion layer also employs a fully connected layer to compress the dimensionality of the features output by the SE module, resulting in fused features. Specifically, based on the 512-dimensional vector obtained by the weighted fusion above, a fully connected layer is used to compress it into a 256-dimensional fused feature, which is then used as the output of the cross-modal fusion layer.
[0065] (4) Based on the fusion characteristics, infer the location and / or magnitude of the force.
[0066] As an optional implementation method, such as Figure 4 , Figure 5 As shown, the multimodal stress detection model uses the output layer to infer the fused features.
[0067] The number of output channels varies depending on the purpose of the inference. For example, if only the magnitude of the force needs to be inferred, there is 1 output channel; if both the location and magnitude of the force need to be inferred simultaneously, and the location of the force is represented by two-dimensional coordinates, then there are 3 output channels, outputting the x-coordinate, y-coordinate, and magnitude of the force, respectively. The output of the 3 channels can be simplified to (x, y, N), where x and y represent the x-coordinate and y-coordinate, respectively, and N represents the magnitude of the force.
[0068] like Figure 5 As shown, specifically, the output layer first uses two fully connected layers (the second fully connected network) for dimensionality compression, compressing the 256-dimensional fused features to 128 dimensions, and then compressing the 128-dimensional features into a 3-dimensional feature vector (number of output channels) for output. The output of the output layer is a continuous value, where x ranges from [0, W], y ranges from [0, H], N ranges from [0, F_max], W and H are the length and width dimensions of the visual-touch sensor, and F_max is the maximum pressure.
[0069] To more clearly illustrate the workflow of the multimodal pressure detection model proposed in this application, such as... Figure 6 As shown in the embodiments of this application, the construction of a multimodal pressure detection model is illustrated by example.
[0070] like Figure 6 As shown, the multimodal stress detection model consists of four layers: an input layer, a feature extraction layer, a cross-modal fusion layer, and an output layer. The input layer is responsible for preprocessing the input data; the feature extraction layer is responsible for extracting features from the multimodal data; the cross-modal fusion layer is responsible for fusing the multimodal features; and the output layer is responsible for task inference based on the fused features.
[0071] The input layer receives two modalities of data: magnetic sensing data and development data. Both are preprocessed. Specifically, the magnetic sensing data is controlled within a desired range by subtracting the mean or dividing by the standard deviation. For example, the magnetic sensing data is normalized to [0,1]. For the development data, it can be scaled and / or cropped to control the resolution to a preset size, such as 64 pixels × 64 pixels. The development data is then converted to grayscale, and finally, the pixel values are normalized to [0,1]. Additionally, the development data can be enhanced and expanded by rotating, flipping, or adding mosaic effects to improve the training sample set.
[0072] The preprocessed magnetic field data and development data are input into the feature extraction layer for multimodal feature extraction. Specifically, the feature extraction layer uses a first fully connected network to extract features from the preprocessed magnetic field data, extracting magnetic field features; and a lightweight CNN is used to extract features from the preprocessed development data, extracting development features. The structure of the first fully connected network can be found in the previous embodiment and will not be repeated here. The lightweight CNN performs multi-layer convolution processing on the input development data, then performs batch normalization using BN (Batch Normalization) layers, performs activation processing using the ReLU function, and finally performs pooling processing using pooling layers. The pooling processing can be max pooling or global average pooling.
[0073] The magnetic and developmental features extracted by the feature extraction layer are input into the cross-modal fusion layer, which then fuses them. Specifically, the cross-modal fusion layer first concatenates the magnetic and developmental features, combining the 256-dimensional features into a 512-dimensional concatenated feature. Then, the SE module applies attention weights to the magnetic and developmental features within the concatenated feature. Finally, a fully connected layer compresses the dimensionality of the weighted concatenated feature, reducing it from 512 dimensions to 256 dimensions, resulting in the fused feature.
[0074] The fused features are finally input into the output layer for task inference. The output layer uses a second fully connected network to reduce the dimensionality of the fused features to the dimension corresponding to the number of task channels. Taking the inference of force location and force magnitude as an example, the second fully connected network uses two fully connected layers to complete the inference from the fused features to the output task using a 256-128-3 dimensional compression rule. The three output channels are the x-coordinate of the force location, the y-coordinate of the force location, and the force magnitude, respectively.
[0075] Multimodal stress detection models require training before they can accurately infer the location and / or magnitude of forces. In one optional implementation, the training method for the multimodal stress detection model includes: (1) Obtain the training sample set.
[0076] A controllable robotic arm or similar device capable of accurately applying pressure continuously presses onto the magnetic film 1 of the visual-tactile sensor. Here, a robotic arm is used as an example. Magnetic and developmental data are acquired during each press, along with the force location and / or force magnitude of the robotic arm press (determined based on the final detection objective). These data are then linked to form a set of training sample data. Multiple sets of training sample data are obtained by controlling the robotic arm to apply different forces at different positions. Furthermore, the developmental data can be augmented through techniques such as flipping, rotating, or masking (i.e., adding mosaic effects), and this augmented data becomes the new developmental data corresponding to the current conditions, also used as training sample data. All training sample data are then merged to obtain a training sample set. For example, assuming simultaneous detection of force location and force magnitude is required, the training sample set must contain at least (magnetic data, developmental data, force location, force magnitude), where the force location and force magnitude are the label values corresponding to the magnetic and developmental data, respectively.
[0077] (2) Model training.
[0078] The training sample set is randomly divided into a training set and a validation set according to a certain ratio. The training set is sequentially input into the initialized multimodal stress detection model for iterative training. Each iteration yields the predicted force location and force magnitude. The model loss is then calculated based on the label values using a weighted MSE loss function. The weighted MSE loss function is designed as follows: ; Where Loss represents the model loss. This represents the weight of the i-th channel. and Let represent the predicted value and label value of the j-th training sample in the i-th channel, respectively, and m represent the number of training samples.
[0079] The weighted MSE loss function assigns weights to the three output channels x, y, and N respectively. These weights are pre-set based on the priority of the detected target / task.
[0080] The stopping condition for model training is: 1) The model loss on the validation set decreases by less than a preset value for K consecutive rounds (K is a positive integer, pre-defined). This preset value is, for example, [value missing]. .
[0081] 2) The model loss on the validation set is reduced to below a preset threshold. For example, the pixel-level loss corresponding to the force location is <0.1, and the loss for the force magnitude is <0.05.
[0082] Once the stopping condition is met, the model parameters are saved, and the trained multimodal pressure detection model can be used to detect the location and magnitude of the force based on the magnetic field data and imaging data.
[0083] The visual-tactile sensor provided in the above embodiments uses a magnetic thin film 1 for force sensing, eliminating the need for a coating layer and effectively improving the sensitivity of force perception. Furthermore, the magnetic thin film 1 covers the top surface region 11 and the side surface region 12, and integrates the magnetic flux acquisition unit 53 and the image acquisition unit 52 into a visual-tactile sensor 5 to simultaneously acquire magnetic and imaging data. Combining these two data for force detection analysis provides extremely high accuracy in detecting both top and side surface forces. The designed multimodal pressure detection model extracts magnetic and imaging features from two branches, then performs attention-weighted fusion. Based on the fused features, multi-task inference is performed, effectively improving the accuracy, richness, and comprehensiveness of task detection.
[0084] Based on the concept of this application, embodiments of this application also propose an electronic skin, which includes a skin body and at least one visual-tactile sensor disposed on the skin body. By distributing / concentrating the visual-tactile sensors on the skin body, and by detecting the location and / or magnitude of the force acting on each visual-tactile sensor, the location and / or magnitude of the force on the electronic skin can be accurately detected.
[0085] Furthermore, this application also proposes a robot whose surface is provided with the aforementioned electronic skin. The robot can be a robotic arm or a humanoid robot, and the electronic skin can be installed at the locations where force detection is required.
[0086] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.
Claims
1. A method for force detection and analysis using a visual-tactile sensor, characterized in that, The visual-tactile sensor includes: A magnetic thin film covering the top and side areas; A support bracket is provided to support the magnetic film. A magnetic graphics card, mounted on the bracket, develops in response to the force applied to the magnetic thin film; The light source is installed inside the bracket; A visual-tactile sensor collects magnetic data of the force applied to the magnetic thin film and development data of the magnetic graphics card. Stress testing and analysis methods include: The magnetic field data and imaging data are input into a trained multimodal pressure detection model, which outputs the location and / or magnitude of the force.
2. The force detection and analysis method of the visual-tactile sensor as described in claim 1, characterized in that, The multimodal pressure detection model is configured as follows: The input magnetic sensing data and development data are preprocessed separately; Magnetic features are extracted from the preprocessed magnetic data, and development features are extracted from the preprocessed development data. The magnetic field feature and the development feature are fused to obtain the fused feature; The location and / or magnitude of the force are inferred based on the fusion features.
3. The force detection and analysis method for visual-tactile sensors as described in claim 2, characterized in that, The preprocessing of the input magnetic sensing data and development data includes: The magnetic sensing data is standardized to convert it into a range that the multimodal pressure detection model can process. For the developed data, its resolution is controlled to a preset size, then the developed data is grayscale processed, and finally the pixel values of the developed data are normalized to the range of [0,1].
4. The force detection and analysis method of the visual-tactile sensor as described in claim 3, characterized in that, The standardization process includes: The magnetic field data can be controlled within the desired range by subtracting the mean or dividing by the standard deviation.
5. The force detection and analysis method for visual-tactile sensors as described in claim 2, characterized in that, Magnetic features are extracted from the preprocessed magnetic field data, including: The magnetic features are obtained by extracting features from the preprocessed magnetic data layer by layer using a multi-layer fully connected layer. Extract development features from the preprocessed development data, including: Feature extraction is performed on the preprocessed development data layer by layer using multiple convolutional layers, where each convolutional layer expands the dimensions of the input data. The convolutional features are then subjected to activation and pooling processes to obtain the development features.
6. The force detection and analysis method of the visual-tactile sensor as described in claim 2, characterized in that, By fusing the magnetic features and the imaging features, a fused feature is obtained, including: A cross-modal fusion layer is used to fuse the magnetic features and the imaging features, wherein: The cross-modal fusion layer first splices the magnetic features and the imaging features; then, it uses the SE module to perform attention fusion on the spliced features; finally, it uses a fully connected layer to compress the dimensions of the features output by the SE module to obtain the fused features.
7. The force detection and analysis method for visual-tactile sensors as described in claim 2, characterized in that, Inferring the location and / or magnitude of the force based on the fused features includes: The fused features are inferred using an output layer. The output layer uses two fully connected layers to compress the dimensionality of the fused features, reducing the feature dimension to the number of output channels. When only the magnitude of the force needs to be inferred, the number of output channels is 1, outputting the magnitude of the force. When both the location and magnitude of the force need to be inferred simultaneously, and the location of the force is represented by two-dimensional coordinates, the number of output channels is 3, outputting the x-coordinate, y-coordinate, and magnitude of the force, respectively.
8. The method for force detection and analysis of a visual-tactile sensor as described in any one of claims 1-7, characterized in that, The training method for the multimodal stress detection model includes: Obtain the training sample set; The training sample set is divided into a training set and a validation set according to a predetermined ratio; the training set is then sequentially input into the initialized multimodal stress detection model for iterative training. Once the model training stops, save the model parameters.
9. The force detection and analysis method for a visual-tactile sensor as described in claim 8, characterized in that, Obtain the training sample set, including: Pressing is performed on the magnetic film of the visual-touch sensor, and magnetic sensing data and imaging data are acquired at the time of pressing, as well as the force location and / or force magnitude. The magnetic sensing data, imaging data, and force location and / or force magnitude are associated with a set of training sample data, where the force location and / or force magnitude are the label values of the magnetic sensing data and imaging data. Multiple sets of training sample data are obtained by pressing with different forces at different locations. All training sample data are merged to obtain the training sample set.
10. The force detection and analysis method for a visual-tactile sensor as described in claim 9, characterized in that, The loss function used to train the multimodal stress detection model is a weighted MSE loss function. When the final detection objective is to simultaneously detect the force location and force magnitude, the weighted MSE loss function is designed as follows: ; Where Loss represents the model loss. This represents the weight of the i-th channel, which has a total of three channels: the x-coordinate of the force application location, the y-coordinate of the force application location, and the magnitude of the force. and Let represent the predicted value and label value of the j-th training sample in the i-th channel, respectively, and m represent the number of training samples.