Vision-based calibration method and device for tactile sensors

By combining simulators and convolutional neural networks, the problem of low calibration efficiency of vision-based tactile sensors is solved, achieving efficient sensor calibration that is suitable for robot grasping and manipulation scenarios.

CN115719385BActive Publication Date: 2025-10-31BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211426179.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-10-31
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

The calibration of vision-based tactile sensors requires the acquisition of a large number of real deformation images, resulting in low calibration efficiency and restricting mass production.

Method used

By establishing and calibrating a simulator, a set of simulated compression images is generated. The target model is then trained using a convolutional neural network to achieve efficient sensor calibration.

Benefits of technology

It greatly improves the calibration efficiency of tactile sensors and reduces the amount of data collected by about one order of magnitude, making it suitable for robot grasping and manipulation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719385B_ABST
    Figure CN115719385B_ABST
Patent Text Reader

Abstract

This application provides a vision-based tactile sensor calibration method and apparatus, relating to the field of tactile sensor calibration. The method includes: acquiring contact surface deformation images collected by a target sensor, and establishing and calibrating a simulator based on the acquired contact surface deformation images; generating a set of simulated pressing images of the target sensor's contact surface in contact with objects of different shapes using the simulator; and training a target model based on the set of simulated pressing images to complete the calibration of the target sensor. The vision-based tactile sensor calibration method and apparatus provided in this application are used to generate deformation images of the sensor's contact surface in batches through simulation, improving the calibration efficiency of vision-based tactile sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of tactile sensor calibration, and more particularly to a vision-based tactile sensor calibration method and apparatus. Background Technology

[0002] Vision-based tactile sensors (such as GelSight) are sensors that capture images of the deformation of the sensor's contact surface and use photometric stereochemistry to reconstruct the three-dimensional geometric deformation of the contact surface to achieve tactile sensing. They are widely used in robot grasping and manipulation scenarios.

[0003] In related technologies, the calibration of vision-based tactile sensors requires acquiring a large number of real deformation images of the contact surface, which is time-consuming and inefficient, severely restricting the mass production of such sensors. Therefore, an efficient calibration method is urgently needed to facilitate the further application and mass production of such sensors. Summary of the Invention

[0004] The purpose of this application is to provide a vision-based tactile sensor calibration method and apparatus, which is used to generate deformation images of the contact surface of the sensor in batches through simulation, thereby improving the calibration efficiency of vision-based tactile sensing.

[0005] This application provides a vision-based tactile sensor calibration method, including:

[0006] The process involves acquiring contact surface deformation images from a target sensor, and establishing and calibrating a simulator based on these images. The simulator generates a set of simulated pressing images of the target sensor's contact surface when it comes into contact with objects of different shapes. A target model is trained based on this set of simulated pressing images to calibrate the target sensor. The establishment and calibration of the simulator includes: establishing and calibrating a near-field camera model, a near-field light source model, a surface reflection model, and a surface deformation model. The trained target model is used to reconstruct the three-dimensional geometric deformation of the contact surface from the deformation images of the target sensor's contact surface.

[0007] Optionally, acquiring the contact surface deformation image collected by the target sensor includes: controlling a calibration ball of a preset size to contact the contact surface of the target sensor, and controlling the calibration ball to be pressed into a hemispherical position on the contact surface; acquiring the deformation image of the contact surface to obtain the contact surface deformation image.

[0008] Optionally, the near-field camera model includes: a camera geometric model; the step of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: acquiring the camera intrinsic parameters of the target sensor, and constructing a camera coordinate system based on the camera position of the target sensor, and constructing a world coordinate system based on the contact surface; determining the extrinsic parameters of the near-field camera based on the three-dimensional rigid transformation from the world coordinate system to the camera coordinate system; establishing the camera geometric model based on the camera intrinsic parameters and the extrinsic parameters of the near-field camera; and calibrating the camera intrinsic parameters and camera extrinsic parameters of the camera geometric model based on the camera geometric calibration method.

[0009] Optionally, the near-field camera model includes: a camera radiation model; the step of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: determining the image value and image illuminance of each color channel based on the contact surface deformation image; determining the photometric response curve of the near-field camera based on the image value and image illuminance of each color channel; determining the vignette effect of the near-field camera based on the image illuminance of each color channel of the contact surface deformation image; establishing the camera radiation model based on the photometric response curve and the vignette effect of the near-field camera; calibrating the photometric response curve of the camera radiation model based on the photometric response curve calibration method; and calibrating the vignette effect of the camera radiation model based on the vignette effect calibration method.

[0010] Optionally, the near-field light source model includes: a light source geometric model and a light source radiation model; the step of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: establishing the light source geometric model based on the position of each light source of the target sensor, and calibrating the position of each light source in the light source geometric model using a light source position calibration method; establishing the light source radiation model based on the principal optical axis direction of each light source of the target sensor and the relative energy intensity of each light source in different directions, and calibrating the radiation of each light source in the light source radiation model based on the brightness of each pixel in the contact surface deformation image.

[0011] Optionally, the step of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: using a generalized Lambertian reflection model as the surface reflection model; and calibrating the surface roughness and surface reflectivity of the surface reflection model using the Levenberg-Marquardt algorithm based on the contact surface deformation image.

[0012] Optionally, the step of establishing and calibrating the simulator based on the collected deformation image of the contact surface includes: calculating the surface deformation of the contact surface when it is pressed based on the intersection relationship between the light and the geometric patch, and smoothing the deformed and non-deformed areas using a Gaussian pyramid.

[0013] Optionally, the target model is a convolutional neural network model; the step of training the target model based on the set of simulated pressing images to complete the calibration of the target sensor includes: calculating the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images; using the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples to train the target model; wherein, the target model is a fully convolutional encoder-decoder architecture, using cosine similarity as the loss function; the encoder includes: 5 convolutional layers, used to reduce the feature scale of the image by a preset factor to reduce the computational load; the decoder includes: 6 convolutional layers and an L2 normalization layer, used to regress the surface normal vector of each pixel and increase the feature scale of the image by the preset factor to restore the image scale.

[0014] Optionally, training the target model using the set of simulated pressing images and the surface normal field corresponding to each simulated pressing image in the set as training samples includes: performing a phase difference operation between the input simulated pressing image and the image of the contact surface in a non-pressed state, retaining the image of the deformed part of the contact surface; inputting the image of the deformed part into the target model to obtain the prediction result of the surface normal field after the contact surface is deformed.

[0015] This application also provides a vision-based tactile sensor calibration device, comprising:

[0016] The system comprises the following modules: an acquisition module for acquiring contact surface deformation images collected by the target sensor; a simulation module for establishing and calibrating a simulator based on the acquired contact surface deformation images; a generation module for generating a set of simulated pressing images of the target sensor's contact surface in contact with objects of different shapes using the simulator; and a calibration module for training a target model based on the set of simulated pressing images to calibrate the target sensor. The establishment and calibration of the simulator includes: establishing and calibrating a near-field camera model, a near-field light source model, a surface reflection model, and a surface deformation model. The trained target model is used to reconstruct the three-dimensional geometric deformation of the contact surface from the deformation images of the target sensor's contact surface.

[0017] Optionally, the control further includes: a control module; the control module is used to control a calibration ball of a preset size to contact the contact surface of the target sensor, and to control the calibration ball to be pressed into a hemispherical position on the contact surface; the acquisition module is specifically used to acquire a deformation image of the contact surface to obtain the deformation image of the contact surface.

[0018] Optionally, the device further includes: a determination module; the near-field camera model includes: a camera geometric model; the acquisition module is further configured to acquire the camera intrinsic parameters of the target sensor; the simulation module is configured to construct a camera coordinate system based on the camera position of the target sensor and a world coordinate system based on the contact surface; the determination module is configured to determine the extrinsic parameters of the near-field camera based on the three-dimensional rigid transformation from the world coordinate system to the camera coordinate system; the simulation module is specifically configured to establish the camera geometric model based on the camera intrinsic parameters and the extrinsic parameters of the near-field camera; the simulation module is further configured to calibrate the camera intrinsic parameters and camera extrinsic parameters of the camera geometric model based on a camera geometric calibration method.

[0019] Optionally, the near-field camera model includes: a camera radiation model; the determining module is further configured to determine the image value of each color channel and the image illuminance of each color channel based on the contact surface deformation image; the determining module is further configured to determine the photometric response curve of the near-field camera based on the image value of each color channel and the image illuminance of each color channel; the determining module is further configured to determine the vignette effect of the near-field camera based on the image illuminance of each color channel of the contact surface deformation image; the simulation module is specifically further configured to establish the camera radiation model based on the photometric response curve of the near-field camera and the vignette effect of the near-field camera; the simulation module is specifically further configured to calibrate the photometric response curve of the camera radiation model based on the photometric response curve calibration method, and calibrate the vignette effect of the camera radiation model based on the vignette effect calibration method.

[0020] Optionally, the near-field light source model includes: a light source geometric model and a light source radiation model; the simulation module is specifically used to establish the light source geometric model according to the position of each light source of the target sensor, and to calibrate the position of each light source in the light source geometric model using a light source position calibration method; the simulation module is also specifically used to establish the light source radiation model according to the principal optical axis direction of each light source of the target sensor and the relative energy intensity of each light source in different directions, and to calibrate the radiation of each light source in the light source radiation model according to the brightness of each pixel in the contact surface deformation image.

[0021] Optionally, the determining module is further configured to use a generalized Lambertian reflection model as the surface reflection model; the simulation module is specifically configured to calibrate the surface roughness and surface reflectivity of the surface reflection model based on the contact surface deformation image using the Levenberg-Marquardt algorithm.

[0022] Optionally, the simulation module is specifically used to calculate the surface deformation of the contact surface when it is pressed based on the intersection relationship between the light and the geometric patch, and to smooth the deformed and non-deformed areas using a Gaussian pyramid.

[0023] Optionally, the device further includes: a calculation module; the target model is a convolutional neural network model; the calculation module is used to calculate the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images; the calibration module specifically uses the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples to train the target model; wherein, the target model is a fully convolutional encoder-decoder architecture, using cosine similarity as the loss function; the encoder includes: 5 convolutional layers, used to reduce the feature scale of the image by a preset factor to reduce the computational load; the decoder includes: 6 convolutional layers and an L2 normalization layer, used to regress the surface normal vector of each pixel and increase the feature scale of the image by the preset factor to restore the image scale.

[0024] Optionally, the calibration module is specifically used to perform a phase difference operation between the input simulated pressing image and the image of the contact surface in the non-pressed state, and retain the image of the deformed part of the contact surface; the calibration module is also specifically used to input the image of the deformed part into the target model to obtain the prediction result of the surface normal vector field after the contact surface is deformed.

[0025] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the vision-based tactile sensor calibration method as described above.

[0026] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the vision-based tactile sensor calibration methods described above.

[0027] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the vision-based tactile sensor calibration method described above.

[0028] The vision-based tactile sensor calibration method and apparatus provided in this application first acquires contact surface deformation images collected by the target sensor, and then establishes and calibrates a simulator based on the acquired contact surface deformation images. Next, the simulator generates a set of simulated pressing images of the target sensor's contact surface in contact with objects of different shapes. Finally, a target model is trained based on the set of simulated pressing images to complete the calibration of the target sensor, greatly improving the calibration efficiency of vision-based tactile sensors. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating the vision-based tactile sensor calibration method provided in this application;

[0031] Figure 2 This is a schematic diagram of the camera perspective projection model provided in this application;

[0032] Figure 3 This is a schematic diagram of the vignette effect provided in this application;

[0033] Figure 4 This is a schematic diagram of the structure of the vision-based tactile sensor calibration device provided in this application;

[0034] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0037] GelSight is a vision-based tactile sensor that acquires images of the contact surface and uses photometric stereochemistry to reconstruct the three-dimensional geometric deformation of the contact surface to achieve tactile sensing. Obtaining the model and parameters required to transform the original sensor image into the three-dimensional geometric deformation is called sensor calibration. GelSight sensors are widely used in robot manipulation tasks.

[0038] In related technologies, GelSight sensors require a large number of devices to perform tactile sensing in many applications (e.g., robotic arms). Furthermore, GelSight calibration necessitates collecting extensive real sensor data (due to the limited size of these sensors, which typically contain only a single camera, requiring even more real deformation images for calibration), hindering efficient mass production. Therefore, calibration efficiency severely restricts the widespread adoption and application of these sensors.

[0039] In view of the above-mentioned technical problems in related technologies, the vision-based tactile sensor calibration method provided in this application aims to solve the problem of small sample geometric deformation calibration of such tactile sensors. The vision-based tactile sensor calibration method provided in this application can reduce the amount of data collection by at least about one order of magnitude.

[0040] The vision-based tactile sensor calibration method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0041] like Figure 1 As shown in the embodiment of this application, a vision-based tactile sensor calibration method is provided, which may include the following steps 101 to 103:

[0042] Step 101: Obtain the contact surface deformation image collected by the target sensor, and complete the establishment and calibration of the simulator based on the collected contact surface deformation image.

[0043] The establishment and calibration of the simulator includes: the establishment and calibration of the near-field camera model, the establishment and calibration of the near-field light source model, the establishment and calibration of the surface reflection model, and the establishment of the surface deformation model.

[0044] For example, the target sensor is a vision-based tactile sensor; the contact surface deformation image is an image of the contact surface of the target sensor coming into contact with an object and deforming.

[0045] It should be noted that the contact surface of vision-based tactile sensors is usually made of silicone. When the contact surface comes into contact with an object, it will deform. The model required to deduce the three-dimensional geometric deformation of the contact surface from the deformation image of the contact surface is called sensor calibration.

[0046] For example, in order to more realistically simulate the deformation images produced when the contact surface comes into contact with different objects, the simulator in this application embodiment needs to include: a near-field camera model, a near-field light source model, a surface reflection model, a projected shadow model, and a surface deformation model.

[0047] For example, the simulator can be built and calibrated based on the camera parameters of the target sensor, the distribution of the light source, and a small number of real deformation images.

[0048] Step 102: Generate a set of simulated pressing images of the contact surface of the target sensor when it comes into contact with objects of different shapes using the simulator.

[0049] For example, after the simulator is built and calibrated, it can be used to generate a large set of simulated pressing images of the contact surface in contact with objects of different shapes, and the target sensor can be calibrated based on the images in the simulated pressing image set.

[0050] Step 103: Train the target model based on the set of simulated pressing images to complete the calibration of the target sensor.

[0051] The trained target model is used to restore the deformation image of the contact surface of the target sensor into the three-dimensional geometric deformation of the contact surface.

[0052] It should be noted that, based on the characteristics of the GelSight sensor, this embodiment first constructs an optical simulator for this type of sensor, namely real2sim, through geometric optical physical modeling and parameter estimation. In this process, physical modeling and parameter estimation are required for elements such as the near-field camera, near-field light source, material reflection, and surface deformation. This process can be completed by collecting a small amount of real data. Establishing the sensor optical simulator is similar to the process of establishing a sensor imaging model (forward problem) in model-based photometric stereo methods, using a physical model to replace real data.

[0053] Then, using this simulator, a large amount of paired data can be generated, showing the sensor readings and the true values ​​of the deformation normal vector field when the contact surface deforms in contact with different objects. A convolutional neural network is then used as the inverse problem solver to solve the inverse problem, ultimately achieving efficient calibration of this type of sensor. This process establishes a mapping from the image to the corresponding geometric deformation normal vector field through the neural network. Since the training data originates from the simulator, rather than data collected by a real sensor, the process of acquiring real images is eliminated, significantly improving the sensor calibration efficiency.

[0054] Finally, by deploying the trained neural network model, geometric deformation estimation can be performed on real sensor data (deformation images of the contact surface collected by the sensor), i.e., sim2real.

[0055] Optionally, in the embodiments of this application, the establishment and calibration of each model in the above simulator can be completed through the following steps.

[0056] For example, before establishing the above model, it is necessary to first collect a certain number of real deformation images.

[0057] Specifically, step 101 above may include steps 101a1 and 101a2:

[0058] Step 101a1: Control the calibration ball of a preset size to contact the contact surface of the target sensor, and control the calibration ball to press the contact surface to a hemispherical position.

[0059] Step 101a2: Acquire the deformation image of the contact surface to obtain the deformation image of the contact surface.

[0060] For example, a calibration ball with a diameter of 2 mm can be selected, with one part (hemisphere) of the calibration ball located below the horizontal line of the contact surface of the target sensor, and the other part (hemisphere) located above the horizontal line of the contact surface of the target sensor. Then, the deformation image of the contact surface is acquired.

[0061] It should be noted that the number of real deformation images required for calibration of different models is not entirely the same, as shown in Table 1 below, which lists the number of real deformation images required for calibration of different models:

[0062]

[0063] Table 1

[0064] It should be noted that, in the embodiments of this application, camera geometry calibration, camera radiation calibration, and surface reflection calibration can be calibrated only once, while the other two types of calibration (light source radiation calibration and surface reflection calibration) need to be performed separately for each sensor. Therefore, the minimum data acquisition requirement of the method proposed in this invention is 6 images, that is, the above-mentioned contact surface deformation images include at least six real deformation images of the contact surface. Based on the 6 real deformation images, a simulator can be built to simulate the pressing process of the sensor, and finally the calibration of the sensor is completed through a convolutional deep neural network.

[0065] For example, the above near-field camera model includes: a camera geometry model and a camera radiation model.

[0066] Specifically, step 102 above, which involves establishing and calibrating the camera geometry model in the near-field camera model, may include the following steps 102a1 to 102a4:

[0067] Step 102a1: Obtain the camera intrinsic parameters of the target sensor, construct a camera coordinate system based on the camera position of the target sensor, and construct a world coordinate system based on the contact surface.

[0068] Step 102a2: Determine the extrinsic parameters of the near-field camera based on the three-dimensional rigid transformation from the world coordinate system to the camera coordinate system.

[0069] Step 102a3: Establish the camera geometric model based on the camera intrinsic parameters and the near-field camera extrinsic parameters.

[0070] Step 102a4: Calibrate the camera intrinsic and extrinsic parameters of the camera geometric model based on the camera geometry calibration method.

[0071] For example, a camera geometry model can be described by a camera perspective projection model, which aims to establish a geometric projection relationship from the three-dimensional world to the two-dimensional camera, such as... Figure 2 As shown, a three-dimensional point The projection point {P}X onto the two-dimensional phase plane can be represented by the following formula:

[0072]

[0073] In this system, {C} and {P} represent the camera coordinate system and pixel coordinate system, respectively, and {W} represents the world coordinate system. The XY plane (i.e., the reference plane) is parallel to the sensor's contact surface. K represents the camera's intrinsic parameters, which can be directly obtained from the sensor's parameters. The three-dimensional rigid transformation from {W} to {C} is called the camera's extrinsic parameters. Based on the camera's extrinsic and intrinsic parameters, a camera geometric model can be established. Subsequently, the camera's intrinsic and extrinsic parameters can be calibrated using camera geometric calibration methods.

[0074] Specifically, step 102 above, which involves establishing and calibrating the camera radiation model in the near-field camera model, may include the following steps 102b1 to 102b5:

[0075] Step 102b1: Determine the image value of each color channel and the image illuminance of each color channel based on the contact surface deformation image.

[0076] Step 102b2: Determine the photometric response curve of the near-field camera based on the image value of each color channel and the image illuminance of each color channel.

[0077] Step 102b3: Determine the vignetting effect of the near-field camera based on the image illuminance of each color channel of the contact surface deformation image.

[0078] Step 102b4: Based on the photometric response curve of the near-field camera and the vignette effect of the near-field camera, establish the camera radiation model.

[0079] Step 102b5: Calibrate the photometric response curve of the camera radiation model based on the photometric response curve calibration method, and calibrate the vignette effect of the camera radiation model based on the vignette effect calibration method.

[0080] For example, the camera radiation model considers the camera's photometric response and vignetting effect, such as... Figure 3 As shown. Because the sensor camera uses an RGB camera, and automatic white balance is turned off during operation, each color channel has its own monotonic camera photometric response, denoted as {F... C} C=R,G,B It is associated with the observed image value {M}. C} C=R,G,B With image illuminance {I C} C=R,G,B M C =F C (I C Therefore, given the image values ​​and image illumination, the camera photometric response {F} for each color channel of the sensor camera can be calculated using the camera photometric response calibration method. C} C=R,G,B The camera's photometric response is used to simulate the camera's photometric response curve. A fifth-order polynomial function is used to describe this curve during calibration. Camera vignetting describes the uneven distribution of image illumination in the image plane and is estimated using a vignetting calibration method.

[0081] In summary, by describing and calibrating the camera's geometric and radiative models, a simulator can be established to simulate the radiance of arbitrary surfaces of the GelSight sensor. Below, the observed image values ​​from the camera. Where (i,j) represents the pixel position of the image, which is determined by the camera geometry model.

[0082] For example, the above near-field light source model includes: a light source geometry model and a light source radiation model.

[0083] Specifically, step 102 above, which involves establishing and calibrating the geometric model and radiation model of the near-field light source, may include the following steps 102c1 and 102c2:

[0084] Step 102c1: Based on the position of each light source of the target sensor, establish the geometric model of the light source, and calibrate the position of each light source in the geometric model of the light source using the light source position calibration method.

[0085] For example, in a GelSight sensor, the light source typically consists of several light-emitting diodes (LEDs) distributed at different locations. During operation, all LEDs are lit simultaneously. Since the size of each LED is much smaller than its working distance, it can be considered a point light source.

[0086] Based on this, the light source geometry model of the aforementioned near-field camera can be described by the point light source positions of the target sensor, i.e. (N is the number of LED light sources, N≥3). The position of each light source can be obtained by the light source position calibration method.

[0087] Step 102c2: Based on the principal optical axis direction of each light source of the target sensor and the relative energy intensity of each light source in different directions, establish the radiation model of the light source, and calibrate the radiation of each light source in the radiation model according to the brightness of each pixel in the deformation image of the contact surface.

[0088] For example, a light source radiation model can describe the relative energy intensity of each LED light source along different directions and the direction of its principal optical axis. This can be expressed by the following formula:

[0089]

[0090] in, This indicates the direction of the principal optical axis of the k-th LED light source, relative to η. k μk and μk are used together as model parameters for the radiation model of the light source. Let k be the direction of the light source received by the surface corresponding to the (i,j)th pixel of the image. The distance between it and the light source is specified. The model parameters of this light source radiation model can be calibrated by analyzing the radiation characteristics of each light source in the model based on the brightness of each pixel in the contact surface deformation image. It should be noted that each LED light source is calibrated individually.

[0091] For example, the above near-field camera model includes: a surface reflection model.

[0092] Specifically, the establishment and calibration of the surface reflection model in step 102 above may include the following steps 102d1 and 102d2:

[0093] Step 102d1: Use the generalized Lambertian reflection model as the surface reflection model.

[0094] Step 102d2: Based on the contact surface deformation image, the surface roughness and surface reflectivity of the surface reflection model are calibrated using the Levenburg-Marquardt algorithm.

[0095] For example, the contact surface of the aforementioned target sensor is made of elastomeric silicone. This silicone material can theoretically be approximated by the Lambertian reflection model. Due to the surface roughness of the silicone, the simulator in this embodiment uses a generalized Lambertian reflection model to describe the silicone reflection. Assuming the contact surface material is spatially uniform, the radiance under a single light source illumination can be expressed by the following formulas three to six:

[0096]

[0097]

[0098]

[0099]

[0100] in, and These represent the incident direction and the exit direction, respectively. σ represents surface roughness, ρ c The surface reflectance, surface roughness, and surface reflectance can be calibrated using the acquired contact surface deformation images.

[0101] Specifically, given the dimensions of the calibration ball and the pressing point, the surface roughness and surface reflectivity in the surface reflection model can be calibrated by fitting the reflection model using the Levenberg-Marquardt algorithm. The overall radiance of the contact surface conforms to the superposition principle, that is, the superposition of various light sources, and can be expressed by the following formula:

[0102]

[0103] For example, in addition to pixel-level local tonal influence factors, global influence factors may be included in the simulator, including secondary reflections and cast shadows. Since GelSight uses a black, low-reflectivity material casing, the effect of secondary reflections is negligible. Therefore, in this embodiment, the classic hidden point removal (HPR) model is used as the cast shadow model.

[0104] Specifically, the establishment and calibration of the surface deformation model in step 102 above may include the following step 102e:

[0105] Step 102e: Calculate the surface deformation of the contact surface when it is pressed based on the intersection relationship between the light and the geometric patch, and use a Gaussian pyramid to smooth the deformed and non-deformed areas.

[0106] For example, when the sensor contact surface is pressed, the simulator first calculates the visible surface of the pressed object using the ray-geometric patch intersection relationship; the non-contact surface can be calculated simultaneously. Due to the continuity of the silicone material, the edges of the contact object are smoothed, and the corresponding smoothing can be simulated using a Gaussian pyramid method.

[0107] Optionally, in this embodiment of the application, after the above-mentioned simulator is established and calibrated, the simulator can be used to generate simulated pressing images, and the target sensor can be calibrated based on the generated simulated pressing images.

[0108] Specifically, the target model mentioned above is a convolutional neural network model, and step 103 may include the following steps 103a and 103b:

[0109] Step 103a: Calculate the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images.

[0110] Step 103b: Train the target model using the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples of the target model.

[0111] The target model is a fully convolutional encoder-decoder architecture, using cosine similarity as the loss function. The encoder includes five convolutional layers to reduce the feature scale of the image by a preset factor to reduce computation. The decoder includes six convolutional layers and an L2 normalization layer to regress the surface normal vector of each pixel and increase the feature scale of the image by the preset factor to restore the image scale.

[0112] Specifically, step 103b above may include steps 103b1 and 103b2:

[0113] Step 103b1: Perform a phase difference operation between the input simulated pressing image and the image of the contact surface in the non-pressed state, and retain the image of the deformed part of the contact surface.

[0114] Step 103b2: Input the image of the deformed part into the target model to obtain the prediction result of the surface normal vector field after the contact surface is deformed.

[0115] For example, the training data for the target model uses 31 different 3D shapes from two datasets: the blobby shape dataset and the tactileshape dataset. For each shape, 432 different pressing angles (3 scales x 12 azimuth angles x 12 elevation angles) are used in the simulator. At each angle, pressing is performed at 27 different positions on the contact surface (3 each for X, Y, and Z). Thus, the simulator simulates pressing at different shapes, angles, and positions a total of 361,584 times, obtaining sensor-simulated pressing images and ground truth data of the surface normal vector field for training. Simultaneously, all samples are divided into training and validation sets at a ratio of 98:2.

[0116] For example, the structure of the target model neural network described above is a fully convolutional encoder-decoder architecture. The sensor-pressed image is first subtracted from its unpressed image, serving as the input to the neural network. The output is the predicted surface normal vector field after pressing. The encoder uses a 5-layer convolutional neural network to reduce the image's feature scale by a factor of 4 to decrease computational load. The decoder then uses a 6-layer convolutional neural network with an L2 normalization layer to regress the surface normal vector of each pixel. The decoder increases the feature scale by a factor of 4 while preserving the original image scale; therefore, this neural network structure can adapt to different sensor image sizes.

[0117] For example, the target model described above, which uses Cosine similarity as the loss function, can be represented by the following formula:

[0118]

[0119] Where H and W represent the height and width of the sensor image, respectively. and These represent the true value and the predicted value of the surface normal vector, respectively.

[0120] For example, the model is trained in PyTorch, and white noise with a variance of 0.01 is added to the images for image augmentation during training. The Adam optimizer is used for solving the problem. The batch size during training is 128.

[0121] For example, the normal vector field of the surface deformation of the GelSight contact surface can be obtained through the trained target model. Furthermore, the geometric deformation of the sensor contact surface can be obtained by integrating the normal vector field. At the same time, the influence of the near-field camera needs to be considered in the integration.

[0122] The vision-based tactile sensor calibration method provided in this application first acquires a contact surface deformation image collected by the target sensor, and then establishes and calibrates a simulator based on the acquired contact surface deformation image. Next, the simulator generates a set of simulated pressing images of the target sensor's contact surface in contact with objects of different shapes. Finally, a target model is trained based on the set of simulated pressing images to complete the calibration of the target sensor, greatly improving the calibration efficiency of vision-based tactile sensors.

[0123] It should be noted that the tactile sensor calibration method based on vision provided in this application can be executed by a tactile sensor calibration device based on vision, or by a control module within that device for executing the calibration method. This application uses the execution of the calibration method by a tactile sensor calibration device based on vision as an example to illustrate the tactile sensor calibration device provided in this application.

[0124] It should be noted that the vision-based tactile sensor calibration methods shown in the accompanying drawings of the embodiments of this application are all illustrated by way of example with reference to one of the accompanying drawings of the embodiments of this application. In specific implementation, the vision-based tactile sensor calibration methods shown in the accompanying drawings of the above methods can also be implemented in conjunction with any other accompanying drawings shown in the above embodiments, which will not be elaborated here.

[0125] The vision-based tactile sensor calibration device provided in this application is described below. The vision-based tactile sensor calibration method described below can be referred to in correspondence with the vision-based tactile sensor calibration method described above.

[0126] Figure 4 A schematic diagram of a vision-based tactile sensor calibration device provided in an embodiment of this application is shown below. Figure 4 As shown, it specifically includes:

[0127] The acquisition module 401 is used to acquire the contact surface deformation image collected by the target sensor; the simulation module 402 is used to complete the establishment and calibration of the simulator based on the acquired contact surface deformation image; the generation module 403 is used to generate a set of simulated pressing images of the contact surface of the target sensor when it comes into contact with objects of different shapes through the simulator; the calibration module 404 is used to train the target model based on the set of simulated pressing images to complete the calibration of the target sensor; wherein, the establishment and calibration of the simulator includes: the establishment and calibration of the near-field camera model, the establishment and calibration of the near-field light source model, the establishment and calibration of the surface reflection model, and the establishment of the surface deformation model; the trained target model is used to restore the deformation image of the contact surface of the target sensor to the three-dimensional geometric deformation of the contact surface.

[0128] Optionally, the control further includes: a control module; the control module is used to control a calibration ball of a preset size to contact the contact surface of the target sensor, and to control the calibration ball to be pressed into a hemispherical position on the contact surface; the acquisition module 401 is specifically used to acquire a deformation image of the contact surface to obtain the deformation image of the contact surface.

[0129] Optionally, the device further includes: a determination module; the near-field camera model includes: a camera geometric model; the acquisition module 401 is further configured to acquire the camera intrinsic parameters of the target sensor; the simulation module 402 is configured to construct a camera coordinate system based on the camera position of the target sensor and a world coordinate system based on the contact surface; the determination module is configured to determine the extrinsic parameters of the near-field camera based on the three-dimensional rigid transformation from the world coordinate system to the camera coordinate system; the simulation module 402 is specifically configured to establish the camera geometric model based on the camera intrinsic parameters and the extrinsic parameters of the near-field camera; the simulation module 402 is also specifically configured to calibrate the camera intrinsic parameters and camera extrinsic parameters of the camera geometric model based on a camera geometric calibration method.

[0130] Optionally, the near-field camera model includes: a camera radiation model; the determining module is further configured to determine the image value of each color channel and the image illuminance of each color channel based on the contact surface deformation image; the determining module is further configured to determine the photometric response curve of the near-field camera based on the image value of each color channel and the image illuminance of each color channel; the determining module is further configured to determine the vignette effect of the near-field camera based on the image illuminance of each color channel of the contact surface deformation image; the simulation module 402 is specifically further configured to establish the camera radiation model based on the photometric response curve of the near-field camera and the vignette effect of the near-field camera; the simulation module 402 is specifically further configured to calibrate the photometric response curve of the camera radiation model based on the photometric response curve calibration method, and calibrate the vignette effect of the camera radiation model based on the vignette effect calibration method.

[0131] Optionally, the near-field light source model includes: a light source geometric model and a light source radiation model; the simulation module 402 is specifically used to establish the light source geometric model according to the position of each light source of the target sensor, and to calibrate the position of each light source in the light source geometric model by means of a light source position calibration method; the simulation module 402 is also specifically used to establish the light source radiation model according to the principal optical axis direction of each light source of the target sensor and the relative energy intensity of each light source in different directions, and to calibrate the radiation of each light source in the light source radiation model according to the brightness of each pixel in the contact surface deformation image.

[0132] Optionally, the determining module is further configured to use a generalized Lambertian reflection model as the surface reflection model; the simulation module 402 is specifically configured to calibrate the surface roughness and surface reflectivity of the surface reflection model based on the contact surface deformation image using the Levenberg-Marquardt algorithm.

[0133] Optionally, the simulation module 402 is specifically used to calculate the surface deformation of the contact surface when it is pressed based on the intersection relationship between the light and the geometric patch, and to use a Gaussian pyramid to smooth the deformed area and the non-deformed area.

[0134] Optionally, the device further includes: a calculation module; the target model is a convolutional neural network model; the calculation module is used to calculate the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images; the calibration module 404 specifically uses the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples to train the target model; wherein, the target model is a fully convolutional encoder-decoder architecture, with cosine similarity as the loss function; the encoder includes: 5 convolutional layers, used to reduce the feature scale of the image by a preset factor to reduce the computational load; the decoder includes: 6 convolutional layers and an L2 normalization layer, used to regress the surface normal vector of each pixel and increase the feature scale of the image by the preset factor to restore the image scale.

[0135] Optionally, the calibration module 404 is specifically used to perform a phase difference operation between the input simulated pressing image and the image of the contact surface in the non-pressed state, and retain the image of the deformed part of the contact surface; the calibration module 404 is also specifically used to input the image of the deformed part into the target model to obtain the prediction result of the surface normal vector field after the contact surface is deformed.

[0136] The vision-based tactile sensor calibration device provided in this application first acquires contact surface deformation images collected by the target sensor, and then completes the establishment and calibration of a simulator based on the acquired contact surface deformation images. Next, the simulator generates a set of simulated pressing images of the target sensor's contact surface when it comes into contact with objects of different shapes. Finally, the target model is trained based on the set of simulated pressing images to complete the calibration of the target sensor, greatly improving the calibration efficiency of vision-based tactile sensors.

[0137] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a vision-based tactile sensor calibration method, which includes: acquiring contact surface deformation images collected by the target sensor, and completing the establishment and calibration of a simulator based on the acquired contact surface deformation images; generating a set of simulated pressing images of the contact surface of the target sensor when it comes into contact with objects of different shapes through the simulator; training a target model based on the set of simulated pressing images to complete the calibration of the target sensor; wherein the establishment and calibration of the simulator includes: the establishment and calibration of a near-field camera model, the establishment and calibration of a near-field light source model, the establishment and calibration of a surface reflection model, and the establishment of a surface deformation model; the trained target model is used to restore the deformation images of the contact surface of the target sensor to the three-dimensional geometric deformation of the contact surface.

[0138] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] On the other hand, this application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by the computer, the computer can execute the vision-based tactile sensor calibration method provided by the above methods. The method includes: acquiring a contact surface deformation image collected by a target sensor, and completing the establishment and calibration of a simulator based on the acquired contact surface deformation image; generating a set of simulated pressing images of the contact surface of the target sensor when it comes into contact with objects of different shapes through the simulator; training a target model based on the set of simulated pressing images to complete the calibration of the target sensor; wherein, the establishment and calibration of the simulator includes: the establishment and calibration of a near-field camera model, the establishment and calibration of a near-field light source model, the establishment and calibration of a surface reflection model, and the establishment of a surface deformation model; the trained target model is used to restore the deformation image of the contact surface of the target sensor to the three-dimensional geometric deformation of the contact surface.

[0140] In another aspect, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the aforementioned vision-based tactile sensor calibration methods. The method includes: acquiring contact surface deformation images collected by a target sensor, and establishing and calibrating a simulator based on the acquired contact surface deformation images; generating a set of simulated pressing images of the target sensor's contact surface in contact with objects of different shapes using the simulator; training a target model based on the set of simulated pressing images to calibrate the target sensor; wherein the establishment and calibration of the simulator includes: establishing and calibrating a near-field camera model, establishing and calibrating a near-field light source model, establishing and calibrating a surface reflection model, and establishing a surface deformation model; the trained target model is used to restore the deformation images of the target sensor's contact surface to the three-dimensional geometric deformation of the contact surface.

[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A vision-based tactile sensor calibration method, characterized in that, Applications in vision-based tactile sensors include: Acquire the contact surface deformation image collected by the target sensor, and complete the establishment and calibration of the simulator based on the acquired contact surface deformation image; The simulator generates a set of simulated pressing images of the target sensor's contact surface coming into contact with objects of different shapes. The target model is trained based on the set of simulated pressing images to complete the calibration of the target sensor; The establishment and calibration of the simulator includes: the establishment and calibration of a near-field camera model, a near-field light source model, a surface reflection model, and a surface deformation model; the trained target model is used to restore the deformation image of the contact surface of the target sensor to the three-dimensional geometric deformation of the contact surface; the near-field camera model includes: a camera geometric model, or a camera radiation model; the camera geometric model is calibrated by calibrating the camera intrinsic and extrinsic parameters of the camera geometric model based on the camera geometric calibration method; the camera radiation model is calibrated by calibrating the camera radiation model based on the photometric response curve calibration method and the vignette effect calibration method; The near-field light source model includes: a light source geometric model and a light source radiation model; the establishment and calibration of the near-field light source model includes: Based on the position of each light source of the target sensor, a geometric model of the light source is established, and the position of each light source in the geometric model is calibrated using a light source position calibration method; based on the principal optical axis direction of each light source of the target sensor and the relative energy intensity of each light source in different directions, a radiation model of the light source is established, and the radiation of each light source in the radiation model is calibrated based on the brightness of each pixel in the contact surface deformation image; The establishment and calibration of the surface reflection model includes: The generalized Lambertian reflection model is adopted as the surface reflection model; based on the contact surface deformation image, the surface roughness and surface reflectivity of the surface reflection model are calibrated using the Levonburg-Marquardt algorithm; The establishment of the surface deformation model includes: The surface deformation of the contact surface when it is pressed is calculated based on the intersection relationship between light rays and geometric patches, and the deformed and non-deformed areas are smoothed using a Gaussian pyramid. The step of training a target model based on the set of simulated pressure images to calibrate the target sensor includes: Calculate the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images; use the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples for the target model to train the target model.

2. The method according to claim 1, characterized in that, The acquisition of the contact surface deformation image collected by the target sensor includes: A calibration ball of a preset size is controlled to contact the contact surface of the target sensor, and the calibration ball is controlled to be pressed into a hemispherical position on the contact surface; The deformation image of the contact surface is obtained by acquiring the deformation image of the contact surface.

3. The method according to claim 2, characterized in that, The near-field camera model includes: a camera geometric model; The process of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: Obtain the camera intrinsic parameters of the target sensor, construct a camera coordinate system based on the camera position of the target sensor, and construct a world coordinate system based on the contact surface; The extrinsic parameters of the near-field camera are determined based on the three-dimensional rigid transformation from the world coordinate system to the camera coordinate system. Based on the camera's intrinsic parameters and the near-field camera's extrinsic parameters, the camera's geometric model is established; The camera intrinsic and extrinsic parameters of the camera geometric model are calibrated based on the camera geometry calibration method.

4. The method according to claim 2, characterized in that, The near-field camera model includes: a camera radiation model; The process of establishing and calibrating the simulator based on the acquired contact surface deformation image includes: Based on the contact surface deformation image, determine the image value and image illuminance of each color channel; The photometric response curve of the near-field camera is determined based on the image value of each color channel and the image illuminance of each color channel; The vignetting effect of the near-field camera is determined based on the image illuminance of each color channel of the contact surface deformation image; Based on the photometric response curve of the near-field camera and the vignetting effect of the near-field camera, a radiation model of the camera is established. The photometric response curve of the camera radiation model is calibrated using the photometric response curve calibration method, and the vignette effect of the camera radiation model is calibrated using the vignette effect calibration method.

5. The method according to claim 1, characterized in that, The target model is a convolutional neural network model; the target model is a fully convolutional encoder-decoder architecture, using cosine similarity as the loss function; the encoder includes: 5 convolutional layers, used to reduce the feature scale of the image by a preset factor to reduce the computational load; the decoder includes: 6 convolutional layers and an L2 normalization layer, used to regress the surface normal vector of each pixel and increase the feature scale of the image by the preset factor to restore the image scale.

6. The method according to claim 5, characterized in that, The step of training the target model using the set of simulated compression images and the surface normal vector field corresponding to each simulated compression image in the set as training samples includes: The input simulated pressing image is compared with the image of the contact surface in the non-pressed state, and the image of the deformed part of the contact surface is retained. The image of the deformed portion is input into the target model to obtain the prediction result of the surface normal vector field after the contact surface is deformed.

7. A vision-based tactile sensor calibration device, characterized in that, The device includes: The acquisition module is used to acquire the contact surface deformation image collected by the target sensor; The simulation module is used to build and calibrate the simulator based on the collected deformation images of the contact surface; The generation module is used to generate a set of simulated pressing images of the contact surface of the target sensor when it comes into contact with objects of different shapes through the simulator; The calibration module is used to train a target model based on the set of simulated pressing images and to complete the calibration of the target sensor. The establishment and calibration of the simulator includes: the establishment and calibration of a near-field camera model, a near-field light source model, a surface reflection model, and a surface deformation model; the trained target model is used to restore the deformation image of the contact surface of the target sensor to the three-dimensional geometric deformation of the contact surface; the near-field camera model includes: a camera geometric model, or a camera radiation model; the camera geometric model is calibrated by calibrating the camera intrinsic and extrinsic parameters of the camera geometric model based on the camera geometric calibration method; the camera radiation model is calibrated by calibrating the camera radiation model based on the photometric response curve calibration method and the vignette effect calibration method; The near-field light source model includes a light source geometric model and a light source radiation model. The simulation module is specifically used to establish the light source geometric model based on the position of each light source in the target sensor, and to calibrate the position of each light source in the light source geometric model using a light source position calibration method. The simulation module is also specifically used to establish the light source radiation model based on the principal optical axis direction of each light source in the target sensor and the relative energy intensity of each light source in different directions, and to calibrate the radiation of each light source in the light source radiation model based on the brightness of each pixel in the contact surface deformation image. The determination module is also used to adopt the generalized Lambertian reflection model as the surface reflection model; the simulation module is specifically used to calibrate the surface roughness and surface reflectivity of the surface reflection model based on the contact surface deformation image using the Levonburg-Marquardt algorithm. The simulation module is specifically used to calculate the surface deformation of the contact surface when it is pressed based on the intersection relationship between the light and the geometric patch, and to use a Gaussian pyramid to smooth the deformed and non-deformed areas. The device further includes: a calculation module; the calculation module is used to calculate the surface normal vector field of the contact surface corresponding to each simulated pressing image in the set of simulated pressing images; the calibration module specifically uses the set of simulated pressing images and the surface normal vector field corresponding to each simulated pressing image in the set of simulated pressing images as training samples of the target model to train the target model.

Citation Information

Patent Citations

  • Device and method for orchestrating display surfaces, projection devices and 2D and 3D spatialized interaction devices for creating interactive environments

    FR3025917A1

  • Method for acquiring normal vector, geometry and material of three-dimensional object employing neural network

    WO2021042277A1