Visual tactile sensor, geometric calibration method and computer equipment

By using a novel structure and calibration method for visual-tactile sensors, combined with a near-light photometric 3D physical model and a lightweight neural network, the problems of equipment dependence and insufficient accuracy of traditional calibration techniques are solved, achieving efficient and universal 3D reconstruction and real-time perception, suitable for a variety of application scenarios.

CN121558113APending Publication Date: 2026-02-24SHANGHAI TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511798221.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing geometric calibration techniques for visual-tactile sensors rely on expensive equipment, involve complex processes, and are inefficient. They cannot accurately model the non-uniform illumination characteristics of near-field point light sources inside the sensor, resulting in limited 3D reconstruction accuracy. Furthermore, they lack universal adaptability to different sensor shapes, which restricts the rapid iteration and practical deployment of high-performance visual-tactile sensors.

Method used

The visual-tactile sensor structure, consisting of a base, a light source module, and an elastic housing, combines a near-light photometric three-dimensional physical model and a lightweight normal prediction neural network. By acquiring image data through multiple presses, a surface normal vector model is established, simplifying the calibration process and improving accuracy. It is suitable for sensors of various geometric shapes.

Benefits of technology

It achieves efficient and low-cost geometric calibration, improves the accuracy and adaptability of 3D reconstruction, supports diverse application scenarios, meets the real-time requirements of robot dexterity operation and medical tactile feedback, and has industrialization potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121558113A_ABST
    Figure CN121558113A_ABST
Patent Text Reader

Abstract

The invention provides a visual tactile sensor, a geometric calibration method and computer equipment, and the visual tactile sensor comprises a base which is internally provided with an imaging unit; the light source module is arranged on the base and comprises 3N single-point light sources which are arranged around the imaging unit and can be independently controlled, and N is larger than or equal to 1; and the elastic housing is arranged on the base in a covering manner, and the imaging unit and the light source module are wrapped in the elastic housing. By constructing the novel curved surface visual tactile sensor and the matched geometric calibration method thereof, the problems that a traditional plane sensor is weak in multi-direction contact sensing capacity, difficult in curved surface calibration and the like are effectively solved, data generated by a physical model are used for training a lightweight neural network, and the accuracy of the curved surface visual tactile sensor is improved. Real-time surface normal prediction under single-frame three-color image input is realized, and the precision and reasoning speed of three-dimensional reconstruction are considered while hardware dependence and calibration cost are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot perception and intelligent control technology, and specifically relates to a visual-tactile sensor, a geometric calibration method, and a computer device. Background Technology

[0002] Curved surface tactile sensors, due to their biomimetic structure that more closely resembles human fingertips, offer significant advantages in dexterous robotic manipulation, envelope grasping, and complex contact tasks. They enable omnidirectional, high-resolution tactile perception and represent an important development direction for service robots, medical, and industrial applications. However, their curved geometry introduces new challenges such as uneven illumination and calibration difficulties, making traditional methods based on planar structures and the assumption of parallel light unsuitable.

[0003] Existing geometric calibration techniques generally suffer from problems such as reliance on expensive equipment (e.g., CNC or standard probes), complex processes, and low efficiency. Furthermore, most methods fail to accurately model the non-uniform illumination characteristics of near-field point sources within the sensor (e.g., distance attenuation, directionality, and occlusion), limiting the accuracy of 3D reconstruction. Simultaneously, these methods lack universal adaptability to different sensor shapes (especially curved surfaces), severely hindering the rapid iteration and practical deployment of high-performance visual-tactile sensors. Therefore, a novel calibration method is urgently needed that is lightweight, efficient, requires no dedicated hardware, and can accurately model the physical characteristics of near-field light sources. Summary of the Invention

[0004] In view of the shortcomings of the prior art, the purpose of this invention is to provide a visual-tactile sensor, a geometric calibration method, and a computer device to solve the problems of existing geometric calibration techniques, such as reliance on expensive equipment, complex processes, and low efficiency. Moreover, most methods cannot accurately model the non-uniform illumination characteristics of the near-field point light source inside the sensor, resulting in limited accuracy of 3D reconstruction. At the same time, these methods lack universal adaptability to different sensor shapes, which seriously restricts the rapid iteration and practical deployment of high-performance visual-tactile sensors.

[0005] To achieve the above and other related objectives, the present invention proposes a visual-tactile sensor, characterized in that it comprises:

[0006] A base in which an imaging unit is installed;

[0007] A light source module, which is disposed on the base, includes K independently controllable single-point light sources arranged around the imaging unit, where K ≥ 3;

[0008] An elastic cover is provided on the base and encloses the imaging unit and the light source module.

[0009] In one embodiment of the present invention, the base includes a base and a sensor housing, the sensor housing has a mounting hole in the middle, and the imaging unit is mounted on the base and located in the mounting hole.

[0010] In one embodiment of the present invention, the sensor housing has a groove formed around the mounting hole, and the light source module is mounted in the groove, with its height lower than the lens surface of the imaging unit.

[0011] In one embodiment of the present invention, the light source module includes a ring-shaped circuit board, which is arranged around the imaging unit, and K independently controllable single-point light sources are uniformly arranged along its circumference.

[0012] In one embodiment of the present invention, the light source module includes 3N independently controlled single-point light sources, where N≥1.

[0013] In one embodiment of the invention, each of the single point light sources may emit white, red, blue, and green light.

[0014] In one embodiment of the present invention, the elastic cover includes a transparent elastomer and a light-shielding layer, the light-shielding layer covering the outer surface of the transparent elastomer.

[0015] In one embodiment of the present invention, the transparent elastomer is made of transparent silicone material.

[0016] This invention also proposes a geometric calibration method based on a visual-tactile sensor, comprising:

[0017] Press the elastic cover to collect multiple sets of light source images under different pressing conditions. The acquisition process for each set of light source images is to press the elastic cover. Under each pressing, one dark field image, K single light source images under a single white light source, and one tri-color illumination image are acquired. Among them, the single light source images are the images acquired when different single light sources are lit in sequence, and K is the number of single light sources.

[0018] A three-dimensional physical model of near-light photometric intensity decays with distance and angle is established, and the surface normal vector under each press is calculated based on the dark field image in multiple light source images and K single light source images.

[0019] The surface normal vectors calculated from each set of light source images and the corresponding tri-color illumination images are used as a set of training data to establish a training dataset;

[0020] A lightweight normal prediction neural network model is constructed, and the three-color illumination images of each training set in the training dataset are used as inputs and the surface normal vectors are used as outputs to train and optimize the lightweight normal prediction neural network model to obtain the optimized lightweight normal prediction neural network model.

[0021] In one embodiment of the present invention, the step of establishing a near-light photometric three-dimensional physical model in which light intensity decreases with distance and angle, and calculating the surface normal vector under each press based on the dark field image in multiple sets of light source images and K single light source images includes:

[0022] Construct a pixel-level lighting-normal relationship model:

[0023]

[0024] Wherein, the observed brightness I of the i-th LED at color channel c relative to pixel p i,c The physical imaging process of (p). Where Ψ t,c ρ represents the calibrated light intensity of the light source in this channel. c (p) is the surface albedo at pixel p; x(p) and Let n(p) represent the three-dimensional positions of a surface point and the light source, respectively, and let n(p) be the normal vector of that point. This represents the main emission direction of the light source. (First term) Used to model the directional attenuation of LEDs, where μ i This is a directional index. The second item... Describe Lambert reflection, and through {} - Clip backlighting to handle shadows and occlusions; denominator This expresses the distance attenuation of a near-light point source;

[0025] Based on logarithmic depth variable Construct the energy function using albedo:

[0026]

[0027] The physical depth is obtained by solving the dark field image and K single light source images from multiple light source images, based on the pixel-level illumination-normal relationship model, energy function and number of depth variables;

[0028] Take the depth derivative of the physical depth: The surface normal vector is obtained by calculation;

[0029] in, This represents the intensity of the observed image. For each pixel p, this is the actual image intensity from the sensor, based on the color channel c and the light source i.

[0030] Represents an optimized depth map and albedo Predicted image intensity.

[0031] This represents the optimized depth value.

[0032] This represents the prior depth used for the initial depth estimate in depth regularization, which is usually derived from calibration data or preliminary estimates.

[0033] The present invention also proposes a computer device including a processor coupled to a memory, the memory storing a lightweight normal prediction neural network model optimized according to any one of the geometric calibration methods based on visual-touch sensors in the above embodiments, wherein when the program instructions stored in the memory are executed by the processor, the normal vector of the contact surface is predicted.

[0034] This invention proposes a visual-tactile sensor, a geometric calibration method, and a computer device, which have the following significant advantages and positive effects:

[0035] The sensor boasts significant structural advantages: by employing a curved, elastic housing and 3N independently controllable point light sources arranged around the imaging unit, omnidirectional tactile perception in a biomimetic form is achieved. This design not only better conforms to the natural contact patterns of human fingertips or palms but also stably acquires high-resolution contact information from multiple directions and angles, significantly improving its adaptability to irregular objects, flexible objects, and envelope-like grasping tasks. Furthermore, the sensor's modular structure facilitates easy assembly, and the elastic housing can be flexibly customized into different geometric shapes such as fingertip, hemispherical, and ellipsoidal to meet diverse application requirements.

[0036] Significantly improved calibration accuracy: The proposed geometric calibration method abandons the traditional parallel light assumption and introduces for the first time in the field of visual and tactile perception a near-light photometric three-dimensional physical model that considers the inverse square attenuation of distance and angle dependence. This model truly reflects the illumination characteristics of the point light source inside the sensor, fundamentally reducing the normal vector reconstruction error caused by the mismatch of illumination modeling, and greatly improving the accuracy of three-dimensional shape restoration.

[0037] The calibration process is low-cost, highly efficient, and highly versatile: the entire calibration process does not rely on a CNC platform, standard probes, or precision indentation devices. High-quality training data can be collected simply by using an ordinary object to perform dozens of random presses. The operation is simple, time-saving, and low-cost. This method is not dependent on a specific sensor shape and is applicable to various curved or non-planar visual-tactile sensors. It has excellent versatility and rapid transfer capabilities, which is conducive to product iteration and large-scale promotion.

[0038] Balancing accuracy and real-time performance with industrialization potential: By combining physics-driven normal vector calculation with lightweight neural networks, a hybrid calibration paradigm of "physics-guided - data-driven" is constructed. The trained model can predict surface normal vectors in real time on embedded devices with millisecond-level latency, meeting the real-time requirements of robot online operation while ensuring high accuracy. This provides a reliable and deployable technical solution for high-value scenarios such as dexterous grasping of service robots, tactile feedback in minimally invasive surgery, and industrial quality inspection, and has broad engineering application prospects and industrialization value. Attached Figure Description

[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 This is a schematic diagram of the overall structure of a near-light photometric stereoscopic visual-tactile sensor in one embodiment of the present invention.

[0041] Figure 2 This is a schematic diagram of the near-light photometric three-dimensional illumination model and its geometric relationship.

[0042] Figure 3 This is a flowchart of the geometric calibration method for visual-tactile sensors. Detailed Implementation

[0043] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0044] It should be noted that the illustrations provided in this embodiment are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components relevant to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0045] The importance of curved surface tactile sensors lies in their better alignment with biomimetic and application requirements. Most existing sensors are planar in structure, which, while providing contact information, are insufficient for dexterous manipulation and complex tasks. Curved sensors, more closely resembling the shape of human fingertips and palms, enable stable contact and grasping, and provide richer tactile information from multiple directions and angles, making them particularly suitable for manipulating enveloped and irregular objects. In service robots, medical robots, and industrial applications, curved surface tactile sensors are considered a crucial direction for future development due to their omnidirectional sensing and higher resolution. However, curved structures also bring new challenges such as uneven illumination distribution and difficulties in geometric calibration, making traditional methods difficult to apply directly.

[0046] Geometric calibration of visual-tactile sensors is a core step in ensuring sensor performance. Visual-tactile sensors can acquire high-resolution geometric shapes during contact and are widely used in robotic grasping, manipulation, human-computer interaction, and medical fields. To achieve high-precision perception, geometric calibration must establish a mapping relationship between illumination intensity and surface normal vectors, thereby ensuring the accuracy of 3D reconstruction. Inaccurate calibration will lead to unstable output, limiting its application in complex environments.

[0047] However, existing calibration methods have limitations in terms of illumination modeling, equipment dependence, adaptability, and efficiency. Most methods are based on the photometric stereo principle, but typically employ the parallel light assumption. In reality, the internal light source of a sensor is a near-field point source, and the light intensity decreases non-linearly with angle and distance, leading to large calibration errors. Existing methods also rely on CNC machining, probes, or indentation devices, which are expensive and complex, hindering widespread adoption. Furthermore, different sensor geometries often require redesigned experimental setups, resulting in insufficient versatility. Coupled with low data acquisition efficiency and the need for numerous repetitive operations, the time and cost are high, hindering rapid iteration and widespread adoption. Specific problems are as follows:

[0048] Most existing calibration methods assume that the light source is parallel light. However, the light source used inside the sensor is actually a near-field point source, which causes the light intensity to decrease non-linearly with angle and distance. This inaccurate assumption leads to significant errors in geometric reconstruction.

[0049] Common calibration methods typically rely on expensive and complex equipment such as CNC (Computer Numerical Control) machines, standard probes, or indentation devices. This not only increases costs but also reduces data acquisition efficiency and limits the widespread adoption and application of these methods.

[0050] For sensors with different geometries, especially curved surfaces, existing methods often require redesigning experimental setups and procedures. This indicates that existing methods are insufficient in terms of adaptability and versatility, making it difficult to support diverse sensor design and application requirements.

[0051] Currently, most sensors still use a planar structure, which cannot meet the needs of robots performing tasks such as envelope grasping, multi-directional contact, and handling complex tasks. Curved surfaces or designs that are more in line with biomimetic principles can provide richer tactile information, improving the flexibility and accuracy of operation.

[0052] To address the aforementioned problems, there is an urgent need to develop a new geometric calibration method. This method should: accurately model the illumination characteristics of near-field point light sources, considering the variation of illumination intensity with angle and distance; reduce reliance on expensive equipment during calibration, simplify the operation process, and improve data acquisition efficiency; possess high adaptability and versatility, applicable to sensors with various geometric shapes; and support the design of curved surfaces or other non-traditional structures to better meet the application needs of biomimetic and complex tasks. Therefore, this invention proposes a visual-tactile sensor, a geometric calibration method, and a computer device.

[0053] Please see Figure 1 and Figure 2 As shown, this invention provides a visual-tactile sensor, the overall structure of which includes a base 1, an imaging unit 2, a light source module 3, and an elastic housing 4. This sensor is specifically designed to solve the problems faced by existing visual-tactile systems in surface geometry calibration, such as inaccurate lighting modeling, strong device dependence, and poor versatility. It is particularly suitable for applications such as robot dexterity manipulation, envelope grasping, and interaction with complex objects.

[0054] Please see Figure 1 and Figure 2 As shown, in this embodiment, the base 1 is made of a transparent rigid material and has an internal mounting cavity for fixing a high-resolution industrial camera as the imaging unit 2. The imaging unit 2 is preferably a global shutter CMOS camera, which features high frame rate and low noise, enabling it to clearly capture deformation images of the contact area under different lighting conditions. The camera's optical axis is perpendicular to the upper surface of the base 1, ensuring that the imaging area covers the entire bottom contact surface of the elastic cover 4.

[0055] Please see Figure 1 and Figure 2As shown, in this embodiment, the light source module 3 is disposed on the upper surface of the base 1 and arranged in a ring around the imaging unit 2. It includes K independently controllable single-point light sources 31, where K≥3. Each single-point light source 31 is preferably an RGB tri-color LED, which can emit red, green, and blue monochromatic light individually or white light. All single-point light sources 31 are precisely time-controlled by a microcontroller to achieve various lighting modes such as individual light-up, group light-up, or full light-up. This layout ensures that at least one single-point light source 31 can effectively illuminate the contact area when contacted from any direction, thereby obtaining multi-view illumination information. Preferably, it includes 3N independently controllable single-point light sources 31 (N≥1, usually N=2~8, i.e., 6~24 LEDs), which is beneficial for grouping and acquiring tri-color illumination images.

[0056] Please see Figure 1 and Figure 2 As shown, in this embodiment, the elastic housing 4 is integrally molded from a high-transmittance, high-elasticity silicone material, such as Solaris, and is placed on the base 1, completely enclosing the imaging unit 2 and the light source module 3. The elastic housing 4 can be processed into a hemispherical, ellipsoidal, fingertip-inspired, or other curved geometric shapes to match the omnidirectional perception requirements of different application scenarios.

[0057] Please see Figure 1 and Figure 2 As shown, in this embodiment, the base 1 includes a base 11 and a sensor housing 12. The base 11 is a rigid support structure made of high-strength engineering plastic, used to support internal electronic components and provide mechanical stability. The sensor housing 12 is fixed above the base 11 and is generally shaped like a ring platform, with a circular or square mounting hole 13 at its center. The imaging unit 2 is securely mounted on the base 11 by screws or adhesive, with its lens optical axis coaxial with the mounting hole 13. The front end of the lens extends into the mounting hole 13, and its photosensitive surface is located in the central area of ​​the mounting hole 13 to ensure that the field of view completely covers the contact area at the bottom of the elastic cover 4.

[0058] Please see Figure 1 and Figure 2As shown, in this embodiment, an annular groove 14 is machined around the inner edge of the mounting hole 13 on the sensor housing 12. The cross-section of the groove 14 can be rectangular, trapezoidal, or arc-shaped, and the depth is designed according to the size of the light source, typically 1-3 mm. Each point light source 31 in the light source module 3 is uniformly embedded around the groove 14 and fixed by a flexible circuit board or PCB bracket. Crucially, the height of the emitting surface of all point light sources 31 is precisely controlled so that it is 0.2-1.0 mm lower than the lens surface of the imaging unit 2. This height difference can effectively avoid specular reflection or saturation overexposure caused by the light source directly hitting the lens, while ensuring that the light can still fully illuminate the contact deformation area after diffuse reflection by the inner wall of the elastic cover 4, thereby maintaining high signal-to-noise ratio image acquisition while suppressing optical interference.

[0059] Please see Figure 1 and Figure 2 As shown, in this embodiment, the light source module 3 includes a ring-shaped circuit board 32, which is a flexible or rigid printed circuit board (FPCB or PCB). Its inner diameter is slightly larger than the outer contour of the imaging unit 2, and its outer diameter matches the mounting area of ​​the base 1. The entire ring-shaped circuit board is concentrically arranged around the imaging unit 2. On the upper surface or side edge of the ring-shaped circuit board 32, 3N independently controllable single-point light sources 31 (N≥1) are evenly distributed along its circumference. Each single-point light source 31 is connected to the driver chip through circuit traces to realize lamp-by-lamp addressing and brightness control. The ring-shaped circuit board 32 is fixed to the corresponding mounting position of the base 1 (such as in the groove 14) by means of clips, adhesives, or screws to ensure the positional accuracy of the light source and long-term working stability. This structure not only simplifies the wiring and assembly process, but also ensures the symmetry of the illumination angle and the uniformity of spatial coverage, providing high-quality multi-view illumination conditions for subsequent geometric calibration based on near-light photometric stereo.

[0060] Please see Figure 1 and Figure 2 As shown, in this embodiment, the elastic housing 4 includes a transparent elastomer 41 and a light-shielding layer 42 covering its outer surface. The transparent elastomer 41 is made of a flexible material with high light transmittance and low hysteresis, covering the imaging unit 2 and the light source module 3 to ensure that external contact force can be effectively transmitted to the internal imaging area. The light-shielding layer 42 is an opaque coating or film, preferably a gray coating, uniformly covering the entire outer surface of the transparent elastomer 41. This light-shielding layer 42 is used to block ambient light interference, prevent external stray light from entering the sensor and affecting image acquisition, and does not affect the deformation characteristics of the elastomer under pressure, thereby ensuring the signal-to-noise ratio and stability of tactile perception.

[0061] Please see Figure 1 and Figure 2As shown, in this embodiment, the transparent elastomer 41 is integrally cast from a transparent silicone material, such as Solaris, which possesses excellent optical transparency, mechanical elasticity, and chemical stability. Its overall geometry can be flexibly designed according to application requirements, specifically as any one of fingertip, pointed, hemispherical, or ellipsoidal shapes. For example, a fingertip-shaped structure is used in human-hand-like operation scenarios to match the curvature of human fingertips; a pointed shape is used in precision insertion tasks to improve local contact resolution; and a hemispherical or ellipsoidal shape is selected in envelope grasping tasks to expand the contact coverage area. Transparent elastomers 41 of different shapes can be compatiblely assembled with the same base 1, demonstrating the advantages of this invention in terms of structural adaptability and application generalization.

[0062] Please see Figure 1 , Figure 2 and Figure 3 As shown, the present invention also proposes a geometric calibration method based on a visual-tactile sensor, comprising:

[0063] S1, elastic cover, collects multiple sets of light source images under different pressing conditions. The acquisition process of each set of light source images is to press the elastic cover. Under this pressing, one dark field image, K single light source images under a single white light source, and one tri-color illumination image are collected. Among them, the single light source image is the image collected when different single light sources are lit in sequence, and K is the number of single light sources.

[0064] S2, a near-light photometric three-dimensional physical model with distance and angle attenuation, and calculates the surface normal vector under each press based on the dark field image in multiple light source images and K single light source images;

[0065] S3. The surface normal vectors calculated from the source image and the corresponding tri-color illumination images are used as a set of training data to establish a training dataset.

[0066] S4. Quantize the normal prediction neural network model, and use the three-color illumination images in each training set of the training dataset as input and the surface normal vector as output to train and optimize the lightweight normal prediction neural network model to obtain the optimized lightweight normal prediction neural network model.

[0067] In step S1, firstly, during the calibration preparation stage, the visual-tactile sensor to be calibrated is fixed to the experimental platform. Then, several everyday objects (such as metal balls, screws, rubber cubes, pen caps, etc.) are used to perform multiple light pressing operations on the elastic cover 4, keeping the object stationary each time and completing the full image acquisition process. For each pressing, the following image acquisition process is executed:

[0068] Acquire a dark field image, i.e., turn off all single point light sources 31, to correct camera noise and ambient light residue;

[0069] K single-point light sources 31 (K=3N,N≥1) are lit sequentially, each light source emits white light individually, and K corresponding single-point light source images are collected respectively;

[0070] Finally, all single-point light sources 31 are simultaneously illuminated and controlled in groups of red, green, and blue (each group contains K / 3 light sources of the same color), and a three-color illumination image (RGB image) containing spatial color coding information is acquired. Thus, each press obtains K+2 images, forming a complete set of observation data.

[0071] In step S2, a photometric stereo physical model considering the near-light effect is established. Since the internal light source of the sensor is a near-distance point source, the traditional parallel light assumption does not hold. Therefore, a near-light photometric stereo (NLiPs) model is adopted to explicitly model the inverse square decay of light intensity with respect to the distance between the light source and the surface, the directional decay of the light source (determined by the LED emission angle), and the Lambertian reflection relationship between the surface normal and the angle between the incident light. Combining the known precise positions and orientations of each single point source 31 in three-dimensional space, this physical model is used to jointly optimize the dark field image and K single-source light source images under each press, solving for the high-precision surface normal vector corresponding to each pixel in the contact area.

[0072] In step S2, the step of establishing a near-light photometric three-dimensional physical model of light intensity attenuation with distance and angle, and calculating the surface normal vector under each press based on the dark field image in multiple sets of light source images and K single light source images includes:

[0073] S21. Construct a pixel-level lighting-normal relationship model:

[0074]

[0075] This formula describes the observed brightness I of the i-th LED at color channel c relative to pixel p under near-light source conditions. i,c The physical imaging process of (p). Wherein, the observed brightness I of the i-th LED at color channel c for pixel p. i,c The physical imaging process of (p). Where Ψ i,c ρ represents the calibrated light intensity of the light source in this channel. c (p) is the surface albedo at pixel p; x(p) and Let n(p) represent the three-dimensional positions of a surface point and the light source, respectively, and let n(p) be the normal vector of that point. This represents the main emission direction of the light source. (First term) Used to model the directional attenuation of LEDs, where μ i This is a directional index. The second item... Describe Lambert reflection, and through {} +Clip backlighting to handle shadows and occlusions; denominator This expresses the distance attenuation of a near-light point source;

[0076] Understandably, this model simultaneously characterizes light source intensity, albedo, directionality, normal constraints, shadows, and distance attenuation, providing a complete physical description of the internal lighting mechanism of a tactile sensor.

[0077] S22, Based on the logarithmic depth variable Construct the energy function using albedo:

[0078]

[0079] This represents the intensity of the observed image. For each pixel p, this is the actual image intensity from the sensor, based on the color channel c and the light source i.

[0080] Represents an optimized depth map and albedo Predicted image intensity.

[0081] This represents the optimized depth value.

[0082] This represents the prior depth used for the initial depth estimate in depth regularization, which is usually derived from calibration data or preliminary estimates.

[0083] The energy function jointly calculates depth and reflectivity by minimizing the error between the observed and predicted images and the depth regularization term.

[0084] S23. Based on the dark field image and K single light source images from multiple light source images, and according to the pixel-level illumination-normal relationship model, energy function and number of depth variables, the physical depth is obtained.

[0085] S24. Perform a depth derivative on the physical depth: The surface normal vector is obtained by calculation.

[0086] The optimization is solved using the Alternating Reweighted Least Squares (ARLS) algorithm: First, the albedo is analytically updated at a fixed depth; then, with the albedo fixed, the illumination model is linearized, and a new depth estimate is obtained by solving the large-scale sparse linear system using the Gauss-Newton method and PCG. After the optimization converges, the depth is calculated using... We obtain the physical depth, and then use the partial derivative of depth: Calculate the surface normal vector for each pixel. Then, backproject the depth map onto the camera coordinate system to obtain a dense 3D point cloud, thereby constructing a high-fidelity model of the contact area.

[0087] In step S3, a training dataset is constructed. The three-color illumination images obtained from each set of presses are used as input samples, and the corresponding surface normal vector maps calculated by the NLiPs model are used as supervision labels (output samples). These are paired one-to-one to form training samples. The entire dataset only requires approximately 50 random presses to cover the main deformation patterns of the sensing area, without requiring knowledge of the object's geometry or a precise positioning device.

[0088] In step S4, a lightweight normal prediction neural network model (denoted as NLiPsNet) is constructed. Its backbone structure uses a three-layer fully connected multilayer perceptron (MLP), with hidden layer dimensions of 256-256-128 respectively. Each layer is followed by BatchNormalization, ReLU activation function, and Dropout (dropout rate p = 0.2) to improve generalization ability. During training, a pixel-level cosine similarity loss function is used as the target. The pixel-level cosine similarity loss function is as follows:

[0089]

[0090] in For predicting the network normal, n i The surface normal vector obtained in step S2 is used to minimize the angular error between the predicted normal vector and the NLiPs reconstructed normal vector. End-to-end training is performed using the AdamW optimizer until the loss converges. After training, the resulting optimized lightweight normal prediction neural network model can be deployed in embedded systems. In actual operation, only a single frame of three-color illumination image needs to be input to output high-precision surface normal vectors in real time, achieving efficient and low-latency 3D tactile reconstruction.

[0091] This embodiment provides a computer device for running the geometric calibration results of a visual-tactile sensor. Its hardware structure includes a processor and a memory coupled thereto. The processor can be a general-purpose central processing unit (CPU), an embedded microcontroller (such as the ARM Cortex-A series), a graphics processing unit (GPU), or a dedicated neural network acceleration chip (such as an NPU or TPU), possessing the ability to perform image processing and neural network inference.

[0092] The memory includes non-volatile storage units (such as Flash, eMMC, or SD card) and RAM (such as DDR4 / DDR5), wherein the non-volatile storage units pre-store a lightweight normal prediction neural network model trained and optimized according to the geometric calibration method described in the above embodiments. This model takes a three-color illumination image (RGB format) as input and outputs a corresponding pixel-level surface normal vector map.

[0093] In actual operation, when an external object contacts the elastic housing of the visual-tactile sensor, the sensor acquires a frame of tri-color illumination image and transmits it to the computer device. The processor loads program instructions from memory, calls the deployed lightweight normal prediction neural network model, performs forward inference on the input image, and generates the surface normal vector distribution of the contact area in real time. This normal vector result can be used for subsequent robot perception tasks such as 3D shape reconstruction, contact force estimation, slip detection, or grasping stability assessment.

[0094] This computer device can be integrated into the robot end effector, edge computing box, or embedded main control board. It features low power consumption, high real-time performance, and strong deployment compatibility, making it suitable for various application scenarios such as service robots, medical operating platforms, and industrial automation production lines.

[0095] In one embodiment, to verify the adaptability of the method of the present invention, various geometric shapes of curved surface visual-tactile sensors were designed and fabricated, including fingertip-shaped, pointed, and ellipsoidal structures. The elastomers of different structures are all made of transparent silicone material. Six to 24 independently controllable RGB point light sources can be arranged internally, distributed in a ring to ensure uniform illumination distribution in different contact areas. Imaging units are installed at the center of the bottom of the elastomer to acquire contact images under multi-light source illumination. The calibration steps are the same as in the above embodiment, establishing geometric mapping relationships through a near-light photometric stereo model, and conducting pressing experiments with various objects. The results show that the method can maintain the stability of normal estimation on surfaces of different shapes, verifying the versatility of the method.

[0096] This invention effectively solves the problems of weak multi-directional contact sensing capability and difficult surface calibration of traditional planar sensors by constructing a novel curved surface visual-tactile sensor and its matching geometric calibration method. The sensor adopts a transparent elastic shell and a ring-shaped distribution of near-field point light sources to support omnidirectional high-resolution tactile sensing; the calibration method is based on a near-light photometric three-dimensional physical model, which does not require expensive equipment and can achieve high-precision normal vector reconstruction simply by pressing an ordinary object.

[0097] Furthermore, by using data generated from the physical model to train a lightweight neural network, real-time surface normal prediction under a single-frame three-color image input was achieved. This significantly reduced hardware dependence and calibration costs while maintaining both the accuracy of 3D reconstruction and inference speed. This technology combines versatility, efficiency, and deployability, providing reliable support for practical applications such as dexterous robot manipulation and medical haptic feedback.

[0098] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A visual-tactile sensor, characterized in that, include: A base in which an imaging unit is installed; A light source module, which is disposed on the base, includes K independently controllable single-point light sources arranged around the imaging unit, where K ≥ 3; An elastic cover is provided on the base and encloses the imaging unit and the light source module.

2. The visual-tactile sensor according to claim 1, characterized in that, The base includes a base and a sensor housing. The sensor housing has a mounting hole in the middle. The imaging unit is mounted on the base and located in the mounting hole.

3. The visual-tactile sensor according to claim 2, characterized in that, The sensor housing has a groove formed around the mounting hole, and the light source module is installed in the groove, with its height lower than the lens surface of the imaging unit.

4. The visual-tactile sensor according to claim 1, characterized in that, The light source module includes a ring-shaped circuit board, which surrounds the imaging unit, and K independently controllable single-point light sources are evenly arranged around it.

5. The visual-tactile sensor according to claim 1, characterized in that, The light source module includes 3N independently controlled single-point light sources, where N≥1.

6. The visual-tactile sensor according to claim 1, characterized in that, Each of the aforementioned point light sources can emit white, red, blue, and green light.

7. The visual-tactile sensor according to claim 1, characterized in that, The elastic cover includes a transparent elastomer and a light-shielding layer, the light-shielding layer covering the outer surface of the transparent elastomer.

8. The visual-tactile sensor according to claim 2, characterized in that, The transparent elastomer is made of transparent silicone material.

9. A geometric calibration method based on a visual-tactile sensor, characterized in that, include: Press the elastic cover to collect multiple sets of light source images under different pressing conditions. The acquisition process for each set of light source images is to press the elastic cover. Under each pressing, one dark field image, K single light source images under a single white light source, and one tri-color illumination image are acquired. Among them, the single light source images are the images acquired when different single light sources are lit in sequence, and K is the number of single light sources. A three-dimensional physical model of near-light photometric intensity decays with distance and angle is established, and the surface normal vector under each press is calculated based on the dark field image in multiple light source images and K single light source images. The surface normal vectors calculated from each set of light source images and the corresponding tri-color illumination images are used as a set of training data to establish a training dataset; A lightweight normal prediction neural network model is constructed, and the three-color illumination images of each training set in the training dataset are used as inputs and the surface normal vectors are used as outputs to train and optimize the lightweight normal prediction neural network model to obtain the optimized lightweight normal prediction neural network model.

10. The geometric calibration method based on a visual-tactile sensor according to claim 9, characterized in that, The steps of establishing a near-light photometric three-dimensional physical model of light intensity attenuation with distance and angle, and calculating the surface normal vector under each press based on the dark field image from multiple sets of light source images and K single-light source images, include: Construct a pixel-level lighting-normal relationship model: Wherein, the observed brightness I of the i-th LED at color channel c relative to pixel p i,o The physical imaging process of (p). Where Ψ i,o This indicates the calibrated light intensity of the light source in this channel. Let x(p) be the surface albedo at pixel p; and x(p) and Let n(p) represent the three-dimensional positions of a surface point and the light source, respectively, and let n(p) be the normal vector of that point. This represents the main emission direction of the light source. (First term) Used to model the directional attenuation of LEDs, where μ i This is a directional index. The second item... Describe Lambertian reflections and use {} to clip backlighting to handle shadows and occlusions; denominator This expresses the distance attenuation of a near-light point source; Based on logarithmic depth variable Construct the energy function using albedo: The physical depth is obtained by solving the dark field image and K single light source images from multiple light source images, based on the pixel-level illumination-normal relationship model, energy function and number of depth variables; Take the depth derivative of the physical depth: The surface normal vector is obtained by calculation; in, This represents the intensity of the observed image. For each pixel p, this is the actual image intensity from the sensor, based on the color channel c and the light source i. Represents an optimized depth map and albedo Predicted image intensity. This represents the optimized depth value. This represents the prior depth used for the initial depth estimate in depth regularization, which is usually derived from calibration data or preliminary estimates.

11. A computer device, characterized in that, The device includes a processor coupled to a memory, the memory storing a lightweight normal prediction neural network model optimized based on a geometric calibration method for a visual-touch sensor according to any one of claims 9 to 10, wherein when the program instructions stored in the memory are executed by the processor, the normal vector of the contact surface is predicted.

Citation Information

Cited By

  • Thin large-area tactile sensor device and method with in-plane light coupling and edge detection

    CN122329530A