Haptic perception system and method, electronic device, storage medium, program product

By combining an active light field with projection time modulation characteristics and an array of acquisition units, the shortcomings of robot tactile perception technology in terms of high frame rate, low power consumption and anti-interference ability are solved, and real-time dynamic tactile perception and high-precision 3D reconstruction are realized.

CN121614037BActive Publication Date: 2026-05-12SHENZHEN RUISHIZHIXIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN RUISHIZHIXIN TECH CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing robot tactile perception technologies struggle to achieve high frame rates, low power consumption, and strong anti-interference capabilities, and are highly sensitive to dynamic contact changes, especially lacking adaptability and reliability in complex interactive tasks.

Method used

By employing a combination of contact components, projection components, and acquisition components, an active light field with time-modulated characteristics is projected and combined with arrayed acquisition units to generate a time-stamped acquisition data sequence, reconstruct the three-dimensional topography, and invert the pressure distribution.

Benefits of technology

It achieves near real-time dynamic tactile perception, reduces system power consumption, and improves the accuracy and robustness of 3D reconstruction, making it suitable for deployment on embedded or mobile platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121614037B_ABST
    Figure CN121614037B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a tactile perception system and method, an electronic device, a storage medium, and a program product. The tactile perception system comprises a contact component, a projection component, a collection component, and a processing component, which are configured to be capable of contacting an external object and producing deformation; the projection component is configured to project an active light field with time modulation characteristics to a first surface of the contact component, the active light field forms a plurality of pattern regions on the first surface, the light intensity of adjacent pattern regions varies based on different time sequence encoding, and the first surface faces away from a second surface of the contact component which contacts the external object; the collection component comprises a plurality of collection units arranged in an array, and is configured to collect the plurality of pattern regions and output a time-labeled collection data sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the fields of robot tactile perception and three-dimensional vision measurement, and more particularly to tactile perception systems, tactile perception methods, electronic devices, non-volatile computer-readable storage media, and computer program products. Background Technology

[0002] With the rapid development of emerging applications such as robot operation, intelligent manufacturing, medical rehabilitation, and virtual reality or augmented reality, the demand for refined and dynamic tactile perception between end effectors and the environment is becoming increasingly urgent.

[0003] Taking robot operation as an example, the robot's fingers not only need to detect whether there is contact with external objects, but also need to accurately perceive the contact position, contact area, pressure distribution information of external objects acting on the contact position, total normal force, shear force, etc., and also need to accurately perceive multi-dimensional information such as the sliding trend and friction changes of external objects relative to the contact position.

[0004] Therefore, the industry urgently needs a new tactile sensing solution that can balance high frame rate, low power consumption, strong anti-interference capability, and high sensitivity to dynamic contact changes, in order to improve the adaptability and reliability of intelligent devices such as robots in complex interactive tasks. Summary of the Invention

[0005] Embodiments of this disclosure provide tactile sensing systems, tactile sensing methods, electronic devices, non-volatile computer-readable storage media, and computer program products that can at least partially solve the problems described above or other problems in the art.

[0006] In a first aspect, embodiments of this disclosure propose a tactile sensing system, comprising: a contact component configured to contact an external object and generate deformation; a projection component configured to project an active light field with time-modulated characteristics onto a first surface of the contact component, the active light field forming multiple patterned regions on the first surface, the light intensity of adjacent patterned regions varying based on different time-series codes, the first surface facing away from a second surface of the contact component that contacts the external object; a acquisition component including multiple acquisition units arranged in an array and configured to acquire multiple patterned regions and output a time-stamped acquisition data sequence; and a processing component configured to perform at least the following: obtaining an observation code sequence for each acquisition unit based on the received time-stamped acquisition data sequence; identifying the multiple observation code sequences to determine the time-series code to which the observation code sequence belongs; reconstructing a three-dimensional topography of one side of the contact component on the first surface based on the time-series code corresponding to each acquisition unit; inverting the pressure distribution of the external object acting on the contact component based on the reconstructed three-dimensional topography; and outputting tactile data based on the pressure distribution.

[0007] Secondly, this disclosure proposes a tactile sensing method, which includes: projecting an active light field with time-modulated characteristics onto a first surface of a contact component by a projection component, the active light field forming multiple patterned regions on the first surface, the light intensity of adjacent patterned regions varying based on different time-series codes, the first surface facing away from a second surface of the contact component that contacts an external object; acquiring multiple patterned regions by an acquisition component and outputting a time-stamped acquisition data sequence, wherein the acquisition component includes multiple acquisition units arranged in an array; and performing at least some of the following based on the received time-stamped acquisition data sequence: obtaining an observation code sequence for each acquisition unit based on the received time-stamped acquisition data sequence; identifying the multiple observation code sequences to determine the time-series code to which the observation code sequence belongs; reconstructing a three-dimensional topography of one side of the contact component on the first surface based on the time-series code corresponding to each acquisition unit; inverting the pressure distribution of the external object acting on the contact component based on the reconstructed three-dimensional topography; and outputting tactile data based on the pressure distribution.

[0008] Thirdly, an electronic device is provided, comprising a processor that can be used to implement the tactile sensing method in the second aspect and any possible implementation thereof.

[0009] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed by a processor, implement the tactile sensing method of the second aspect and any possible implementation thereof.

[0010] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements at least a portion of the steps of the tactile sensing method in the second aspect and any possible implementation thereof.

[0011] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

[0012] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0013] Some exemplary embodiments of this disclosure will now be described with reference to the accompanying drawings. These drawings are for illustrative purposes only and are not intended to limit the claimed technical solutions, wherein:

[0014] Figure 1This is a schematic diagram illustrating the structure and working principle of a tactile sensing system provided according to an exemplary embodiment of this disclosure;

[0015] Figure 2 This is a schematic diagram of the cavity structure provided according to an exemplary embodiment of the present disclosure;

[0016] Figure 3 and Figure 4 These are schematic diagrams illustrating the working principle of the processing components provided according to exemplary embodiments of this disclosure;

[0017] Figure 5 This is a timing diagram of time-series encoded projected structured light and time-stamped acquisition data provided according to an exemplary embodiment of this disclosure;

[0018] Figure 6 This is a timing diagram of time-series encoded projected structured light and time-stamped acquisition data provided according to an exemplary embodiment of this disclosure;

[0019] Figure 7 This is a schematic diagram illustrating the principle of spot aggregation and center positioning according to an exemplary embodiment of this disclosure;

[0020] Figure 8 This is a schematic diagram illustrating the principle of three-dimensional reconstruction based on spatial intersection according to an exemplary embodiment of this disclosure;

[0021] Figure 9 This is a flowchart illustrating a tactile sensing method provided according to an exemplary embodiment of the present disclosure;

[0022] Figure 10 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0023] To better understand this disclosure, various aspects of this disclosure will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of exemplary embodiments of this disclosure and are not intended to limit the scope of this disclosure in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0024] It should be noted that in this specification, the terms "first," "second," etc., are used only to distinguish one feature from another and do not imply any limitation on the features. Therefore, without departing from the teachings of this disclosure, the first surface discussed below may also be referred to as the second surface.

[0025] It should also be understood that expressions such as “comprising,” “including,” “having,” “containing,” and / or “comprising” are open-ended rather than closed expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Additionally, the use of “exemplary” is intended to refer to an example or illustration.

[0026] Unless otherwise specified, all terms used in this disclosure (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms (e.g., those defined in common dictionaries) shall be understood to have the meaning consistent with their meaning in the context of the relevant art, and shall not be interpreted in an idealized or overly formal sense, unless expressly so specified in this disclosure.

[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. Furthermore, unless explicitly limited or contradicted by the context, the specific steps included in the methods described in this disclosure are not limited to the order in which they are described, but can be performed in any order or in parallel. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0028] Figure 1 This is a schematic diagram illustrating the structure and working principle of a tactile sensing system 1000 provided according to an exemplary embodiment of the present disclosure.

[0029] like Figure 1As shown, embodiments of this disclosure provide a tactile sensing system 1000, which includes: a contact component 100, a projection component 200, a acquisition component 300, and a processing component 400. The contact component 100 is configured to contact an external object and deform. The projection component 200 is configured to project an active light field with time-modulated characteristics onto a first surface of the contact component 100. The active light field forms multiple patterned regions on the first surface, and the light intensity of adjacent patterned regions changes based on different time-series encodings. The first surface faces away from the second surface of the contact component 100 that contacts the external object. The acquisition component 300 includes... The system includes multiple acquisition units arranged in an array and configured to acquire multiple pattern regions and output time-stamped acquisition data sequences; and a processing component 400 is configured to perform at least the following: obtaining an observation coding sequence for each acquisition unit based on the received time-stamped acquisition data sequences; identifying multiple observation coding sequences to determine the time series coding to which the observation coding sequence belongs; reconstructing the three-dimensional topography of one side of the contact component 100 on the first surface based on the time series coding corresponding to each acquisition unit; inverting the pressure distribution of an external object acting on the contact component 100 based on the reconstructed three-dimensional topography; and outputting tactile data based on the pressure distribution.

[0030] The tactile sensing system 1000 proposed in this disclosure is designed based on a projection component 200 and a acquisition component 300. The acquisition component 300 includes multiple acquisition units arranged in an array and generates a time-stamped acquisition data sequence based solely on changes in light intensity, thus making it suitable for tactile sensing such as high-speed sliding or vibration. Furthermore, because the acquisition component 300 generates the time-stamped acquisition data sequence based solely on changes in light intensity, it detects less redundant data, significantly reducing the operating power consumption of the tactile sensing system 1000.

[0031] In addition, the contact component 100 includes a first surface and a second surface arranged opposite to each other, wherein an external object can contact the second surface, and the projection component 200 can project an active light field onto the first surface, wherein the active light field is projected onto the first surface to form multiple patterned regions, and the light intensity of adjacent patterned regions changes based on different time series encodings. The acquisition component 300 can acquire the light intensity changes of the reflected light from the active light field on the first surface and output a time-stamped acquisition data sequence.

[0032] When the contact component 100 is not deformed, the reflected light is affected by the intensity change of the active light field, resulting in a corresponding intensity change. The acquisition component 300 can output a time-stamped acquisition data sequence based on the intensity change of the reflected light. When the contact component 100 is deformed, the reflected light is affected not only by the intensity change of the active light field but also by the deformation of the first surface. In this case, the acquisition component 300 can still output a time-stamped acquisition data sequence based on the intensity change of the reflected light. The output time-stamped acquisition data sequence includes the deformation information of the contact component 100. Therefore, the tactile sensing system 1000 can reconstruct the three-dimensional shape of one side of the first surface of the contact component 100 based on the time-stamped acquisition data sequence. Based on this, the tactile sensing system 1000 can also obtain the pressure distribution of an external object acting on the contact component 100 or tactile data of an external object acting on the contact component 100.

[0033] The time-stamped acquisition data sequence may include a data stream acquired by an event camera or other photoelectric detection array that generates synchronous or asynchronous acquisition data based on changes in light intensity. The acquired data in the time-stamped acquisition data sequence includes at least: the time information, location information, and light intensity change information characterizing the direction or magnitude of the change in the intensity of the reflected light. When the acquisition unit receives the light intensity change caused by the reflected light, and the change magnitude exceeds a preset light intensity threshold, a data acquisition point can be generated.

[0034] For example, the time-stamped data sequence may include any of the following different data sequences: a light intensity change event sequence output by an event camera (hereinafter referred to as an event sequence); a light intensity pulse signal sequence output by a pulse camera; a photon-triggered timestamp sequence output by a photodetector array or SPAD (Single-Photon Avalanche Diode) array; a sequence of transition moments obtained based on light intensity acquisition by a high-speed image sensor or an in-pixel (ADC-in-pixel) converter array; or other non-frame-based data sequences that can record the timing or phase information of changes in reflected light signals.

[0035] The time-stamped data sequence acquired by the event camera can be described as an event sequence. An event sequence may include multiple events, which are asynchronous events independent of each other in terms of time, spatial location, and event polarity. Specifically, the event camera may include multiple event pixels (hereinafter referred to as pixels) with different pixel positions, and these pixels can form a pixel array. Reflected light illuminating the pixel array can form a light spot area. Based on the reflected light acquired by one or more pixels in the light spot area, an event sequence characterizing the spatiotemporal information of the reflected light can be generated. The events in the event sequence may include the time information, location information, and polarity information of changes in the intensity of the reflected light. An event can be generated when a pixel receives a brightness change caused by reflected light, and the change exceeds a preset threshold.

[0036] Event cameras can operate in both asynchronous and synchronous modes. In asynchronous mode, each pixel of the event camera operates independently, triggering an event instantly when the light intensity change exceeds a set threshold. This event can contain an independent timestamp. When the event camera performs reception sampling in synchronous mode, all pixels in the pixel array can be sampled simultaneously, calculating the light intensity difference between the current and previous rounds to generate a sparse event matrix. This means that each frame shares a unified timestamp. Therefore, in synchronous mode, the event sequence generated by the event camera can be understood as a sequence of sampled data obtained at a fixed sampling frequency.

[0037] Alternatively, an event sequence may include multiple events, such as multiple events occurring in chronological order. Events may include timestamps of changes in the intensity of reflected light, location information, and polarity information. For example, an event sequence may be represented as a data set { t_1, t_2, ..., t_n},in t_1, t_2 and t_n These can be multiple events that occur in chronological order. Alternatively, multiple events in an event sequence can be arranged in ascending order of their timestamps.

[0038] Taking robot operation as an example, the robot's fingers, as end effectors of intelligent devices, not only need to detect whether there is contact with external objects, but also need to accurately perceive the contact position, contact area, pressure distribution information, total normal force, shear force, etc., of the external object acting on the contact position. They also need to accurately perceive multi-dimensional information such as the sliding trend and friction changes of the external object relative to the contact position. The skin at the tip of the robot's finger can be understood as a specific instance of a contact component 100, which can contact external objects and deform. The first surface of the contact component 100 can be understood as the inner surface of the contact component 100 facing the projection component 200 and the acquisition component 300. This surface is the direct projection target of the active light field and the main interface for acquiring deformation optical information. In contrast, the second surface of the contact component 100 can be understood as the contact surface between the contact component 100 and external objects in the external environment.

[0039] One side of the contact component 100 is exposed to the external environment for interaction with an external object. The external object can be a solid object (e.g., a workpiece grasped by a robot, a human finger, etc.), and when the external object contacts the contact component 100 and applies pressure, the contact component 100 can convert the pressure into its own deformation. It is understood that causing deformation does not necessarily depend on direct physical contact. Non-contact physical actions, such as air pressure, liquid pressure (e.g., in underwater applications), or forces applied by a changing electromagnetic field, can also cause measurable deformation of the contact component 100.

[0040] Existing tactile sensing technologies include array sensors based on resistance or capacitance changes, piezoelectric material sensors, fiber optic sensors, and three-dimensional vision measurement solutions based on frame cameras and active light fields (e.g., structured light).

[0041] Taking a 3D vision measurement scheme as an example, due to its reliance on frame-based cameras for image acquisition, the inherent exposure and readout delays result in high motion latency and a strict frame rate limit. In high-speed sliding or high-frequency vibration scenarios, the system is prone to failure due to image blurring, making it difficult to achieve true real-time dynamic tactile perception. Furthermore, this 3D vision measurement scheme utilizes 3D reconstruction algorithms (e.g., grayscale-related matching), which are computationally complex, time-consuming, and have large matching windows, leading to low 3D data output rates and high system power consumption, making it difficult to deploy on embedded or mobile platforms sensitive to computing power and power consumption. Additionally, for tactile surfaces covered with soft materials, relying solely on single deformation or height information for mechanical inversion makes its accuracy and robustness susceptible to changes in the material, geometry, and operating conditions of external objects, limiting its reliable application in real-world, variable environments.

[0042] To at least address the aforementioned technical problems, the tactile sensing system 1000 proposed in this disclosure projects a time-series encoded active light field onto the first surface of the contact component 100 via a projection component 200, and detects it using a collection component 300. The collection component 300 includes multiple collection units arranged in an array, and generates a time-stamped collection data sequence based solely on changes in light intensity. Since the collection units respond only to changes in light intensity and output collection data with precise timestamps, the system has no exposure delay or fixed frame rate limitations, thus enabling near real-time capture of microsecond-level deformation signals caused by high-speed sliding or vibration, thereby achieving real-time dynamic tactile sensing of dynamic contact processes.

[0043] The time-stamped data sequence only includes information about changes in light intensity, resulting in low data redundancy. This significantly reduces the amount of data the system needs to transmit and process, and also lowers system power consumption. Furthermore, the timestamp-based temporal decoding method can replace traditional, complex grayscale correlation matching algorithms, thereby greatly reducing the computational complexity and power required for 3D reconstruction, making this system highly suitable for deployment on embedded or mobile platforms.

[0044] Furthermore, regardless of whether the contact component 100 deforms, the acquisition component 300 can output a time-stamped acquisition data sequence based on the intensity change of the reflected light. Specifically, when the contact component 100 is not deformed, the reflected light is affected by the intensity change caused by the time-series encoding of the active light field, resulting in a corresponding intensity change. The acquisition component 300 can output a time-stamped acquisition data sequence based on this intensity change. When the contact component 100 deforms, the reflected light is affected not only by the intensity change of the active light field but also by the deformation of the first surface. In this case, the acquisition component 300 can still output a time-stamped acquisition data sequence based on the intensity change of the reflected light, and the output time-stamped acquisition data sequence includes the deformation information of the contact component 100. This consistently effective data acquisition mechanism provides a reliable foundation for subsequent stable and high-precision deformation field reconstruction and pressure distribution inversion.

[0045] Existing 3D vision measurement schemes are based on static structured light methods, whose 3D reconstruction relies on the correlation matching of texture or grayscale spatial features in an image. When the soft material surface of the contact component deforms and its reflectivity changes due to contact with different objects, these spatial features change accordingly, leading to matching errors or failures. In the tactile sensing system 1000 proposed in the embodiments of this disclosure, the projection component 200 projects an active light field onto the first surface of the contact component 100, wherein the active light field forms multiple patterned regions on the first surface, and the light intensity of adjacent patterned regions changes based on different time-series encoded changes. This transfers the features on which the aforementioned correlation matching depends from the unstable spatial domain to the stable temporal domain, and the matching process is based on a preset temporal light intensity change pattern of the patterned regions. As long as the light intensity change of the patterned region can be detected, it can be accurately identified, and its matching robustness is basically unaffected by surface reflectivity, contact material, or deformation geometry.

[0046] Furthermore, 3D reconstruction relies on matching and associating the pixel coordinates on the imaging plane of the acquisition component 300 with the coordinates of a preset coded pattern in the projection component 200. This robust matching based on time-series coding ensures the uniqueness and accuracy of this association, thereby guaranteeing high precision and stability of the 3D topography reconstruction results.

[0047] The active light field may include structured light. The structure and working principle of the tactile sensing system 1000 will be described below using structured light as an example. However, those skilled in the art will understand that the embodiments described herein are merely exemplary and intended to clearly demonstrate the technical principles and solutions of the present invention, and are not intended to limit the scope of protection of the present invention in any way.

[0048] Figure 2 This is a schematic diagram of the cavity 1001 provided according to an exemplary embodiment of the present disclosure.

[0049] like Figure 1 and Figure 2 As shown, in some embodiments of this disclosure, the projection component 200 and the acquisition component 300 may be at least partially placed within the cavity 1001. Furthermore, considering that open optical path designs are susceptible to ambient light interference, affecting the stability and reliability of measurements, the cavity 1001 may be a closed cavity, and the projection component 200 and the acquisition component 300 may be at least partially placed within this closed cavity 1001, thereby reducing the interference of ambient light on structured light projection and time-stamped data sequence acquisition.

[0050] Optionally, the contact component 100 may include an outer layer 110 and an inner layer 120, the outer layer 110 being usable for receiving contact from an external object and including a second surface; the inner layer 120 may be stacked on top of the outer layer 110 and includes a first surface. Alternatively, the inner layer 120 may include at least one of a diffuse reflection layer, a scattering layer with a microstructure, a deformable speckle layer, and a texture with anisotropic optical characteristics.

[0051] In addition, the contact assembly 100 may also include an intermediate layer 130, which may be located between the outer layer 110 and the inner layer 120, wherein the hardness of the portion of the intermediate layer 130 near the outer layer 110 is less than the hardness of the portion of the intermediate layer 130 near the inner layer 120. The hardness of the intermediate layer 130 may increase in a gradient.

[0052] Specifically, the side of the contact component 100 that contacts an external object can be defined as the outer side, and the opposite side can be defined as the inner side. The contact component 100 can serve as an end effector of a smart device and can be implemented as a robot finger, a tactile pad surface (e.g., a robotic arm tactile pad), a gripper liner, a tactile structure within a flexible wearable device, and so on.

[0053] The outer layer 110 can be made of an elastomer material, such as silicone / TPU (Thermoplastic Polyurethane) / PDMS (Polydimethylsiloxane). Exemplarily, the diffuse reflection layer can be at least one of a white diffuse reflection coating and a microtextured film; the microstructured scattering layer can be at least one of microbumps, microprisms, and random textures; a deformable speckle layer can be embedded within the intermediate layer 130 to ensure the stability of the structured light features. Furthermore, the inner layer 120 can also introduce a texture with anisotropic optical characteristics, which has directional optical features such that when shear displacement occurs on its surface, the shape or orientation of the reflected light spot will undergo a specific change related to the displacement direction. This design facilitates the detection and quantification of in-plane displacement caused by shear force, thereby directly serving the estimation of shear force and slip information in tactile data.

[0054] The above measures can optimize optical measurement conditions, ensure that structured light can be effectively detected and analyzed, thereby improving the accuracy and reliability of three-dimensional reconstruction or surface topography measurement, and enhancing diffuse reflection while reducing specular reflection.

[0055] Specular reflection or highlights can cause overexposure or saturation in some areas, resulting in the loss of structured light information. The direction of light reflection depends on the surface normal, which may prevent the reception of light signals at certain angles. A diffuse reflection layer allows incident light to be scattered uniformly in all directions, avoiding interference from specular reflection. Furthermore, the first surface can be white; white surfaces have high reflectivity, reflecting more structured light back to the acquisition component 300, enhancing the brightness and contrast of stripes in the image. The diffuse reflection characteristic means that the light intensity distribution mainly depends on the incident light intensity and surface geometry, facilitating subsequent processing.

[0056] Diffuse reflective layers, microstructured scattering layers, and deformable speckle layers can possess the following optical properties, including but not limited to: high reflectivity and spectral flatness; high reflectivity to the wavelength of the structured light used, and consistent reflection across different wavelengths to avoid introducing color deviations; near-ideal diffuse reflectors, with reflected light intensity independent of viewing angle; surface roughness much smaller than the period of the structured light fringes; and uniform and smooth thickness. Furthermore, textures with anisotropic optical characteristics can possess the following optical properties, including but not limited to: random and uniform texture, spatial frequency higher than the camera resolution, and avoidance of moiré fringes.

[0057] It should be noted that the outer layer 110, inner layer 120 and middle layer 130 mentioned above are functional layers, and at least two of them can be formed by physical layering or integral molding.

[0058] In some embodiments of this disclosure, at least some modules of the projection component 200 and at least some modules of the acquisition component 300 may be disposed within the cavity 1001. Exemplarily, the cavity 1001 may be a closed cavity.

[0059] At least a portion of the projection component 200 and the acquisition component 300 can be optically accessible within a closed cavity to the inner surface (e.g., the first surface) of the contact component 100. Optically accessible means that the projection component 200 can project structured light onto the first surface of the contact component 100, and the acquisition component 300 can capture light intensity changes on the first surface of the contact component 100. Alternatively, the optically accessible projection component 200 can be positioned on one side of the first surface of the contact component 100, with its optical axis pointing towards the first surface at a certain angle, thereby projecting a structured light pattern carrying encoded information onto the first surface; the acquisition component 300 is also positioned on the same side of the first surface, spatially separated from the projection component 200, with its optical axis pointing towards the first surface from a different angle, thereby observing the first surface illuminated by structured light and capturing changes in the intensity of its reflected light. The contact component 100, as the observed target, has its first surface located at the intersection of the projection and acquisition optical paths.

[0060] Optionally, to reduce the space of the cavity 1001 and achieve miniaturization of the tactile sensing system 1000, the projection component 200 and the acquisition component 300 may be arranged side by side in the same direction. In this embodiment, the tactile sensing system 1000 may also include at least one optical mirror (not shown) disposed in the cavity 1001, for refraction and reflection of the structured light emitted by the projection component 200 to the first surface, and / or refraction and reflection of the reflected light from the first surface to the acquisition component 300, thereby also achieving a triangular optical path layout.

[0061] At least some modules of the projection component 200 disposed within the cavity 1001 may include modules related to directly projecting structured light; at least some modules of the acquisition component 300 disposed within the cavity 1001 may include modules related to directly capturing changes in light intensity. Optionally, the processing component 400 may also reconstruct the three-dimensional topography by processing the time-stamped acquisition data sequence using a spatial intersection algorithm based on fixed position parameters between the projection component 200 and the acquisition component 300. In other words, the modules in the acquisition component 300 that directly capture changes in light intensity and the modules in the projection component 200 that directly project structured light have a fixed geometric relationship, which facilitates the calibration of the intrinsic and extrinsic parameters of the acquisition component 300 for spatial intersection point cloud recovery.

[0062] Spatial intersection algorithms can utilize two-dimensional information acquired by multiple observation devices with known spatial positions and attitudes to determine the unique coordinates of a target point in three-dimensional space through the intersection of geometric models. Therefore, in some embodiments of this disclosure, the spatial intersection algorithm can utilize the fixed spatial geometric relationship (i.e., "position parameters") between the projection component 200 and the acquisition component 300 to inversely calculate three-dimensional spatial points from two-dimensional observation data. Spatial intersection algorithms may include triangulation, least squares solving, epipolar geometric constraint-based solving, bundle adjustment, and other equivalent stereo reconstruction methods.

[0063] In this invention, regardless of the specific algorithm used, as long as it utilizes a fixed projection-acquisition spatial geometric relationship to intersect the two-dimensional observation position of the coded pattern matched in the time-stamped data sequence with the observation line of sight in space through the projection ray, thereby calculating the three-dimensional shape data, it should be regarded as a specific implementation of the "spatial intersection algorithm" and fall within the protection scope of this invention.

[0064] Furthermore, the remaining modules in the projection component 200 or the acquisition component 300 used for driving, control and signal processing can be flexibly selected to be all set inside the cavity 1001, partly set inside the cavity 1001 or all set outside the cavity 1001 according to the design requirements of the tactile sensing system 1000. This disclosure does not limit this.

[0065] The processing component 400 is communicatively connected to the acquisition component 300 to receive the time-stamped acquisition data sequence output by the acquisition component 300 and process the time-stamped acquisition data sequence into tactile event data. The processing component 400 can be disposed inside or outside the cavity 1001. Alternatively, the processing component 400 may also include multiple processing modules (not shown), some of which can be disposed inside the cavity 1001 and others outside the cavity 1001.

[0066] Optionally, the inner wall of cavity 1001 may be coated / attached with a light-absorbing material to absorb stray reflected light and prevent spurious events that interfere with measurements. Similarly, cavity 1001 may also be provided with at least one of a light-absorbing wall 1004 and a light-shielding structure 1005, wherein the light-absorbing wall 1004 may be made of a light-absorbing material, and the light-absorbing wall 1004 and the light-shielding structure 1005 may be arranged in the optical path where stray light reflection may occur to absorb stray light and suppress spurious events.

[0067] Refer again Figure 1 and Figure 2 In some embodiments of this disclosure, the projection component 200 is used to project time-series encoded structured light onto a first surface of the contact component 100. The structured light is projected onto the first surface to form multiple patterned regions, and the light intensity of adjacent patterned regions varies based on different time-series encodings.

[0068] Optionally, the projection component 200 can vary the light intensity of two patterned regions that are spatially greater than or equal to a predetermined distance according to the same time-series encoding. Optionally, the projection component 200 can also set the Hamming distance between the time-series encodings corresponding to adjacent patterned regions to be greater than or equal to a predetermined threshold.

[0069] Specifically, the projection assembly 200 may include an emitter, optical elements, a drive controller, and a timing control module. The emitter in the projection assembly 200 may be one or a combination of two or more of the following: a VCSEL (Vertical-Cavity Surface-Emitting Laser) array, a laser diode array, a microprojector, or a multi-channel LED (Light-Emitting Diode).

[0070] Alternatively, the structured light pattern can be one or more coded patterns selected from speckle patterns, stripe patterns, Gray code patterns, and binarized dot matrix patterns.

[0071] For example, VCSEL arrays feature low power consumption, high beam quality, ease of two-dimensional integration into dense dot lattices, and high modulation bandwidth, making them suitable for generating high-density speckle patterns. Laser diode arrays can provide higher single-point optical power. Microprojectors based on digital micromirror devices (DLP) or liquid crystal on silicon (LCOS) can generate complex patterns (e.g., Gray code fringes) through dynamic programming, exhibiting good programmability. Multi-channel LEDs offer advantages in terms of cost and system integration. Furthermore, hybrid projection schemes can be employed, such as combining a VCSEL array to generate background speckle and using a DLP projector to project specific auxiliary coded patterns to achieve multi-scale, robust depth sensing. Therefore, in practical designs, a suitable emitter can be selected by comprehensively considering factors such as projection distance, spatial resolution, system cost, and power consumption.

[0072] The optical elements in the projection assembly 200 may include DOEs (Diffractive Optical Element), diffusers, collimating lenses, etc., to form the projected structured light into a multi-spot array (e.g., a dot matrix), speckle, or stripe pattern. DOEs can be used to generate random speckles or regular dot matrices, possessing micro- or nano-structures internally or on their surfaces, which can diffract light to form a certain number of tiny spots with unique distribution characteristics. Diffusers can be used to homogenize the beam and widen the illumination angle, ensuring a uniform illumination field. Collimating lenses can be used to convert divergent light source output into parallel light, ensuring that the pattern formed by the structured light remains undistorted within the effective distance.

[0073] Alternatively, the aforementioned transmitter and optical elements, as related modules for directly projecting structured light in the projection assembly 200, can be disposed in the cavity 1001.

[0074] The drive controller in the projection component 200 can control the output timing of multiple independent "logical channels," where the coded light of each logical channel is projected onto the first surface of the contact component 100, illuminating and forming a corresponding "pattern area." In other words, the drive controller can be used to generate multi-channel time-series encoding to independently control the light intensity output timing of multiple logical channels. Multi-channel time-series encoding can be understood as multiple logical channels being assigned a distinct M-bit binary code, and the light intensity state (on or off) of each logical channel being driven in chronological order to sequentially present the values ​​of the corresponding bits in its encoding sequence, where the total number of bits M can be a natural number greater than 1.

[0075] Specifically, multi-channel time series encoding can be understood as encoding a time series in M ​​bits. Within this context, M codes are assigned to logical channel 1, for example, [0,1,0,1,1,0,0,…,1]. Therefore, in consecutive time steps… to Within this sequence, the light intensity of logic channel 1 will change sequentially according to the light intensity variation time sequence of "off, on, off, on, on, on, off, off, ..., on". In addition, within the above time sequence, M codes are assigned to logic channel 2, for example, [1,0,1,0,0,1,1,...,0], where the light intensity variation time sequence of logic channel 2 may be different from that of logic channel 1.

[0076] When the independently encoded structured light emitted from the projection component 200 illuminates the first surface of the contact component 100, the light from different logic channels will form a corresponding pattern area on the first surface, where the pattern can be a light spot or stripes, etc. Since the light intensity of adjacent pattern areas is based on different time-series encoded changes, the feature matching basis relied upon by existing 3D reconstruction can be shifted from "spatial grayscale or texture features" which are easily affected by surface reflection characteristics to stable "time-series light intensity change features". The feature matching process is based on the light intensity of the pattern area being encoded by a preset time series. As long as the light intensity change of the pattern area can be detected, it can be accurately identified. Its matching robustness is basically unaffected by surface reflection characteristics, contact material, or deformation geometry.

[0077] The multi-channel time-series encoding of the drive controller is designed to drive the projected structured light spots / speckles / stripes to change brightness / dullness according to the encoded pulses. For ease of description, the spot / speckle / stripe will be referred to as a light-emitting unit. One or more light-emitting units can be programmed into a logic channel, and light-emitting units within the same logic channel can present the same light-emitting rhythm according to the same time-series encoding.

[0078] Alternatively, the drive controller can control the encoding of each logic channel independently or in groups. Independent control can be understood as driving and controlling each logic channel independently, which enables fully parallel signal output. Under the independent control strategy, to enhance the distinguishability of different channels during decoding, the Hamming distance between the time-series codes corresponding to adjacent logic channels can be set to be greater than or equal to a predetermined threshold. In this implementation, the Hamming distance between the time-series codes corresponding to adjacent pattern areas can be set to be greater than or equal to the predetermined threshold. Maintaining a sufficient Hamming distance can effectively prevent the confusion of different logic channels during decoding due to signal interference.

[0079] The Hamming distance is the number of times corresponding bits differ between two coded sequences of equal length. For example, the time series codes "1011" and "1101" differ in two corresponding bits, so their Hamming distance is 2. A predetermined threshold for the Hamming distance can be used to determine whether two time series codes are distinguishable or whether there are valid errors. For example, a predetermined threshold of 2 can be set for the Hamming distance; if the Hamming distance between two logical channel codes is greater than or equal to 2, they are considered significantly distinguishable.

[0080] Group control can be understood as dividing multiple logical channels into multiple groups, with a unified control strategy within each group and a differentiated strategy between groups. For example, assigning the same time-series code to the logical channels corresponding to two spatially distant pattern regions (greater than or equal to a predetermined distance) and grouping them together for unified driving, while assigning different time-series codes to the logical channels corresponding to spatially adjacent or close pattern regions. This spatial distance-based grouping method significantly optimizes the system's coding resources and control complexity, achieving a balance between performance and efficiency.

[0081] Furthermore, to enhance robustness, embodiments of this disclosure may employ one or more of the following coding design methods:

[0082] (i) By optimizing the design of the time series coding, the number of light intensity state (on / off) transitions within a coding cycle is increased. More state transitions allow the acquisition component 300 to capture light intensity change events more densely, thus making the coding features more significant. For example, the original light intensity change sequence "on-on-off-off" is optimized to "on-off-on-off", changing from 1 transition to 3 transitions, which significantly increases the event occurrence rate.

[0083] (ii) Repeatedly project the same complete time series encoding multiple times. This temporal redundancy, through comparison and fusion of multiple periodic data, can suppress random noise interference and improve decoding accuracy. For example, by replicating the 4ms encoding period three times, the total projection time of the structured light is 12ms. Applying mean filtering to the three sets of time series encodings can eliminate encoding errors caused by single-pulse noise.

[0084] (iii) Increase the length of the time series code, which can be understood as the total number of bits M described above. A longer code provides a larger coding space, thus accommodating more unique codes while satisfying the minimum Hamming distance constraint, significantly improving the system's anti-aliasing capability. For example, upgrading a 4-bit binary code to 8 bits expands the coding space from 16 to 256.

[0085] Furthermore, in some embodiments of this disclosure, the encoding of structured light may include periodic codes, Gray codes, binary codes, pseudo-random codes, etc.

[0086] Specifically, periodic encoding can be used as the encoding scheme for time series encoding. In this implementation, periodic encoding can use a fixed period as a unit, allowing the emitting unit to switch between on / off or bright / dark states according to a repetitive temporal pattern. The temporal sequence within the period serves as a unique identifier, and redundancy verification is achieved through multi-period acquisition. For example, the period of a complete time series encoding can be set to 4 frames, and the light intensity change sequence of a certain emitting unit can be "on-off-on-off". This sequence can be periodically repeated in subsequent projections.

[0087] Optionally, Gray code encoding can be used as the encoding scheme for time series encoding. In this implementation, a complete Gray code sequence is preset for the luminous unit, so that when it is driven in consecutive time steps, only one bit of the binary number changes between adjacent light intensity states. For example, the sequence of adjacent light intensity states may be 000→001→011→010… This characteristic of only one bit changing between adjacent states brings a fault tolerance advantage. Even if the data acquisition fails at a certain time step, the complete encoding sequence can still be uniquely derived from the correct states of the time steps before and after it.

[0088] Optionally, binary code encoding can be used as the encoding scheme for time series encoding. In this implementation, the light intensity state (on / off) of the emitting unit within each encoded time slot (or time step) is mapped to a binary bit (e.g., 1 / 0). The sequence composed of the states of multiple consecutive time slots constitutes the unique time series code for that emitting unit. During decoding, bitwise operations can be used to match the observed sequence with the preset time series code. For example, using a 4-bit binary code, the time sequence of a emitting unit "on-on-off-on" corresponds to binary "1101" (or "0010"), and another light point "off-on-on-off" corresponds to "0110" (or "1001").

[0089] Optionally, pseudo-random code encoding can be used as the encoding scheme for time series encoding. In this implementation, pseudo-random code encoding can be based on a pseudo-random sequence generation algorithm to assign a time sequence with no obvious pattern to each emitting unit. High robustness matching is achieved by utilizing the autocorrelation of the sequence, avoiding confusion with the encoding sequences of adjacent logic channels. For example, using an M-bit sequence pseudo-random code, the light intensity change time sequence of a certain emitting unit is "on-off-on-on-off-off-on," which has characteristics similar to white noise. Even if light spots spatially overlap, different emitting units can still be distinguished by calculating the autocorrelation coefficient between the acquired sequence and the preset sequence.

[0090] In some embodiments of this disclosure, the tactile sensing system 1000 further includes a timing control module. This module aligns the timing of the time-stamped acquired data sequence with the timing of the time-series encoded data, providing a basis for accurate decoding by the processing component 400. Optionally, this timing control module can be integrated into the projection component 200, the acquisition component 300, or the processing component 400, or it can exist as a separate hardware unit. This disclosure does not limit the specific hardware structure, implementation circuit, or software organization of the timing control module; it can be a dedicated clock generator, a programmable logic unit, a timer module in a microcontroller, or functional logic composed of a synchronization algorithm running on a processor.

[0091] It should be noted that any functional module or mechanism that can realize or maintain the above-mentioned timing correspondence by generating clock signals, issuing hardware triggers, performing software timestamp calibration, or using any other technical means, and ensure that the encoded state and event acquisition can be reliably associated on the timeline, falls within the conceptual scope of the timing control module described in this disclosure and can be used to implement the technical solution of this invention.

[0092] For example, the following description uses a projection component 200 that includes the timing control module described above. The timing control module in the projection component 200 can be used to establish and maintain the timing alignment relationship between the projection component 200 and the acquisition component 300.

[0093] Specifically, the processing component 400 can reconstruct the light intensity state (on or off) observed by each acquisition unit of the acquisition component 300 within the encoding time window defined by the projection component 200 based on the timestamps in the acquired data sequence, thus forming an observation encoding sequence. If there is a discrepancy between the timing reference of the projection component 200 and the acquisition component 300, the acquired data cannot be accurately mapped to its corresponding encoding time window, resulting in errors in state judgment and decoding.

[0094] Therefore, in order to ensure that the acquired data accurately reflects the coding characteristics of structured light, the timestamps of the acquired data sequence can be time-stamped and the coding time sequence of the projection component 200 can be synchronized. Timing alignment. Optionally, one or a combination of the following methods can be used: on the one hand, it can be achieved through hardware timing alignment, such as using a global clock signal, hardware trigger line or precision time protocol to unify the clock reference of the projection component 200 and the acquisition component 300; on the other hand, it can be achieved through software calibration, by having the acquisition component 300 record the known reference timing signal emitted by the projection component 200, calculate the time deviation and perform software compensation, thereby establishing a time alignment relationship at the data processing level.

[0095] Refer again Figure 1 and Figure 2As described above, the acquisition component 300 can be used to capture the light intensity changes caused by the reflection of structured light through the first surface and output a time-stamped acquisition data sequence.

[0096] It should be noted that "time-stamped acquisition data sequence" is a general term for non-frame-based sampling data generated based on changes in the intensity of reflected light signals, capable of recording the timing of changes, and is not limited to any specific hardware or data format. Each acquisition data in this sequence includes at least: the time information, location information, and intensity change information representing the direction or amplitude of the change in reflected light intensity. When the acquisition unit receives the aforementioned intensity change caused by the reflected light, and the magnitude of the change exceeds a preset intensity threshold, an acquisition data can be generated.

[0097] Taking the acquisition component 300, which includes an event camera, as an example, the event camera may include a standalone EVS (Event-based Vision Sensor) or a hybrid sensor (HVS) that integrates an event sensor and an active pixel sensor (APS). The HVS has the function of outputting event data and frame images. In the embodiments of this disclosure, the HVS can analyze event data to achieve tactile perception, but it is not excluded that it can fuse frame image data for auxiliary analysis and verification.

[0098] For example, an event camera may include an optical system, an event vision sensor, an analog front-end, and a processor. The optical system images reflected light signals onto the photosensitive surface of the sensor and may include a lens group, an aperture, etc. The event vision sensor includes multiple pixels arranged in an array, where each pixel can be understood as a collection unit of the acquisition component 300. Each pixel can operate independently and triggers and outputs event data only when it detects a change in the intensity of the reflected light it receives exceeding a preset intensity threshold. Therefore, the redundancy of the detected data is low. This significantly reduces the amount of data that the system needs to transmit and process, and also reduces the system's operating power consumption.

[0099] Events can be represented in data form. As an option, one can use ( x, y, t, p An event is represented in the form of a quadruple. t The timestamp represents the time of an event, and it can be accurate to the millisecond or microsecond level; x, y () indicates the pixel location where the event occurred; p (Polarity) represents the polarity of an event, characterizing the trend of light intensity change. For example, when p is +1, it indicates an increase in light intensity; when p is 0, it indicates no change in light intensity; and when p is -1, it indicates a decrease in light intensity.

[0100] The output of an event camera can be either a 1-bit or 2-bit matrix representing the states of all pixels. In a 1-bit matrix, the value of any element indicates whether the corresponding pixel has output an event; if an event occurs, the element is 1; otherwise, it is 0. In a 2-bit matrix, the value of any element indicates whether the corresponding pixel outputs a positive or negative event, or no event at all. For example, a positive event is represented by a +1 element, a negative event by a -1 element, and no event by a 0 element. The first bit of any element in a 2-bit matrix can represent the sign.

[0101] Therefore, in at least one embodiment of this disclosure, since the acquisition component 300 generates a time-stamped acquisition data sequence based solely on the intensity change of reflected light, its acquisition unit has a fast response speed, and the system is not limited by the exposure delay and fixed frame rate of traditional cameras. Thus, it can capture microsecond-level deformation signals caused by high-speed slippage or high-frequency vibration in near real-time.

[0102] Furthermore, the time-stamped data sequences contain only information on light intensity changes, resulting in extremely low data redundancy. This significantly reduces the system's data transmission load, processing burden, and overall power consumption. The processing component 400 can perform time-series decoding based on precise timestamps, replacing the computationally complex grayscale correlation matching algorithm. This reduces the computational complexity and power required for 3D reconstruction, making the system highly suitable for deployment on power- and computing-intensive embedded or mobile platforms.

[0103] Furthermore, regardless of whether the contact component 100 deforms, the acquisition component 300 can output a time-stamped acquisition data sequence based on the intensity change of the reflected light. In the absence of contact, the time-stamped acquisition data sequence reflects the encoding changes of the structured light itself; in the presence of contact deformation, the time-stamped acquisition data sequence reflects both the encoding changes of the structured light itself and the deformation information of the first surface. This data acquisition mechanism, effective under any condition, provides a reliable foundation for subsequent stable and high-precision deformation field reconstruction and pressure distribution inversion.

[0104] Figure 3 and Figure 4 These are processing components 400 provided according to exemplary embodiments of this disclosure (such as...). Figure 1 The diagram shows the working principle of the diagram.

[0105] like Figure 1 , Figure 3 and Figure 4As shown, in some embodiments of this disclosure, the processing component 400 can perform EVS-based temporal structured light feature extraction and recognition. Specifically, the processing component 400 obtains the observation coding sequence for each acquisition unit based on the received time-stamped acquisition data sequence; it identifies multiple observation coding sequences to determine the time series coding to which the observation coding sequence belongs. Alternatively, before obtaining the observation coding sequence for each acquisition unit, the processing component 400 can also filter the received time-stamped acquisition data sequence. This filtering may include: counting the acquisition data of each acquisition unit within the coding period of the time series coding; and excluding acquisition units whose acquisition data count is lower than the response rate threshold to form a set of acquisition units for subsequent processing.

[0106] Optionally, identifying multiple observation coding sequences to determine the time series coding sequence to which the observation coding sequence belongs may include determining the coding sequence to which the observation coding sequence belongs through Hamming distance metric, correlation operation, frequency domain demodulation, statistical similarity metric, maximum likelihood, or other equivalent time series feature identification methods. To simplify the above description, this content will be described below as matching the time series coding sequence corresponding to each acquisition unit, or "matching".

[0107] The processing component 400 can be used for real-time processing, timing calibration, and format conversion of time-stamped acquired data sequences. The processing component 400 can be built based on high-performance chips such as FPGAs (Field-Programmable Gate Arrays), MCUs (Microcontroller Units), DSPs (Digital Signal Processors), GPUs (Graphics Processing Units), or NPUs (Neural Processing Units), and possesses parallel computing capabilities. The processing component 400 can also integrate a parameter configuration interface, allowing for dynamic adjustment of sensor thresholds, filtering parameters, etc., from external sources to adapt to different encoding schemes and application scenarios. Therefore, the processing component 400 can perform low-latency processing on acquired data from large amounts of time-stamped acquired data sequences.

[0108] Optionally, the processing component 400 may also perform at least one of the following processing: perform timing calibration on the acquired data in the time-stamped acquisition data sequence, and, in conjunction with the clock reference provided by the timing control module in the projection component 200, correct the timestamp deviation caused by hardware transmission to ensure that the timestamp of the acquired data is aligned with the encoded time sequence of the projection component 200; and filter the received time-stamped acquisition data sequence, etc.

[0109] Optionally, the processing component 400 may perform a series of processes on the received time-stamped acquisition data sequence and output tactile data reflecting the tactile perception of the contact component 100. Exemplarily, the tactile data may include at least one of the following: contact state information between the contact component 100 and an external object; pressure distribution information of the external object acting on the contact component 100; total normal force of the external object acting on the contact component 100; shear force of the external object acting on the contact component 100; and slippage information of the external object relative to the contact component 100.

[0110] The processing component 400 has optional hardware implementation and can employ different integrated architectures. For example, the processing component 400 can be a highly integrated single-chip architecture, implemented as a System on Chip (SoC). This SoC integrates a central processing unit, dedicated hardware accelerators, memory controllers, and other modules to form a complete computing system. Alternatively, the processing component 400 can also be a distributed, collaborative, modular architecture, comprising multiple processing modules (not shown), which can be discrete but collaborative hardware and software modules. The processing performed by the processing component can be executed by dedicated chips, programmable chips, general-purpose processors, etc.

[0111] For example, processing component 400 may be configured to perform a series of processes to extract tactile information from a time-stamped acquisition data sequence. This processing may include: filtering the time-stamped acquisition data sequence, wherein acquisition data for each acquisition unit is counted within the encoding period of the time-series encoding; and excluding acquisition units whose acquisition data count is lower than a response rate threshold to form a set of acquisition units for subsequent processing. By setting a response rate threshold, data corresponding to acquisition units with an average data volume lower than the threshold are excluded from subsequent calculations, thereby obtaining the set of acquisition units. This directly sparsifies the data to be processed, significantly reducing the complexity of subsequent operations.

[0112] Optionally, the above processing may further include: performing time-series decoding on the acquisition unit set, wherein based on the acquisition data of each acquisition unit in the acquisition unit set, the observation coding sequence representing the light intensity state observed by it can be restored, the observation coding sequence can be a light intensity state sequence, and the observation coding sequence is matched with the multi-channel time series coding of the projection component 200 to determine the logical channel to which each acquisition unit belongs.

[0113] Optionally, the above processing may further include: performing spot aggregation, aggregating multiple acquisition units that are determined to match the same logical channel and are spatially adjacent into a single spot region. The center coordinates of each spot region are calculated using center positioning.

[0114] Optionally, the above processing may further include: comparing the center coordinates of the light spot regions corresponding to the same logical channel with the preset projection coordinates of that channel in the projection component 200 through parallax calculation to obtain parallax information. Using the principle of spatial intersection, the coordinates of the center of each light spot region in three-dimensional space are calculated, where the centers of all light spot regions together constitute the three-dimensional shape of one side of the contact component 100 on the first surface. In other words, based on the fixed position parameters between the projection component 200 and the acquisition component 300, the above-mentioned three-dimensional shape can be reconstructed by processing the time-stamped acquisition data sequence through a spatial intersection algorithm.

[0115] Optionally, the above processing may further include: constructing a deformation field based on the three-dimensional morphology of one side of the contact component 100 on the first surface, and quantifying the deformation and deformation gradient of the first surface relative to the non-contact state of the contact component 100.

[0116] Optionally, the above processing may further include: performing pressure inversion based on the deformation field, converting the deformation information into pressure / pressure distribution of the external object acting on the contact component 100 through a physical model or a data-driven model, and outputting tactile data.

[0117] Specifically, taking the acquisition component 300, which includes an event camera, and the time-stamped acquisition data sequence as an event sequence as an example, the event can be... Marked as formula (1).

[0118] (1)

[0119] in, For the first One event; The event sequence number, when encoded as a whole frame. k Refers to the first k frame; The first in the image coordinate system The x-coordinate of the pixel where the event occurred; The first in the image coordinate system The y-coordinate of the pixel where the event occurred; For the first The timestamp of the event; For the first The polarity of an event.

[0120] In the step of filtering time-stamped data sequences, the event response rate can be statistically analyzed to determine the set of acquisition units, where the event response rate can be defined as the ratio of the number of valid events to the number of frames per cycle. Each pixel... During the encoding cycle Event count (or event rate) within. It can be expressed using formula (2).

[0121] .

[0122] in, Represents pixels; Indicates the first The pixel coordinates of each event; For the first The timestamp of the event; Indicates the encoding period duration of the time series encoding; Represents pixels During the encoding cycle Event count within; Indicates encoding period The starting time; Indicates encoding period The end moment.

[0123] Based on this, the event response rate It can be expressed using formula (3).

[0124] (3)

[0125] in, Represents pixels; For pixels Event response rate; Represents pixels During the encoding cycle The event count within the time series encoding period; M is the time series encoding period. The total number of bits within the code, also known as the number of coding steps.

[0126] Figure 5 This is a timing diagram of time-series encoded projected structured light and time-stamped acquisition data provided according to an exemplary embodiment of this disclosure.

[0127] like Figure 1 and Figure 5 As shown, the encoding step size M can be set to 6, that is, 6 bits can be used to represent an encoding. One encoding cycle The projection component 200 can be encoded according to a time sequence [011010] in Within a given time period, light spots are projected in a sequence of "off-on-on-off-on-off".

[0128] The event camera can detect the following events sequentially at the same pixel location where the aforementioned light spot is projected: 0 (event polarity is set to 0 in the first frame), +1, 0, -1, +1, -1. The event count is 4, and the event response rate of this pixel is 4 / 6 = 0.667.

[0129] Optionally, a response rate threshold can be set. When the event response rate Below the response rate threshold Exclude the corresponding pixel when... hour, Pixels Included in the collection unit set .

[0130] In other words, the event sequence and the time series encoding do not directly correspond numerically. Based on the acquired data of the pixel, the observation encoding sequence representing the light intensity state can be reconstructed. This observation encoding sequence can be a light intensity state sequence. By matching the observation encoding sequence with the multi-channel time series encoding of the projection component 200, the logical channel to which the pixel belongs can be determined.

[0131] Figure 6 This is a timing diagram of time-series encoded projected structured light and time-stamped acquisition data provided according to an exemplary embodiment of this disclosure.

[0132] like Figure 1 and Figure 6 As shown, structured light from different logic channels is projected into different regions of the first surface. The encodings of different channels can be different from each other. Furthermore, the Hamming distance between the encodings of different channels can be greater than or equal to a predetermined threshold.

[0133] Furthermore, embodiments of this disclosure employ structured light with brightness abrupt changes, thus ensuring at least one abrupt change occurs within a single encoding cycle. Locations not affected by structured light projection will not trigger events from the event camera, and pixels at these locations are excluded. Alternatively, before obtaining the observation encoding sequence for each acquisition unit, the received time-stamped acquisition data sequence is filtered, wherein the acquisition data for each acquisition unit within the encoding cycle of the time-series encoding is counted; and acquisition units with acquisition data counts below a response rate threshold are excluded to form a set of acquisition units for subsequent processing.

[0134] In some implementations of this disclosure, the response rate threshold used for the above screening can be adaptive based on the "contact state". For example, contact between the contact component 100 and an external object can cause a change in the intensity of the reflected light. Dynamically adjusting the response rate threshold can balance the sensitivity and false detection rate of the tactile sensing system 1000.

[0135] Optionally, the tactile sensing system 1000 may have a first operating mode and a second operating mode, and the material of the contact component 100 may include a light-transmitting material. In the first operating mode, the processing component 400 may acquire a data sequence based on a time stamp and detect changes in the light intensity transmitted through the light-transmitting material to the side where the first surface is located; in response to the change in light intensity satisfying a preset condition, the tactile sensing system 1000 is triggered to switch to the second operating mode, so that the projection component 200 projects structured light onto the first surface.

[0136] The response rate threshold can be defined as At different times Different filtering thresholds are used. For example, the material of the contact component 100 may include a light-transmitting material, thus the contact component 100 is light-transmitting. If an external object approaches the second surface of the contact component 100, it will block some external light from entering the side where the first surface of the contact component 100 is located, resulting in a decrease in light intensity at that location. In response to this change in light intensity satisfying a preset condition, the projection component 200 can project structured light onto the first surface. Therefore, the event camera can normally detect at a relatively low frame rate, and once a change in the contact state is detected, it can immediately switch to high frame rate operation and dynamically redefine the response rate threshold.

[0137] In other implementations of this disclosure, "local ROI (Region of Interest) adaptation" can be incorporated. On the imaging plane of the acquisition component 300, pixel areas predicted to potentially correspond to physical contact can employ a lower response rate threshold to maintain high sensitivity; while other areas on the imaging plane can employ a higher response rate threshold to suppress noise and reduce the operating power consumption of the tactile sensing system 1000. In this embodiment, the response rate threshold can be defined as... It has different response rate thresholds in different regions of the 300-degree imaging plane of the acquisition component. It should be noted that... It can be understood as a data pair that represents the position coordinates in the imaging plane.

[0138] In some embodiments of this disclosure, timing decoding of the acquisition unit set may include: decoding and matching the pixels in the acquisition unit set on a time axis. One flash corresponds to one whole frame, and timing decoding and matching are performed on multiple frames (e.g., 4 or 6 frames) of events on the same pixel.

[0139] logical channel j Time series encoding is represented as The time series encoding is an M-bit sequence of 0s and 1s. With M=6, example time series encodings could be [011010], [110110], [100100], etc., resulting in 64 different encodings in the encoding space. Different encodings can be used for different logical channels, and as described previously, the Hamming distance between the time series encodings corresponding to adjacent pattern regions can be set to be greater than or equal to a predetermined threshold to ensure that the light intensity changes of adjacent pattern regions are significantly distinguishable.

[0140] Pixels The observation coding sequence is represented as The values ​​in the sequence are obtained by determining whether the event is triggered at each encoding step. As shown in Table 1 below, the time sequence encoding of logic channel 1 can be... For [011010], the position of the modulated structured light projected onto the first surface corresponds to the pixel coordinates of the event camera ( a, b At point ), the flickering of the structured light causes the event camera to record the pixel coordinates ( a, b The event sequence at point ) will be encoded with the time series if the record is accurate. Corresponding. The other position on the first surface corresponds to the coordinates of the event camera ( c, d At point ), the flickering of the structured light causes the event camera to record the pixel coordinates ( c, d The event sequence at point () is [001010], which is related to the time series encoding. Not a match.

[0141]

[0142] Table 1

[0143] In one implementation, as shown in formula (4), regarding pixel coordinates ( x, y Is it by logical channel? j Structured light projection patterns, also known as pixels ( x, y Is it related to the logical channel? j Matching can be determined using Hamming distance.

[0144] .

[0145] in, Represents pixels; It can be used to represent pixels The observed encoded sequence is a binary code after decoding, wherein the above Table 1 contains... The event sequence is [0, +1, 0, -1, +1, -1]. After being reset before encoding, the event changes begin with the light spot being off. Therefore, the binary code [011010] can be recovered from the event sequence. For simplicity, in this formula... This can refer to the recovered binary code [011010]. Similarly, The event sequence is [0, 0, +1, -1, +1, -1], which can be deduced from the binary code [001010].

[0146] In addition, in formula (4) Indicates the first Preset time series encoding for each logical channel; This represents the function for calculating Hamming distance; Indicates the first Calculate the Hamming distance for each logical channel, and take the channel number with the smallest distance as the optimal matching result.

[0147] As an alternative, pixel-by-pixel comparison can be used. x, y ) observation coding sequence Time-series encoding with logical channels The Hamming distance between the nodes is used to determine the time series code corresponding to each pixel, or in other words, the logical channel matched to each pixel. For example, there are a total of N logical channels, where the logical channel number is... j =1, 2, ..., N, where N is a natural number greater than 1, calculate the Hamming distance for each number. Take the Hamming distance. Logic channel number with the smallest value j The value of is determined as a value related to the pixel ( x, y The logical channels for matching. Furthermore, the calculation process of the Hamming distance can be represented as... Among them, there are two encoded sequences of length M. and Their Hamming distance is equal to the sum of the absolute values ​​of the differences between their corresponding bit values. The sequence... The value of the m-th position with sequence The value of the m-th position Subtracting the two codes and taking their absolute values, then summing the results for all M bits, determines the two encoded sequences. and Hamming distance.

[0148] The above Time series encoding with logic channel 1 They are completely identical, therefore the Hamming distance between them is 0. This confirms the pixel ( a, b This matches logical channel 1. (The above...) Time series encoding with logic channel 1 Since one bit is different, the Hamming distance between the two is 1.

[0149] In addition, to suppress noise and enhance anti-interference capabilities, a decision threshold can be set:

[0150] (5)

[0151] in, This represents the function for calculating Hamming distance; Indicates the Hamming distance from the decision threshold; The observed coding sequence representing a pixel; Represents the optimal matching logical channel The preset time series encoding.

[0152] In other words, when using the above method to determine the minimum distance value for channel matching by determining the Hamming distance between the corresponding observation coding sequence and the time series coding of each logical channel for any pixel in the acquisition unit set, a threshold can be set to suppress the influence of noise. This ensures that the matching results are valid.

[0153] Response to the observation coding sequence corresponding to that pixel With the serial number j Time-series encoding of logical channels Hamming distance between Less than or equal to the threshold It can be determined that the pixel is related to the sequence number. j Logical channel matching.

[0154] For example, in pixels ( c, d ) corresponding observation coding sequence After calculating the Hamming distance with the time series codes of other logical channels, in response to, for example, the time series code of logical channel 3, the pixel ( c, d The Hamming distance between the observed encoded sequences corresponding to a pixel is 0, which indicates that the pixel ( c, d ) matches channel 3. Furthermore, in response to this pixel ( c, d ) corresponding observation coding sequence After calculating the Hamming distance with the time series codes of all logical channels, the time series code of logical channel 1 is related to the pixel ( c, d The Hamming distance 1 between the corresponding observed coded sequences is the minimum of all Hamming distances, and , can determine pixels ( c, d Matched with logical channel 1; in response to the time-series encoding of logical channel 1 and the pixel ( c, dThe Hamming distance 1 between the corresponding observed coded sequences is the minimum of all Hamming distances, and , can determine pixels ( c, d It does not match logical channel 1.

[0155] In one implementation, regarding pixel coordinates ( x, y Is it indexed by serial number? j The structured light projection pattern of the logical channel, i.e., the pixel ( x, y Is it consistent with the serial number? j Logical channel matching can be determined using binary decoding. Binary decoding interprets a binary sequence directly as an integer value and then compares these integer values. For example, the decimal integer value of [011010] is 26, and the decimal integer value of [001010] is 10. When using binary decoding, if the difference occurs in different positions, the decoded decimal values ​​may differ significantly. Generally, the two values ​​need to be exactly equal to determine if they match. For example, the two codes mentioned above differ only in the second bit from the left, and the two decimal values ​​differ by 16.

[0156] In one implementation, regarding pixel coordinates ( x, y Is it indexed by serial number? j The structured light projection pattern of the logical channel, i.e., the pixel ( x, y Is it consistent with the serial number? j Logical channel matching can be determined using Gray code decoding. The Gray code sequence belonging to a specific channel after decoding can then be used to determine the pixel (…). x, y ) and the serial number is j Logical channel matching.

[0157] Figure 7 This is a schematic diagram illustrating the principle of spot aggregation and center positioning according to an exemplary embodiment of this disclosure.

[0158] like Figure 1 and Figure 7 As shown, the processing component 400 can also perform spot aggregation, aggregating multiple acquisition units that are determined to be matched with the same logical channel and are spatially adjacent into a single spot region. The center coordinates of each spot region are calculated through center positioning.

[0159] Specifically, taking the acquisition component 300 as including an event camera and the time-stamped acquisition data sequence as an event sequence as an example, the luminous units projected by the projection component 200 will affect multiple related pixels in the luminous area to generate events. For example, a speckle may cover a diameter of 0.5 mm, and the event camera may record dozens of pixels that cause the event by the speckle flicker.

[0160] After filtering, decoding, and channel matching, a logical channel matching each pixel has been obtained. Multiple acquisition units that are determined to match the same logical channel and are spatially adjacent can be aggregated into a single spot area. Spatially adjacent units can be determined using an aggregation distance threshold. For example, the aggregation distance threshold could be 10 pixels. It should be noted that this aggregation distance threshold is only an example and can be adjusted according to actual needs.

[0161] In other words, within the aggregated pixel set, except for the condition of matching the same logical channel, the difference between the horizontal and vertical coordinates of any two pixels does not exceed the aggregation distance threshold. Alternatively, a Euclidean distance metric can be used, where the distance between any two pixels does not exceed the aggregation distance threshold.

[0162] Specifically, for the same channel j Collection of acquisition units By using methods such as connected component clustering, the first... i A spot area .

[0163] Taking the connected component clustering method as an example, aggregating multiple acquisition units that are determined to match the same logical channel and are spatially adjacent after decoding into a spot region may include: determining the neighboring pixels of each pixel; and clustering based on the neighborhood relationship between pixels to form a cluster.

[0164] In some implementations, two adjacent pixels can be considered to belong to the same neighborhood. For example, either the horizontal or vertical coordinates of the two pixels differ by 1; or, within the pixel ( x, y When a pixel is a non-boundary pixel, its neighboring pixels are Furthermore, the range of neighboring pixels can be expanded; for example, the eight pixels surrounding a non-boundary pixel can be considered as neighboring pixels. Alternatively, the Euclidean distance between pixels can be used as a measure; if the Euclidean distance between two pixels is less than a set threshold, the two pixels are considered neighboring pixels.

[0165] After determining the neighboring pixels of each pixel, clusters can be formed based on the neighborhood relationships between pixels, where the connectivity between pixels can be determined.

[0166] For example, algorithms such as flood fill (flood algorithm) and union-find can be used to cluster pixels based on their neighborhood relationships. The following explanation uses flood fill as an example. This process may include: initialization; traversing unvisited pixels; expanding neighboring pixels; and generating clusters.

[0167] During initialization, all pixels can be marked as "unvisited" to create an empty cluster set C={}; during the traversal of unvisited pixels, one unvisited pixel can be selected. Use it as a seed point to initialize a temporary cluster. and mark the pixel. "Visited"; during the expansion of neighboring pixels, temporary clusters can be traversed. For each pixel in the cluster, add unvisited pixels from its neighborhood set N(p) to cluster c and mark them as "visited"; repeat this step until no new pixels can be added to the temporary cluster c; during the generation of clusters, the temporary cluster c can be added to the cluster set C, and the process returns to the step "traversing unvisited pixels" until all pixels have been visited, and the spot region is obtained through connected component clustering. .

[0168] After multiple acquisition units that have been decoded and determined to be matched with the same logical channel and are spatially adjacent are aggregated into a single spot area, the center coordinates of each spot area can be calculated through center positioning.

[0169] For example, calculating the coordinates of the center pixel of the spot region can include defining the weights using formula (6):

[0170] .

[0171] in, For pixels In an encoding cycle The event count (or event rate) within the event; For pixels The weight.

[0172] According to formula (7), the center coordinates (centroid) of the spot region can be calculated based on the weights:

[0173] (7)

[0174] in, For the first The first logical channel The center coordinates of each light spot region; The x-axis coordinate of the center of the non-homogeneous spot in the image coordinate system; The y-axis coordinate of the center of the non-homogeneous spot in the image coordinate system; For the first The first logical channel The set of pixels corresponding to each spot area; For pixels The weights; For the light spot area Pixel coordinates within; Sum the weights × coordinates of all pixels within the spot area; Sum the weights of all pixels within the spot area.

[0175] In other words, based on the number of events per pixel within the spot area, the weight of each event, and the total number of events within the spot area, the coordinates of the center pixel of the spot area can be determined. The pixel with the most events indicates that the flickering events of the projected spot have been collected most fully, thus preserving the encoded information to the greatest extent. This calculation method assigns a higher weight to the pixel with the most events, thus shifting the center towards the pixel that retains the most encoded information.

[0176] Figure 8 This is a schematic diagram illustrating the principle of three-dimensional reconstruction based on spatial intersection according to an exemplary embodiment of this disclosure.

[0177] like Figure 1 and Figure 8 As shown, in some embodiments of this disclosure, the processing component 400 can also calculate parallax by comparing the center coordinates of the light spot regions corresponding to the same logical channel with the preset projection coordinates of that channel in the projection component 200 to obtain parallax information. Using the principle of spatial intersection, the coordinates of the center of each light spot region in three-dimensional space are calculated, where the centers of all light spot regions together constitute the three-dimensional shape of one side of the contact component 100 on the first surface. In other words, based on the fixed position parameters between the projection component 200 and the acquisition component 300, the aforementioned three-dimensional shape can be reconstructed by processing the time-stamped acquisition data sequence through a spatial intersection algorithm.

[0178] Specifically, the center coordinates of the spot areas corresponding to the same logical channels and the original projection coordinates are used to form camera rays and projection rays in the camera coordinate system, and the intersection of these rays in three-dimensional space is used to obtain three-dimensional points.

[0179] The center coordinates of the light spot region obtained above Coordinates corresponding to the camera coordinate system It can be expressed using formula (8).

[0180] (8)

[0181] in, For the first The first logical channel The coordinates of the center of each light spot on the z=1 plane in the event camera coordinate system; This is the intrinsic parameter matrix of the event camera; For the first The first logical channel The center coordinates of each light spot region; Here are the pixel coordinates of the center of the light spot; the superscript T indicates the transpose of the matrix / vector. The center coordinates of the light spot area The image coordinate representation. This indicates that the image coordinates are projected onto the plane at z=1. The center image coordinates of the projected light spot region are... Through camera internal parameters Convert to points in the camera coordinate system.

[0182] The camera ray can be represented using the point method, which can be expressed by formula (9).

[0183] (9)

[0184] in, The coordinates of any point on the event camera ray (point method). The optical center of the camera is usually (0, 0, 0). It is a point on the plane with z=1 in the event camera coordinate system, obtained by transforming the coordinates of the center of the spot area through the intrinsic parameters of the event camera. It is a scaling factor that corresponds to the center coordinates of the light spot region in the camera coordinate system. Adjust to generate a series of points in the camera coordinate system. Formula (9) indicates that the camera ray is formed by a ray drawn from the camera optical center to these points.

[0185] Similarly, structured light can be represented by formula (10) in the event camera coordinate system.

[0186] (10)

[0187] in, The coordinates of any point on the structured light projection ray (point method); The projection center is represented in the event camera coordinate system. The two can be obtained by converting the fixed positional relationship between the acquisition component 300 and the projection component 200 through calibration parameters. It represents the projection coordinates of the point in the event camera coordinate system. Similarly, the coordinates of the projection component can be transformed to the event camera coordinate system through calibration parameters. This is a scaling factor used to adjust the position of a point along the direction of the projected light beam. >0.

[0188] like Figure 8 As shown, two rays can be calculated in space. and The distance. When The distance between the two light rays changes when different values ​​are taken. This can be determined by finding the shortest distance between the two spatial lines. (Least squares closed-form solution) can yield three-dimensional points. :

[0189] (11)

[0190] in, The optimal scaling factor is used to minimize the distance between the projected light and the camera light. The optimal scaling factor is used to minimize the distance between the event camera's light and the projected light. These are the coordinates of the optimal position point on the projected ray in the event camera coordinate system; These are the coordinates of the optimal position point on the event camera ray in the event camera coordinate system; It is a logical channel j The i By combining the 3D representation of the center of each spot region with the encoding time of the spot, all 3D point clouds can be represented as 3D point clouds with time information. .

[0191] In some embodiments of this disclosure, the processing component 400 may also construct a deformation field based on the three-dimensional topography of one side of the contact component 100 on the first surface, and quantify the deformation and deformation gradient of the first surface relative to the non-contact state of the contact component 100.

[0192] Specifically, a local coordinate system can be defined on one side of the first surface of the contact assembly 100. Define the reference surface for zero load on this side as .

[0193] Alternatively, the height field can be obtained by fitting / interpolating the three-dimensional point cloud using formula (12). ,

[0194] (12)

[0195] in, The local planar coordinate system of the first surface of the contact component is mapped to the event camera coordinate system; The current moment; For the logical channel at time t j The i A three-dimensional representation of the center of each light spot region; It can be one of Delaunay triangulation, RBF (Radial Basis Function Interpolation), or local plane fitting; for Local coordinates at time The height field below.

[0196] Furthermore, using formula (13), the deformation of the first surface in the non-contact state relative to the contact assembly 100 can be the aforementioned height field. With reference surface The difference ,

[0197] (13)

[0198] in, The local planar coordinate system of the first surface of the contact component is mapped to the event camera coordinate system; It serves as a reference height field when the contact components are not in contact (zero load), and can be pre-stored through calibration; for Local coordinates at time The height field below; for Local coordinates at time The deformation can also be understood as the height change of the first surface of the contact component relative to the non-contact state of the contact component.

[0199] Using formula (14), based on deformation... Change the threshold to segment the contact area.

[0200] (14)

[0201] in, The local planar coordinate system of the first surface of the contact component is mapped to the event camera coordinate system; for Local coordinates at time The deformation can also be understood as the height change of the first surface of the contact component relative to the non-contact state of the contact component; The height change threshold for contact determination; for The contact area at any given moment.

[0202] In addition, the contact area can be segmented according to curvature or deformation gradient, etc., and this disclosure does not limit this.

[0203] Alternatively, using formula (15), the tangential displacement (under the small deformation approximation) can be estimated by the displacement of the spot center on the image plane over time:

[0204] (15)

[0205] in, The local planar coordinate system of the first surface of the contact component is mapped to the event camera coordinate system; The Jacobian mapping obtained from the calibration can map pixel displacement to the tangential displacement field of the first surface of the contact component 100. In other words, Local coordinates at time t The tangential displacement field below; Let be the pixel displacement of the center of the light spot at time t.

[0206] After obtaining the tangential displacement field Then, the slip index can be defined using formula (16). ,

[0207] (16)

[0208] in, Let be the slip index at time t; This represents the area of ​​the contact region between the contact component and the external object; Let L2 be the norm of the vector; It is a micro-element of area; The time rate of change of the tangential displacement; double integral for Contact area at any moment The inner integral.

[0209] Referring to formula (17), in response to the slip index Furthermore, the contact area boundary has a large rate of change, which triggers a slip alarm output for use in tasks such as capture closed-loop control.

[0210] .

[0211] in, for The contact area at any given moment; This represents the area of ​​the contact region between the contact component 100 and the external object; A threshold representing the rate of change of the area of ​​the contact region; Let be the rate of change of the area of ​​the contact region at time t.

[0212] In some embodiments of this disclosure, the processing component 400 may also perform pressure inversion based on the deformation field, converting deformation information into pressure / pressure distribution of external objects acting on the contact component 100 through a physical model or a data-driven model, and outputting tactile data. The pressure inversion may be achieved through data-driven calibration mapping, physical model inversion, or a fusion inversion of data-driven calibration mapping and physical model inversion.

[0213] Specifically, in data-driven calibration mapping, calibration data can be acquired by a standard force sensor and different indenters under different forces. For example, the standard force sensor can be a six-dimensional force / torque sensor; the indenter can be any known shape, such as a spherical indenter or a cylindrical indenter.

[0214] A six-dimensional force / torque sensor can acquire three-dimensional force. With three-dimensional torque The contact end of a spherical indenter is a standard sphere (e.g., a carbide sphere with a diameter of 5mm to 20mm), suitable for pressure calibration of point contact or small-area surface contact. The contact end of a cylindrical indenter is a cylindrical plane (e.g., a circular end face with a diameter of 10mm to 30mm), suitable for pressure calibration of large-area uniform contact.

[0215] Data acquisition under different forces may include: controlling the indenter to act vertically or obliquely on the first surface at a preset force (e.g., 0N~300N, with a gradient interval of 10N). A standard force sensor records the force / torque data at each force level, converts it into pressure, and acquires the deformation characteristic data of the first surface. The deformation characteristic data may include: Δh, Δh, local curvature, spot displacement vector, contact area, etc., where Δh is the height change. Δh is the gradient of height variation.

[0216] Using formula (18), the pressure distribution is discretized into grid vectors. Deformation characteristics are discretized into (For example Local curvature, spot displacement vector, elliptic second moment, contact area, etc., are used to give linear / nonlinear mappings.

[0217] (18).

[0218] in, This is the pressure distribution vector; The mapping model can be at least one of linear regression, piecewise polynomial, lightweight neural network, lookup table interpolation, etc. This is the deformation eigenvector.

[0219] In some implementations, as shown in formula (19), ridge regression can also be used.

[0220] (19)

[0221] in, This is the regression coefficient matrix, where each column corresponds to the regression coefficient of one dimension of the pressure grid vector; is the Frobenius norm, which characterizes the overall error of matrix elements; This is the ridge regression regularization coefficient, used to balance fitting accuracy and parameter complexity; This is the pressure vector predicted by regression. It is the deformation feature vector; For the regression coefficient matrix Find the minimum value; This is the pressure distribution vector.

[0222] As shown in formula (20), by differentiating the above objective function with respect to W and setting the derivative to 0, a closed-form solution can be obtained.

[0223] (20)

[0224] in, This is the optimal regression coefficient matrix for ridge regression; For multiple sets of calibration sample feature matrices; For multiple sets of calibration sample feature matrices Transpose of; To correspond to the actual pressure matrix; This is the ridge regression regularization coefficient, used to balance fitting accuracy and parameter complexity; It is an identity matrix.

[0225] In addition, in data-driven calibration mapping, the mapping model can also be at least one of the following: linear regression, piecewise polynomial, lightweight neural network, and lookup table interpolation.

[0226] Optionally, the linear regression model is suitable for scenarios where deformation and pressure have an approximately linear relationship, as shown in formula (21), and a multivariate linear equation can be constructed.

[0227] (twenty one)

[0228] in, To predict the pressure value; The number of deformation features; For the first k One deformation characteristic parameter; For feature weights; For bias terms; To The weights of each deformation feature are summed by multiplying their weights by the features themselves. Formula (21) obtains the optimal weight parameters by fitting a calibration dataset using the least squares method. The advantages of this model are low computational complexity, strong interpretability, and ease of real-time deployment in engineering.

[0229] Optionally, the piecewise polynomial model can divide the calibration force range into multiple sub-intervals to address the nonlinear relationship between deformation and pressure, and use a polynomial model to fit within each sub-interval.

[0230] Optionally, lightweight neural network models are suitable for complex scenarios with highly nonlinear deformation features and pressure distributions. For example, lightweight network structures such as MobileNet, ShuffleNet, or custom small fully connected networks can be used. The input layer can be multidimensional deformation features, the hidden layers can be 2-3 layers (32-128 neurons), and the output layer can be a pressure distribution matrix. Furthermore, network parameters can be trained using a calibrated dataset, and the loss function (e.g., mean squared error loss) can be optimized using backpropagation. The advantages of this model are strong generalization ability, the ability to handle high-dimensional, strongly coupled feature data, and a lightweight structure that meets the real-time computing requirements of embedded devices.

[0231] Optionally, the lookup table interpolation model is suitable for real-time calibration scenarios with extremely high computational efficiency requirements. For example, the calibration dataset can be pre-stored in a grid according to key feature parameters (e.g., height variation, contact area) to form a calibration lookup table. In practical applications, approximate pressure values ​​are obtained by looking up the table using the input real-time deformation feature parameters, and then linear interpolation or cubic spline interpolation algorithms are used to calculate the precise pressure distribution. The advantage of this model is its fast computation speed, the absence of complex iterative calculations, and its suitability for scenarios with limited hardware resources.

[0232] Furthermore, as shown in formula (22), the ridge regression model described above can be used to target the deformation characteristics of the input. Predicted pressure ,

[0233] (twenty two)

[0234] in, This is the predicted pressure vector based on the new deformation characteristics; This is the optimal regression coefficient matrix; This represents the newly acquired deformation feature vector. Alternatively, other models can be used to target the input deformation features. Predicted pressure .

[0235] Furthermore, based on the obtained pressure distribution The total normal force can be obtained according to the following formula (23). Output,

[0236] (twenty three)

[0237] in, These are points within the contact region that meet the conditions for contact region segmentation. The total normal force can be obtained by multiple integrations or summations of the pressure. In other words, Let be the total normal force exerted by the external object on the contact component at time t; Local coordinate system at time t Pressure distribution function under; For area infinitesimal elements; double integral This is the integral of the pressure over the contact area; For the first in the contact area The pressure of each grid; The area of ​​each grid cell; This indicates that the summation range covers all grid cells within the contact area.

[0238] In some embodiments of this disclosure, physical model inversion may include: treating the first surface as an elastic body surface, as shown in formula (24), where pressure and normal displacement... They satisfy a convolutional relationship.

[0239] (twenty four)

[0240] in, Let r be the normal displacement at point r; Let the coordinates of the target point on the first surface of the contact component be in the local coordinate system. Let be the integral variable of the first surface of the contact component in the local coordinate system; The pressure at the target point on the first surface of the contact component in the local coordinate system; It is the Green's function; This is the total integral of the first surface of the contact component.

[0241] At point Apply a unit point force (normal concentrated force) at the point The function that causes the normal displacement is the Green's function. ; It can be determined by "elasticity theory / finite element method / empirical fitting".

[0242] As shown in formula (25), after discretizing the pressure, the normal displacement vector With pressure vector satisfy,

[0243] (25)

[0244] in, The parameters are determined by material parameters (e.g., Young's modulus E, Poisson's ratio ν) and geometry (e.g., thickness). The system matrices, as determined by parameters such as the system matrix, can be obtained through analytical approximation or offline finite element calculation. For example, a half-space model, a thin film model, or a finite element approximation can be used, and material parameters can be corrected through a small amount of calibration. This is the normal displacement. This is the pressure distribution vector.

[0245] As shown in formula (26), the inversion can be performed using regularized least squares.

[0246] (26)

[0247] in, Use a gradient operator (encouraging pressure smoothing) and explicitly add nonnegativity constraints. (Non-negative pressure enhances physical plausibility). This can be achieved using NNLS (Non-Negative Least Squares) or projected gradient; α can be understood as a regularization parameter, which can be used to control the weight of the regularization term. The system matrix, determined by material parameters (e.g., Young's modulus E, Poisson's ratio ν) and geometry (e.g., thickness d), can be obtained through analytical approximation or offline finite element calculation. Let be the pressure vector to be determined; The square of the L2 norm; This is the predicted normal displacement vector; The gradient of the pressure vector; This is the normal displacement.

[0248] As shown in formula (27), in some embodiments of this disclosure, the fusion inversion may include: providing prior information based on the above physical model, and then using data-driven correction.

[0249] (27)

[0250] in, The pressure vector obtained by the final inversion is consistent with the pressure distribution vector shown in formula (18); The mapping model can be at least one of linear regression, piecewise polynomial, lightweight neural network, lookup table interpolation, etc. It is the deformation feature vector; The initial pressure vector obtained from the physical model inversion can be determined by formula (26).

[0251] Alternatively, as shown in formula (28), the material parameters can also be regarded as learnable parameters, and a small number of calibration updates can be used to solve the errors caused by material aging, temperature changes and batch differences.

[0252]

[0253] in, The parameters are determined by material parameters (e.g., Young's modulus E, Poisson's ratio ν) and geometry (e.g., thickness). The system matrices, as determined by parameters such as (e.g., ) and (followed by) analytical approximation or offline finite element calculation, can be obtained. The Young's modulus of the contact component; The Poisson's ratio of the contact component; The thickness of the contact component; This is the pressure distribution vector; This is the normal displacement.

[0254] Therefore, the processing component 400 can output tactile data after the above processing, wherein the tactile data may include at least one of the following: contact state information between the contact component 100 and the external object; pressure distribution information of the external object acting on the contact component 100; total normal force of the external object acting on the contact component 100; shear force of the external object acting on the contact component 100; and slippage information of the external object relative to the contact component 100.

[0255] In addition, this disclosure also provides specific embodiments of the tactile sensing system 1000.

[0256] Example 1

[0257] refer to Figure 1 The contact component 100 can form the soft skin of the robot fingertip. The contact component 100 may include an outer layer, a middle layer and an inner layer stacked together. The inner layer may include a first surface projected by the structured light emitted by the projection component 200, and the outer layer may include a second surface in contact with the external object.

[0258] Alternatively, the outer layer can be made of silicone with a Shore A hardness of 10-20 and a thickness of 2-5 mm; the middle layer can be made of silicone with a Shore A hardness of 30-40 and a thickness of 1-3 mm; and the inner layer can be a white diffuse reflection layer formed on the middle layer, wherein the white diffuse reflection layer can have microtextures with a feature of 50-200 µm.

[0259] The projection component 200 can employ an 850nm VCSEL array, where it can control the output timing of multiple independent "logic channels." The encoded light from each logic channel is projected onto the first surface of the contact component 100, illuminating and forming a corresponding "pattern area." The number of logic channels can be 8 to 32. The number of steps in a time-series encoding cycle can be 6. The Hamming distance between time-series codes corresponding to adjacent pattern areas can be greater than or equal to a predetermined threshold, where the predetermined threshold can be 2. Replicating the encoding cycle allows the emitting units to switch between on / off or bright / dark states according to a repetitive timing pattern. The timing sequence within the encoding cycle can serve as a unique identifier, and redundancy verification is achieved through multi-cycle acquisition to enhance robustness.

[0260] The acquisition component 300 can be fixedly installed inside the cavity. The acquisition component 300 is positioned opposite to the first surface, or the optical path can be adjusted using a mirror so that the field of view of the acquisition component 300 completely covers the first surface.

[0261] One flash of the projection component 200 corresponds to a complete frame, and a sequence of 6 frames of events on the same acquisition unit of the acquisition component 300 can be used for time decoding matching. The sequence number is... j The time-series encoding of the logical channel is represented as Its encoding is a 6-bit sequence of 0s and 1s.

[0262] To ensure the correctness of time series decoding, a known initial reference state needs to be determined for the acquisition unit before structured light projection. This can be achieved by resetting the state of the acquisition unit in hardware or logic; or by pre-setting a specific starting state for time series encoding. For example, all time series encoding can be set to start with an "off" state.

[0263] For example, multiple time series codes may include [011010], [110110], [100100], etc. There can be 64 time series codes with a total of 6 bits in the coding space. Different time series codes can be used for different logical channels. 8 to 32 codes can be selected from the 64 codes and applied to the aforementioned 8 to 32 logical channels.

[0264] Optionally, the processing component 400 may filter the received time-stamped acquisition data sequence before obtaining the observation coding sequence of each acquisition unit. The filtering may include: counting the acquisition data of each acquisition unit within the coding period of the time series coding; and excluding acquisition units whose acquisition data count is lower than the response rate threshold to form a set of acquisition units for subsequent processing.

[0265] Therefore, statistics can be collected for each acquisition unit. In an encoding cycle Collect data within the system and calculate the event response rate. Events with response rates lower than the response rate threshold are excluded. After identifying the acquisition units, a collection of acquisition units can be obtained.

[0266] By performing time-series decoding on the data collected by the acquisition units in the acquisition unit set, the observation coding sequence of each acquisition unit can be obtained; based on the observation coding sequence, the time series code corresponding to each acquisition unit can be matched, or in other words, each acquisition unit can be matched to the corresponding logical channel.

[0267] For example, in the determination acquisition unit Is it related to logical channels? jMatching can be done using Hamming distance as the decision-making method. This can be achieved by comparing pixels one by one. Observation coding sequence Time-series encoding with logical channels The Hamming distance between them is used to determine the time series code corresponding to each pixel, or the logical channel matched to each pixel.

[0268] Multiple acquisition units that are determined to match the same logical channel and are spatially adjacent can be aggregated into a single spot region. For example, multiple acquisition units matching the same logical channel can be clustered according to a distance of at least 10 pixels between them. Centroid / ellipse fitting is then performed on the clustered pixels to obtain the center coordinates of each spot region.

[0269] Based on the center coordinates of each obtained light spot region, a three-dimensional point cloud or height field of one side of the contact component on the first surface is obtained using a disparity map and spatial intersection algorithm. Optionally, a six-dimensional force / torque sensor and a spherical / cylindrical indenter are used to collect the height change Δh and the actual pressure distribution (or the average pressure is calculated using the known indenter area), and a lightweight model is trained to output the pressure vector P(u,v).

[0270] Example 2

[0271] Building upon Embodiment 1, a texture with anisotropic optical characteristics can be introduced into the inner layer of the contact component 100, causing shear displacement to alter the shape / direction of the light spot. Therefore, the changes in the principal axis direction and major and minor axes of the ellipse are output during light spot region fitting, and the shear force is estimated by combining this with planar displacement. Slip detection is achieved based on the rate of change of the contact region boundary over time and the change in the displacement vector. In other words, the tactile data output by this embodiment can include the shear force exerted on the contact component 100 by an external object; and the slip information of the external object relative to the contact component 100.

[0272] Figure 9 This is a flowchart illustrating a tactile sensing method 2000 provided according to an exemplary embodiment of the present disclosure.

[0273] like Figure 9 As shown, embodiments of this disclosure also provide a tactile sensing method 2000, which includes:

[0274] Step S1: An active light field with time modulation characteristics is projected from the projection component onto the first surface of the contact component. The active light field is projected onto the first surface to form multiple patterned regions. The light intensity of adjacent patterned regions changes based on different time series encodings. The first surface faces away from the second surface of the contact component that is in contact with the external object.

[0275] Step S2: The acquisition component acquires multiple pattern areas and outputs a time-stamped acquisition data sequence, wherein the acquisition component includes multiple acquisition units arranged in an array.

[0276] Step S3, based on the received time-stamped acquisition data sequence, performs at least the following: based on the received time-stamped acquisition data sequence, obtain the observation coding sequence for each acquisition unit; identify multiple observation coding sequences to determine the time series coding to which the observation coding sequence belongs; based on the time series coding corresponding to each acquisition unit, reconstruct the three-dimensional morphology of one side of the contact component on the first surface; based on the reconstructed three-dimensional morphology, invert the pressure distribution of the external object acting on the contact component; and output tactile data based on the pressure distribution.

[0277] The tactile sensing method 2000 proposed in this disclosure is based on a projection component and a acquisition component. The acquisition component includes multiple acquisition units arranged in an array and generates a time-stamped acquisition data sequence based solely on changes in light intensity, making it suitable for tactile sensing such as high-speed sliding or vibration. Furthermore, because this acquisition component generates a time-stamped acquisition data sequence based solely on changes in light intensity, it detects less redundant data, significantly reducing the operating power consumption of the tactile sensing method 2000.

[0278] In addition, the contact component includes a first surface and a second surface arranged opposite to each other, wherein an external object can contact the second surface, and the projection component can project an active light field onto the first surface, wherein the active light field is projected onto the first surface to form multiple patterned regions, and the light intensity of adjacent patterned regions changes based on different time series encodings. The acquisition component can acquire the light intensity changes of the reflected light from the active light field on the first surface and output a time-stamped acquisition data sequence.

[0279] When the contact component is not deformed, the reflected light is affected by the intensity changes of the active light field, resulting in corresponding intensity changes. The acquisition component can output a time-stamped acquisition data sequence based on these intensity changes. When the contact component is deformed, the reflected light is affected not only by the intensity changes of the active light field but also by the deformation of the first surface. In this case, the acquisition component can still output a time-stamped acquisition data sequence based on the intensity changes of the reflected light. This output time-stamped acquisition data sequence includes the deformation information of the contact component. Therefore, the tactile sensing method 2000 can reconstruct the three-dimensional shape of one side of the first surface of the contact component, the pressure distribution of an external object acting on the contact component, or the tactile data of an external object acting on the contact component based on this time-stamped acquisition data sequence.

[0280] Regardless of whether the contact component deforms, the acquisition component can output a time-stamped acquisition data sequence based on the change in the intensity of the reflected light. This consistently effective data acquisition mechanism provides a reliable foundation for subsequent stable and high-precision deformation field reconstruction and pressure distribution inversion.

[0281] The time-stamped acquisition data sequence may include a data stream acquired by an event camera or other photoelectric detection array that generates synchronous or asynchronous acquisition data based on changes in light intensity. The acquired data in the time-stamped acquisition data sequence includes at least: the time information, location information, and light intensity change information characterizing the direction or magnitude of the change in the intensity of the reflected light. When the acquisition unit receives the light intensity change caused by the reflected light, and the change magnitude exceeds a preset light intensity threshold, a data acquisition point can be generated.

[0282] Taking an event camera as an example, the time-stamped data sequence acquired by the event camera can be described as an event sequence. An event sequence may include multiple events, which are asynchronous events independent of each other in terms of time, spatial location, and event polarity. Specifically, an event camera may include multiple event pixels (hereinafter referred to as pixels) with different pixel positions, and these pixels can form a pixel array. Reflected light illuminating the pixel array can form a light spot area. Based on the reflected light acquired by one or more pixels in the light spot area, an event sequence characterizing the spatiotemporal information of the reflected light can be generated. The events in the event sequence may include the time information, location information, and polarity information of changes in the intensity of the reflected light. When a pixel receives a brightness change caused by reflected light, and the change exceeds a preset threshold, an event can be generated.

[0283] The tactile sensing method 2000 proposed in this disclosure projects a time-series encoded active light field onto a first surface of a contact component via a projection component, and detects it using a collection component. The collection component 300 includes multiple collection units arranged in an array, and generates a time-stamped collection data sequence based solely on changes in light intensity. Since the collection units respond only to changes in light intensity and output collection data with precise timestamps, there are no exposure delays or fixed frame rate limitations, enabling near real-time capture of microsecond-level deformation signals caused by high-speed sliding or vibration, thereby achieving real-time dynamic tactile sensing of dynamic contact processes.

[0284] In some embodiments of this disclosure, reconstructing the three-dimensional morphology of one side of the contact component on the first surface based on the time-series encoding corresponding to each acquisition unit may include: processing the time-stamped acquisition data sequence through a spatial intersection algorithm based on fixed position parameters between the projection component and the acquisition component to reconstruct the three-dimensional morphology.

[0285] In some embodiments of this disclosure, an active light field is projected from the projection component onto the first surface of the contact component, wherein the active light field is projected onto the first surface to form a plurality of patterned regions, and the light intensity of adjacent patterned regions varies based on different time series encodings, including: varying the light intensity of two patterned regions that are spatially greater than or equal to a predetermined distance according to the same time series encoding.

[0286] Multi-channel time-series encoding of the projection component is used to drive the projected active light field's spots / speckles / stripes to change brightness / dullness according to the encoded pulses. For ease of description, the spot / speckle / stripe will be referred to as a luminous unit. One or more luminous units can be encoded into a logical channel, and luminous units within the same logical channel can present the same luminous rhythm according to the same time-series encoding.

[0287] The light intensity of two pattern regions spatially greater than or equal to a predetermined distance is varied according to the same time-series encoding. In other words, multiple logical channels are divided into multiple groups, with a unified control strategy within each group and a differentiated strategy between groups. For example, the logical channels corresponding to two spatially distant pattern regions that are greater than or equal to a predetermined distance are assigned the same time-series encoding and grouped together for unified driving. Different time-series encodings are assigned to the logical channels corresponding to spatially adjacent or close pattern regions. This spatial distance-based grouping method significantly optimizes the system's encoding resources and control complexity, achieving a balance between performance and efficiency.

[0288] In some embodiments of this disclosure, an active light field is projected from the projection component onto the first surface of the contact component, wherein the active light field forms multiple patterned regions on the first surface, and the light intensity of adjacent patterned regions varies based on different time-series codes. This may further include setting the Hamming distance between the time-series codes corresponding to adjacent patterned regions to be greater than or equal to a predetermined threshold. Maintaining a sufficient Hamming distance can effectively prevent the confusion of different logic channels during decoding due to signal interference.

[0289] In some embodiments of this disclosure, the tactile sensing method 2000 may further include: filtering the received time-stamped acquisition data sequence before obtaining the observation encoding sequence of each acquisition unit, including: counting the acquisition data of each acquisition unit within the encoding period of the time series encoding; and excluding acquisition units whose acquisition data count is lower than a response rate threshold, to form a set of acquisition units for subsequent processing. By setting a response rate threshold, data corresponding to acquisition units with an average data volume lower than the threshold are excluded from subsequent calculations, thereby obtaining a set of acquisition units. This directly sparsifies the data to be processed, significantly reducing the complexity of subsequent operations.

[0290] In some embodiments of this disclosure, the material of the contact component includes a light-transmitting material. The tactile sensing method 2000 may further include: acquiring a data sequence based on a time stamp, detecting changes in light intensity transmitted through the light-transmitting material to the side containing the first surface; and triggering the projection component to project an active light field onto the first surface in response to a preset condition met by the change in light intensity. In other words, the aforementioned response rate threshold can be dynamically adjusted based on changes in the intensity of reflected light caused by the contact component's contact with an external object, thus balancing the sensitivity and false detection rate of the tactile sensing method 2000.

[0291] In some embodiments of this disclosure, the tactile sensing method 2000 may further include: aligning the timing of a time-stamped data sequence with the timing of a time-series encoded sequence.

[0292] Specifically, the processing component can reconstruct the light intensity state (on or off) observed by each acquisition unit of the acquisition component within the encoding time window defined by the projection component, based on the timestamps in the time-stamped data sequence, to form an observation encoding sequence. If there is a discrepancy between the timing references of the projection component and the acquisition component, the acquired data cannot be accurately mapped to its corresponding encoding time window, leading to errors in state judgment and decoding. Therefore, to ensure that the acquired data accurately reflects the encoding characteristics of structured light, the timestamps of the time-stamped data sequence and the encoding time sequence of the projection component can be correlated. Timing alignment. Optionally, one or a combination of the following methods can be used: on the one hand, it can be achieved through hardware timing alignment, such as using a global clock signal, hardware trigger line or precision time protocol to unify the clock reference of the projection component and the acquisition component; on the other hand, it can be achieved through software calibration, by having the acquisition component record the known reference timing signal emitted by the projection component, calculate the time deviation and perform software compensation, thereby establishing a time alignment relationship at the data processing level.

[0293] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0294] Figure 10 A schematic block diagram of an example electronic device 3000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobility methods, such as personal digital processing, cellular phones, smartphones, wearable devices, and other similar computing methods. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0295] like Figure 10 As shown, device 3000 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 302 or a computer program loaded from storage unit 308 into random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of device 3000. The computing unit 301, ROM 302, and RAM 303 are interconnected via bus 304. Input / output (I / O) interface 305 is also connected to bus 304.

[0296] Multiple components in device 3000 are connected to I / O interface 305, including: input unit 306, such as a keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as a disk, optical disk, etc.; and communication unit 309, such as a network card, modem, wireless transceiver, etc. Communication unit 309 allows device 3000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. Computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Computing unit 301 performs the various devices and processes described above, such as tactile sensing methods. For example, in some embodiments, the tactile sensing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on device 3000 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by computing unit 301, one or more steps of the tactile sensing method described above may be performed. Alternatively, in other embodiments, computing unit 301 may be configured to perform the tactile sensing method by any other suitable means (e.g., by means of firmware).

[0297] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input method, and at least one output method, and transferring data and instructions to the storage system, the at least one input method, and the at least one output method.

[0298] Program code for implementing the apparatus of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing method, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0299] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, method, or apparatus. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, methods, or apparatuses, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0300] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display method for showing information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing method (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of methods can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0301] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0302] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0303] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, apparatuses, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0304] Various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0305] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0306] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A tactile sensing system, characterized in that, include: Contact components are configured to come into contact with external objects and deform. The projection component is configured to project an active light field with time modulation characteristics onto a first surface of the contact component. The active light field forms a plurality of patterned regions on the first surface. The light intensity of adjacent patterned regions changes based on different time series codes. The first surface faces away from a second surface of the contact component that is in contact with the external object. The acquisition component includes multiple acquisition units arranged in an array and is configured to acquire multiple pattern regions and output a time-stamped acquisition data sequence; as well as The processing component is configured to perform at least some of the following: Based on the received time-stamped data sequence, the observation coding sequence of each acquisition unit is obtained; Identify multiple observation coding sequences to determine the time series coding to which the observation coding sequence belongs; Based on the time series encoding corresponding to each of the acquisition units, the three-dimensional morphology of the contact component on one side of the first surface is reconstructed; Based on the reconstructed three-dimensional topography, the pressure distribution exerted by the external object on the contact component is inverted; and Tactile data is output based on the pressure distribution.

2. The system according to claim 1, characterized in that, The tactile data includes at least one of the following: The contact state information between the contact component and the external object; The pressure distribution information of the external object acting on the contact component; The total normal force exerted by the external object on the contact component; The shear force exerted by the external object on the contact assembly; and The sliding information of the external object relative to the contact component.

3. The system according to claim 1, characterized in that, The processing component is further configured to: Based on the fixed position parameters between the projection component and the acquisition component, the time-stamped acquisition data sequence is processed by a spatial intersection algorithm to reconstruct the three-dimensional shape.

4. The system according to claim 1, characterized in that, The projection component is also configured to: The light intensity of two patterned regions that are spatially greater than or equal to a predetermined distance will vary according to the same time-series encoding.

5. The system according to claim 1, characterized in that, The projection component is also configured to: The Hamming distance between the time series codes corresponding to the adjacent pattern regions is set to be greater than or equal to a predetermined threshold.

6. The system according to claim 1, characterized in that, The processing component is further configured to: Before obtaining the observation encoding sequence for each acquisition unit, the received time-stamped acquisition data sequence is filtered, including: The data acquired by each acquisition unit within the encoding period of the time series encoding is counted; and The acquisition units whose data count is lower than the response rate threshold are excluded to form a set of acquisition units for subsequent processing.

7. The system according to claim 1, characterized in that, The contact component includes: An outer layer for receiving contact from the external object, and including the second surface; and The inner layer, stacked on top of the outer layer, includes the first surface. The inner layer includes at least one of the following: a diffuse reflection layer, a scattering layer with a microstructure, a deformable speckle layer, and a texture with anisotropic optical characteristics.

8. The system according to claim 7, characterized in that, The contact assembly further includes: The intermediate layer is located between the outer layer and the inner layer. Wherein, the hardness of the intermediate layer near the outer layer is less than the hardness of the intermediate layer near the inner layer; and The hardness of the intermediate layer increases in a gradient.

9. The system according to claim 1, characterized in that, The system has a first operating mode and a second operating mode, and the material of the contact component includes a light-transmitting material. In the first operating mode, the processing component is further configured to acquire a data sequence based on the time stamp and detect changes in light intensity transmitted through the light-transmitting material to the side containing the first surface; and In response to the change in light intensity satisfying a preset condition, the system is triggered to switch to the second working mode, so that the projection component projects the active light field onto the first surface.

10. The system according to claim 1, characterized in that, The active light field includes: Structured light, wherein the pattern of the structured light is selected from at least one of speckle pattern, stripe pattern, Gray code pattern, and binarized dot matrix pattern.

11. The system according to claim 1, characterized in that, The system also includes: A timing control module is integrated into the projection component, the acquisition component, or the processing component, and is configured to align the timing of the time-stamped acquired data sequence with the timing of the time-series encoded data.

12. A tactile perception method, characterized in that, include: An active light field with time modulation characteristics is projected from the projection component onto the first surface of the contact component. The active light field forms multiple patterned regions on the first surface. The light intensity of adjacent patterned regions changes based on different time series encodings. The first surface faces away from the second surface of the contact component that is in contact with an external object. Multiple pattern regions are acquired by an acquisition component, and a time-stamped acquisition data sequence is output, wherein the acquisition component includes multiple acquisition units arranged in an array; as well as Based on the received time-stamped data sequence, perform at least some of the following: Based on the received time-stamped data sequence, the observation coding sequence of each acquisition unit is obtained; Identify multiple observation coding sequences to determine the time series coding to which the observation coding sequence belongs; Based on the time series encoding corresponding to each of the acquisition units, the three-dimensional morphology of the contact component on one side of the first surface is reconstructed; Based on the reconstructed three-dimensional topography, the pressure distribution exerted by the external object on the contact component is inverted; and Tactile data is output based on the pressure distribution.

13. The method according to claim 12, characterized in that, The reconstructing of the three-dimensional morphology of the contact component on one side of the first surface based on the time-series encoding corresponding to each of the acquisition units includes: Based on the fixed position parameters between the projection component and the acquisition component, the time-stamped acquisition data sequence is processed by a spatial intersection algorithm to reconstruct the three-dimensional shape.

14. The method according to claim 12, characterized in that, An active light field with time-modulated characteristics is projected from the projection component onto the first surface of the contact component. The active light field forms multiple patterned regions on the first surface, and the light intensity of adjacent patterned regions changes based on different time-series encoded variations, including: The light intensity of two patterned regions that are spatially greater than or equal to a predetermined distance will vary according to the same time-series encoding.

15. The method according to claim 12, characterized in that, The projection of an active light field with time-modulated characteristics onto the first surface of the contact component by the projection component, wherein the active light field forms multiple patterned regions on the first surface, and the light intensity of adjacent patterned regions changes based on different time-series encoded variations, further includes: The Hamming distance between the time series codes corresponding to the adjacent pattern regions is set to be greater than or equal to a predetermined threshold.

16. The method according to claim 12, characterized in that, The method further includes: Before obtaining the observation encoding sequence for each acquisition unit, the received time-stamped acquisition data sequence is filtered, including: The data acquired by each acquisition unit within the encoding period of the time series encoding is counted; and The acquisition units whose data count is lower than the response rate threshold are excluded to form a set of acquisition units for subsequent processing.

17. The method according to claim 12, characterized in that, The material of the contact component includes a light-transmitting material, and the method further includes: Based on the time-stamped data sequence, the change in light intensity transmitted through the light-transmitting material to the side containing the first surface is detected; and In response to the change in light intensity satisfying a preset condition, the projection component is triggered to project the active light field onto the first surface.

18. The method according to claim 12, characterized in that, The method further includes: Align the timing of the time-stamped data sequence with the timing of the time-series encoding.

19. An electronic device, characterized in that, include: A processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the tactile sensing method as described in any one of claims 12 to 18.

20. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the tactile sensing method according to any one of claims 12 to 18.

21. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the tactile sensing method according to any one of claims 12 to 18.