Digital touch with artificial robot fingertips

By integrating high-resolution sensors and omnidirectional optical systems in artificial finger sensors and combining AI neural network accelerators, the problem of difficulty in providing multimodal digital touch sensing in existing systems is solved, and high sensitivity detection and real-time processing of touch information are achieved.

CN120201977APending Publication Date: 2025-06-24CTRL-LABS CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380078364.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-07
Filing Date
2023-11-09
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing systems have difficulty providing rich multimodal digital touch sensing capabilities while retaining the shape elements of human fingers.

Method used

Using artificial finger sensors with high resolution sensors, including omnidirectional optical systems and AI neural network accelerators, can acquire multimodal signals and process data in real time.

Benefits of technology

High sensitivity detection of spatial features is achieved, features as small as 7μm can be distinguished, normal forces, shear forces, vibration, odor and heat are sensed, and the reflection arc of humans is imitated, improving the performance of touch digitization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201977A_ABST
    Figure CN120201977A_ABST
Patent Text Reader

Abstract

In one embodiment, a system includes a silicone hemispherical dome and an omnidirectional optical system. The dome includes a surface comprising a reflective silver film layer. The optical system includes a lens including a plurality of lens elements, with a first lens element in direct contact with the hemispherical dome without an air gap. The lens is configured to collect scattering of internal incident light generated by the reflective silver film layer. The optical system also includes an image sensor configured to generate image data from data acquired by the lens. The system also includes a non-image sensor. The system also includes processors and a non-transitory memory coupled to the processors, the non-transitory memory including instructions executable by the processors, the processors being operable when executing the instructions to: access image data from the omnidirectional optical system and sense data from the non-image sensors, and generating, by a machine learning model, a touch digitization based on the accessed image data and sensing data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 383,069, filed on Nov. 9, 2022, under 35 U.S.C. § 119(e), which is incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to robotics, and more particularly, to hardware and software for intelligent robot sensing. Background Art

[0004] Artificial intelligence (AI) is the intelligence of machines or software. AI technologies are widely used in industrial, governmental, and scientific fields, such as advanced web search engines, recommendation systems, understanding human speech, self-driving cars, generative or creative tools, and competing at the highest levels in strategic games.

[0005] Robotics is an interdisciplinary branch of electronics and communication, computer science, and engineering. Robotics involves the design, construction, operation, and use of robots. The goal of robotics is to design machines that can help and assist humans. Robotics integrates fields such as mechanical engineering, electrical engineering, information engineering, mechatronics engineering, electronics, biomedical engineering, computer engineering, control systems engineering, software engineering, mathematics, etc. Machines developed in the field of robotics can perform tasks automatically and complete various jobs that humans may not be able to do. Some robots require user input to operate, while other robots operate autonomously. Summary of the Invention

[0006] Touch is a sensing modality that can provide rich information about object properties and interactions with the physical environment. Both humans and robots benefit from using touch to sense and interact with their surroundings. However, no existing system can provide rich multimodal digital touch sensing capabilities while preserving the form factor of a human finger. Embodiments disclosed herein can improve the digitization of touch through technological advancements embodied in artificial finger-shaped sensors with enhanced sensing capabilities. In a particular embodiment, an artificial fingertip can include high-resolution sensors (e.g., approximately 8.3 million contact pixels (taxels)) that respond to omnidirectional touch, acquire multimodal signals, and use on-device artificial intelligence to process data in real time. For example, in a particular embodiment, evaluations have shown that an artificial fingertip can resolve spatial features as small as 7 μm, sense normal and shear forces with a resolution between, e.g., 1 mN and 1.3 mN, perceive vibrations up to, e.g., 9 kHz to 11 kHz, sense odors, and even sense heat. Additionally, an on-device AI neural network accelerator can act as a peripheral nervous system on a robot and mimic a human reflex arc. These results indicate that embodiments disclosed herein can digitize touch with enhanced performance. Embodiments disclosed herein can be applied to fields including robotics (industrial, medical, agricultural, and consumer levels), virtual reality and telepresence, prosthetics, and e-commerce. Although this disclosure describes digitizing a particular modality in a particular way, this disclosure contemplates digitizing any suitable modality in any suitable way.

[0007] In a particular embodiment, a system for touch digitization can include a silicone hemispherical dome that includes a surface containing a reflective silver film layer. The system can additionally include an omnidirectional optical system that includes: a lens that includes a plurality of lens elements; and an image sensor that is configured to generate image data based on data acquired by the lens. In a particular embodiment, a first lens element of the plurality of lens elements can be in direct contact with the silicone hemispherical dome without an air gap. The lens can be configured to acquire scattering of internally incident light produced by the reflective silver film layer. The system can also include one or more non-image sensors disposed below the omnidirectional optical system. The system can also include one or more processors and a non-transitory memory coupled to the processors, the non-transitory memory including instructions executable by the processors. In a particular embodiment, the processors can be operative, when executing the instructions, to: access image data from the omnidirectional optical system and sensing data from the one or more non-image sensors, and generate the touch digitization based on the accessed image data and sensing data by one or more machine learning models.

[0008] In certain embodiments, an artificial fingertip for touch digitization may include a silicone hemispherical dome. The artificial fingertip may also include an omnidirectional optical system that includes: a lens that includes a plurality of lens elements; and an image sensor that is configured to generate image data based on data collected by the lens. In certain embodiments, a first lens element of the plurality of lens elements may be in direct contact with the silicone hemispherical dome without an air gap. The artificial fingertip may also include one or more non-image sensors disposed below the omnidirectional optical system.

[0009] Certain embodiments disclosed herein may provide one or more technical advantages. The technical advantages of these embodiments may include increased sensitivity to input stimuli by reconstructing touch surface topology signals (images) using dynamic illumination with variable wavelengths and directions, as the systems disclosed herein go beyond traditional Lambertian scattering paradigms towards a surface with a controlled degree of scattering, and also employ fundamental methods of optimizing material properties to achieve the highest sensitivity to spatial features, which is achieved by developing a new process for chemically growing a silver thin film layer on the fingertip surface. Another technical advantage of these embodiments may include high spatial resolution in capturing minute details on the fingertip surface, which allows bypassing traditional methods of capturing normal and shear forces by using markers. The embodiments disclosed herein utilize a 3-zone solid immersion hyperfisheye lens design for a touch digitization system in a non-air-gap configuration of PDMS material. Due to the design and specific requirements of non-human imaging, the embodiments disclosed herein are capable of maintaining excellent performance of the modulation transfer function (MTF) across the entire field of view of the lens, thus ensuring spatial performance along the tip, sides, and edges of the fingertip. Additionally, an omnidirectional surface using a single camera and lens system can achieve a full field of view of over 200 degrees. Another technical advantage of these embodiments may include high temporal resolution based on an interchangeable and modular electronic stack for touch digitization (visual acquisition, multimodal acquisition, on-device AI processing), which combines all electronic components and sensing elements into a size and shape familiar to the human thumb. Each individual system of the stack can be interchanged with newer designs to shorten the hardware development lifecycle, and additionally, components in the stack can be excluded to further reduce the size of the sensor based on the individual requirements of the application. Another technical advantage of these embodiments may include digitized touch signals for on-device AI processing to provide human-like reflex arc actions to a robotic manipulator, as a neural network processor for direct inference of input data is located within the fingertip and the processor can provide a direct output to control auxiliary devices such as a robotic end effector. Certain embodiments disclosed herein may not provide the above technical advantages, provide some of the above technical advantages, or provide all of the above technical advantages. Given the figures, specification, and claims of the present disclosure, one or more other technical advantages may be apparent to those skilled in the art.

[0010] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited to these embodiments. A particular embodiment may include all components, elements, features, functions, operations, or steps of the embodiments disclosed herein, some components, elements, features, functions, operations, or steps, or no components, elements, features, functions, operations, or steps. Embodiments in accordance with the present invention are particularly disclosed in the appended claims directed to methods, storage media, systems, and computer programs, where any feature mentioned in one claim category (e.g., method) may also be claimed in another claim category (e.g., system). The dependency or citation relationships in the appended claims are merely chosen for formality reasons. However, any subject matter resulting from an intentional citation of any previous claim (in particular multiple dependent claims) may also be claimed, such that any combination of the claims and their features is disclosed and may be claimed regardless of the dependency relationships chosen in the appended claims. The subject matter that may be claimed includes not only combinations of the features set forth in the appended claims, but also any other combinations of the features in the claims, where each feature mentioned in the claims may be combined with any other feature in the claims, or a combination of other features. Additionally, any embodiment and feature described or depicted herein may be claimed in separate claims, and / or any combination of any embodiment and feature described or depicted herein with any embodiment or feature described or depicted herein, or any combination with any feature in the appended claims may be claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 FIG. shows an exemplary cross-sectional view of an artificial fingertip and its modular electronics disclosed herein.

[0012] Figure 2A FIG. shows an exemplary evaluation of normal and shear forces in three separate regions of a sensor from the tip towards the side.

[0013] Figure 2B FIG. shows an exemplary prediction of normal and shear forces based on a visual-tactile image output.

[0014] Figure 2C FIG. shows an exemplary evaluation of spatial resolution by indenting an artificial fingertip of different widths using a dual-pronged microindenter.

[0015] Figure 2D FIG. shows an exemplary method of creating the volume of an artificial fingertip.

[0016] Figure 2EAn example simulation showing the effect of increasing surface scattering along the internal reflection layer of an artificial fingertip.

[0017] Figure 3A The ability to determine the volume of water in a container by tapping the opaque container with a finger and recording the response on the finger in contact with the container is shown.

[0018] Figure 3B An example spectrogram of surface audio texture recorded for different objects is shown.

[0019] Figure 3C An example sensitivity of an artificial fingertip to a thermal gradient is shown.

[0020] Figure 3D An example accuracy of object recognition by local gas sensing is shown.

[0021] Figure 3E An example accuracy of object classification by odor is shown.

[0022] Figure 3F An example positioning of a finger on an object during a moving transient in the case of an empty volume liquid and a full volume liquid is shown.

[0023] Figure 4 An example touch sensitivity surface mapping plot for Ec = 5.0 MPa and Eg = 3.0 MPa is shown.

[0024] Figure 5 An example visualization of finite element analysis (FEM) results using a surface mapping plot is shown.

[0025] Figure 6 An example maximum normal force and shear force applied to the surface of an artificial fingertip before separation from the body are shown.

[0026] Figure 7 An example cross-section of a fingertip gel coating is shown, showing 3 layers: an outer layer, a silver layer, and a base gel layer.

[0027] Figures 8A to 8H An example hyperfisheye lens designed specifically for acquiring omnidirectional tactile sensing images is shown.

[0028] Figure 9 An example inference delay measurement result comparing a device with a host, incorporating a tactile data preprocessing stage and a transmission stage, is shown.

[0029] Figure 10 An example evaluation of the non-uniformity index is shown.

[0030] Figure 11Shows an example data acquisition pipeline for the vision system of the disclosed artificial fingertip and touch information from external stimuli.

[0031] Figure 12 Shows example data collection and prediction.

[0032] Figure 13 Shows the example simulated performance of the vision system of the artificial fingertip disclosed herein from on-axis contact to far-field contact to increase the number of line pairs per millimeter of spatial resolution converted to sagittal response and tangential response.

[0033] Figure 14 Shows the example average pipeline latency of the MobileNetV2 network from acquiring tactile data, transmitting data, preprocessing, inference to providing actions for different image and channel width sizes.

[0034] Figure 15 Shows an example comparison of touch information from the artificial fingertip disclosed herein tapping a water bottle in which the volume of water varies from empty, half-full, and full.

[0035] Figure 16 Shows an example 6-DoF robotic indentor for testing the force resolution of tactile sensors.

[0036] Figures 17A to 17B Shows example image snapshots taken from the shear force data collection at two critical moments of the artificial fingertip disclosed herein.

[0037] Figure 18 Shows the example normal force prediction error distribution by surface type and region.

[0038] Figure 19 Shows an example acquisition in a vision-based touch sensor.

[0039] Figure 20A Shows an example simulation of the human reflex arc that rapidly processes sensing inputs within the fingertip and directly controls the actuators of the robotic hand to retract in response to a touched object.

[0040] Figure 20B Shows an example tactile processing and control paradigm for transmitting sensing data to a remote computer for processing.

[0041] Figure 20C Shows an example local processing for simulating the reflex arc with a system fingertip.

[0042] Figure 20D Shows the example mean and standard deviation of the event-to-action delay. Detailed Description

[0043] Touch is a sensing modality that can provide rich information about object properties and interactions with the physical environment. Both humans and robots benefit from using touch to sense and interact with their surroundings. However, no existing system can provide rich multimodal digital touch sensing capabilities while preserving the shape elements of a human finger. Embodiments disclosed herein can improve the digitization of touch through technological advancements embodied in artificial finger-shaped sensors with enhanced sensing capabilities. In a particular embodiment, an artificial fingertip can include high-resolution sensors (e.g., approximately 8.3 million contact pixels) that respond to omnidirectional touch, acquire multimodal signals, and use on-device artificial intelligence to process data in real time. For example, in a particular embodiment, evaluations show that an artificial fingertip can resolve spatial features as small as 7 μm, sense normal and shear forces with resolutions between, e.g., 1 mN and 1.3 mN respectively, perceive vibrations up to, e.g., 9 kHz to 11 kHz, sense odors, and even sense heat. Additionally, an on-device AI neural network accelerator can act as a peripheral nervous system on a robot and mimic a human reflex arc. These results indicate that embodiments disclosed herein can digitize touch with enhanced performance. Embodiments disclosed herein can be applied to fields including robotics (industrial, medical, agricultural, and consumer levels), virtual reality and telepresence, prosthetics, and e-commerce. Although this disclosure describes digitizing a particular modality in a particular way, this disclosure contemplates digitizing any suitable modality in any suitable way.

[0044] In a particular embodiment, a system for touch digitization can include a silicone hemispherical dome that includes a surface containing a reflective silver film layer. The system can additionally include an omnidirectional optical system that includes: a lens that includes a plurality of lens elements; and an image sensor that is configured to generate image data based on data acquired by the lens. In a particular embodiment, a first lens element of the plurality of lens elements can be in direct contact with the silicone hemispherical dome without an air gap. The lens can be configured to acquire scattering of internally incident light generated by the reflective silver film layer. The system can further include one or more non-image sensors disposed below the omnidirectional optical system. The system can also include one or more processors and a non-transitory memory coupled to the processors, the non-transitory memory including instructions executable by the processors. In a particular embodiment, the processors can be operative, when executing the instructions, to: access image data from the omnidirectional optical system and sensing data from the one or more non-image sensors, and generate the touch digitization based on the accessed image data and sensing data through one or more machine learning models.

[0045] In certain embodiments, an artificial fingertip for touch digitization may include a silicone hemispherical dome. The artificial fingertip may further include an omnidirectional optical system that includes: a lens including a plurality of lens elements; and an image sensor configured to generate image data based on data collected by the lens. In certain embodiments, a first lens element of the plurality of lens elements may be in direct contact with the silicone hemispherical dome without an air gap. The artificial fingertip may further include one or more non-image sensors disposed below the omnidirectional optical system.

[0046] Of all human senses, touch is perhaps the most crucial in how humans interact with the world. Touch enables humans to measure forces and identify object properties such as shape, weight, density, texture, friction, elasticity. Touch also plays an important role in social relations and cognitive development. In contrast, the earliest efforts to endow robots with this sense have been only a rough approximation. There is currently no solution to digitize touch with the rich sensory spectrum that humans take for granted. With the development of robotic hand manipulation, the embodiments disclosed herein can mimic the familiar features of the human hand: fingers. Touch digitization in the embodiments disclosed herein can enable intelligent systems to discern significantly higher levels of physical information during environmental interaction.

[0047] Digitizing touch can depend on two characteristics, which include an actual temporal characteristic and a spatial characteristic. The temporal characteristic can process information from a basic signal that varies over time, while the spatial characteristic can process a discrete multi-dimensional array of temporal signals. The embodiments disclosed herein combine these methods within a unified platform to improve the ability to digitize touch, namely, a modular, finger-shaped, multimodal tactile sensor with on-device artificial intelligence (AI) capabilities and superhuman performance.

[0048] In certain embodiments, an artificial fingertip may include a multimodal modular sensor. The multimodal modular sensor may include a body mechanical housing. By way of example and not limitation, the body mechanical housing may have a shape and size similar to that of a human thumb. In certain embodiments, a silicone hemispherical dome, an omnidirectional optical system, one or more non-image sensors, one or more processors, and non-transitory memory may be disposed within the housing. The multimodal modular sensor may further include a soft silicone solid body fingertip. By way of example and not limitation, the soft silicone may be composed of polydimethylsiloxane (PDMS) material. In other words, the silicone hemispherical dome may be based on polydimethylsiloxane (PDMS) material. The solid body fingertip may include a chemically grown metallic silver reflective layer and an outer layer that protects the reflective layer. The multimodal modular sensor may additionally include multimodal electronics. The multimodal electronics may go beyond the traditional limitations of vision-based modalities, which may be limited by the acquisition rate of CMOS. Instead, the optical systems disclosed herein may acquire at variable frames per second (e.g., 240 fps, depending on the capabilities of the CMOS), while the multimodal inclusion may allow for the acquisition of time data up to 10,000 Hz. This is a system-level contribution using off-the-shelf components, but assembled in a way that allows for the direct sampling or acquisition of signals generated due to input stimuli on the fingertip surface. In certain embodiments, the multimodal electronics may be realized through customized molding and casting techniques by placing the electronics directly into a mold, injecting silicone, and placing the mold under vacuum to ensure that the liquid silicone can fill all the orifices and gaps around the off-the-shelf sensors.

[0049] In certain embodiments, one or more non-image sensors may include one or more of an inertial measurement unit (IMU) sensor, a microphone, an environmental sensor, a gas sensor, a pressure sensor, or a temperature sensor. The multimodal modular sensor may include an inertial measurement unit (IMU) located in the main housing to process electronics. The IMU may be firmly coupled to the enclosure and the fingertip. The IMU may be used to sense vibrations, rotations, and positions of the fingertip. The multimodal modular sensor may also include a MEMS-based microphone in a multimodal sensing printed circuit board (PCB) that may be directly molded onto a soft silicone solid fingertip. In one example embodiment, there may be two microphones with top apertures and two microphones with bottom apertures. The two microphones may be digital microphones, and the two microphones may be analog microphones, with different sensing frequency ranges to cover a larger bandwidth. These microphones may provide surface audio textures similar to those perceived by a human fingertip when scratching an object surface, sampling object-object interactions, or object-environment interactions.

[0050] In certain embodiments, the multimodal modular sensor may additionally include an environmental sensor located in the main housing. An internal air flow may be provided by a micro air intake fan that generates an air flow through the device. The environmental sensor is capable of sampling local air for air pressure, temperature, moisture, and humidity. The multimodal modular sensor may also include a gas sensor located within the main housing. An internal air flow may be provided by a micro air intake fan that generates an air flow through the device. The gas sensor may identify chemical compounds and constituents of a local air sample to provide object state information and for classifying different objects. The multimodal modular sensor may also include a pressure / temperature sensor in the multimodal sensing PCB that may be directly molded onto a soft silicone solid fingertip. The pressure / temperature sensor may measure the absolute value of the compressive force applied to the fingertip and capture the thermal gradient caused by a heat source on the fingertip and the heat flow due to a metal reflective layer.

[0051] In certain embodiments, the multimodal modular sensor may include processing electronics. The system may also include a stack that includes multiple printed circuit boards for one or more processors, an omnidirectional optical system, and a data transfer system, where the multiple printed circuit boards share a common electrical interface and a connector stack. In other words, the processing electronics may be located in a stack that combines five independent and distinct PCBs sharing a common electrical interface and a connector stack in a limited space. The processing electronics may include a microprocessor, a neural network accelerator, an image acquisition system, and a data transfer system.

[0052] In certain embodiments, one or more processors may include one or more of a microprocessor or an accelerator. The microprocessor may be responsible for obtaining all data from sensors, lightweight processing, and configuration of processing and sampling parameters. In certain embodiments, one or more machine learning models may include one or more neural network models. Accordingly, one or more processors may include one or more neural network accelerators configured to accelerate real-time inference on accessed image data and sensed data through one or more neural network models. The neural network accelerator may have a direct path to the microprocessor to obtain an alternative data stream, whether the data stream is image data or non-image data. The neural network accelerator may provide neural network acceleration that may be uploaded based on offline training to perform real-time inference on the input data stream. The neural network accelerator may also have a direct path to provide an output control signal to an auxiliary device. In one example embodiment, the auxiliary device may include a robotic end effector. By way of example and not limitation, the auxiliary device may be a single finger on a robotic hand, which may enable rapid response of the finger based on the input data. The image acquisition system may be responsible for acquiring image frames from a complementary metal-oxide semiconductor (CMOS) and configuring sampling parameters for different frames per second and resolutions. The data transfer system may provide a USB interface to a host computer that combines all multimodal information and image / visual information on a single connector and also powers the device.

[0053] The embodiments disclosed herein may be used in a variety of applications. By way of example and not limitation, these applications may include medical applications (such as prosthetics, palpation, sensing and localization (e.g., breast cancer, testicular cancer, etc.), remote surgery, etc.). As another example and not limitation, these applications may include advertising (such as e-commerce, acquiring the feel of fabric, understanding the feel of skin and application of cosmetics, etc.). As yet another example and not limitation, these applications may include agriculture (such as fruit / vegetable picking, food quality determination, food processing, etc.). As yet another example and not limitation, these applications may include haptics (such as accurately acquiring physical words using an artificial fingertip).

[0054] Figure 1Shows an example cross-sectional view 100 of the artificial fingertip and its modular electronics disclosed herein. Traditional efforts in the field may involve subsets of sensing modalities optimized for cost, iteration of fingertip geometry, or maximizing a particular design metric. However, these choices may have limitations in performance. The primary sensing modality of vision-based tactile sensors can acquire the geometry of the object being touched. Based on the geometric data, the normal force and shear force can be reconstructed. However, this particular modality may not conform to the multimodal nature of human skin, as human skin uses many different types of receptors (e.g., mechanoreceptors, thermoreceptors, and nociceptors). Compared with traditional efforts, the disclosed artificial fingertip (as Figure 1 shown) may include 5 printed circuit boards located directly within the fingertip. This is a self - contained design. All of these printed circuit boards are custom - designed. From top to bottom, the disclosed artificial fingertip includes a multimodal sensor system, CMOS, an image acquisition system, a processing and neural network accelerator, and a data transfer system.

[0055] When deploying a high - end modular research platform for studying touch and its digitalization, the embodiments disclosed herein introduce new methods. The platform disclosed herein (identified as the artificial fingertip) may belong to the family of vision - based tactile sensors. Figure 2A Shows an example evaluation of the normal force and shear force in three separate regions of the sensor from the tip towards the side. The dots and error bars show the median and 95th percentile of the error, respectively. For example, in a particular embodiment, the median error from the deep - learning model for the three regions is 1.01 mN, 1.09 mN, 1.41 mN for the normal force and 1.27 mN, 1.48 mN, 1.64 mN for the shear force. Figure 2B Shows an example prediction of the normal force and shear force based on the vision - tactile image output. A particular embodiment may train a deep - learning model to output predicted normal force and shear force based on the vision - tactile image. For example, in a particular embodiment, the normal force and shear force may be predicted to have median errors of 1.01 mN and 1.27 mN, respectively. Compared with traditional methods, predicting the shear force may require the use of labels. However, as the spatial resolution increases, more features can be extracted from the vision - tactile image, which helps in shear - force prediction. Figure 2C Shows an example evaluation of the spatial resolution by pressing an artificial fingertip with different widths using a two - pronged micro - indentor. Visual verification and examination of the profile intensity of tactile pixels confirm the ability to clearly distinguish features as small as, for example, 7 μm. Figure 2D Shows an example method of creating the artificial fingertip volume. The top row corresponds to the internal - structure - based method, while the bottom row corresponds to the solid - gel - based method with an immersion lens. Through the internal structure, illumination artifacts are visible, while through the solid volume, the quality of the resulting image is much higher and there are fewer illumination artifacts.Figure 2E An example simulation showing the effect of increasing surface scattering in the internal reflection layer along the artificial fingertip is presented. From left to right, machine polishing to 1° Lambertian scattering, the embodiments disclosed herein optimize image contrast while constraining the background illumination uniformity.

[0056] The platform disclosed herein can have an elastomer serving as a touch sensing interface and a subcutaneous camera that measures the deformation of the elastomer through structured light. However, in addition to this capability, the platform disclosed herein is capable of sensing multiple tactile modalities (see Figure 2B ), including contact intensity and geometry, static and dynamic forces, surface audio texture and vibration, thermal variations, fingertip speed / acceleration / orientation, and even identifying some airborne compounds. Additionally, to simulate the human reflex arc, the embodiments disclosed herein introduce an on-device AI neural network accelerator that provides next-generation touch processing capabilities, enabling real-time local processing to minimize reaction latency and reduce communication bandwidth.

[0057] To effectively capture the nuances in touch interactions with the world, the artificial fingertips disclosed herein can be sensitive in both the time domain and the spatial domain. These domains can be obtained through vision, audio, vibration, pressure, thermal, and gas sensing as modal signals incorporated within the artificial fingertips. Typically, traditional vision-based touch sensors using off-the-shelf imaging systems may be limited by a slow visual acquisition rate, which reduces the amount of sequential information generated from frame encoding times, thereby limiting the temporal characteristics of non-static touch interactions encountered during manipulation. Increasing the temporal frequency of the visual system may not be beneficial without an increase in the spatial resolution of dynamic motion. The embodiments disclosed herein incorporate a touch digitization system in the form of an artificial fingertip having a geometry similar to that of a human finger. The surface of the fingertip can encode touch information from the depressions in the reflection layer, and the camera captures the internal light reflections of the reflection layer. A visuotactile system can be utilized to resolve the smallest possible spatial features presented by an object interacting with the fingertip at a high temporal rate. The artificial fingertips disclosed herein can advance spatial, temporal, and multimodal performance through embodiments of a modular platform for studying touch digitization. To achieve the high spatial and temporal performance demonstrated by the artificial fingertips disclosed herein, the embodiments disclosed herein utilize methodological breakthroughs in five different subsystems (elastomer interface, optical system, illumination system, multimodal sensing, and on-device AI processing).

[0058] When the reflective fingertip surface layer is subjected to impression stimuli, the material properties of the surface layer may be directly related to the spatial resolution ability. The embodiments disclosed herein develop experimental design techniques to identify these six material parameters that affect the sensitivity of the sensor to input stimuli: Rg (fingertip radius), Tc (surface reflection coating thickness), Tg (surface layer thickness), h (height), Ec (coating Young's modulus), and Eg (fingertip bulk Young's modulus). If Tc and Tg are too thick or exhibit low compliance, the low-pass filtering effect may be evident at the edges of discrete objects. Similarly, if the object has rich spatial information and a fractal dimension, the fingertip surface may resolve fewer features due to local gradients in material compliance. Specific embodiments may avoid specifying any constraints on the coating thickness layer to best capture small input stimuli while maintaining a suitable parameter range for the general size of the fingertip. Forming this layer on the fingertip surface may involve hand painting, spray painting, or dip coating techniques. However, when generating touch images, they may be far from optimal and result in a large coating thickness and inconsistent yield due to manufacturing variations. The embodiments disclosed herein address this problem by developing a new chemical deposition technique for directly growing a silver thin film onto the surface of the fingertip, resulting in a much smaller coating thickness than previous methods and thus achieving better sensitivity.

[0059] Common visual tactile sensors can acquire input stimuli at a planar surface, using multiple cameras that are difficult to integrate and process together, or by default using common off-the-shelf cameras optimized for human-centric imaging, which results in degraded optical performance for touch. Specific embodiments can avoid using standard image sensor features (such as automatic exposure control, automatic white balance, and autofocus), which are designed to respond to changes in the natural environment, since the fingertip chamber disclosed herein can be a closed and controlled environment. To model an isotropic representation of a size similar to that of a human fingertip, a new approach may be needed to optimize the acquisition of a hemispherical surface. When optimizing input stimuli from the touch interaction layer, the imaging system disclosed herein can perform within the finite element method simulation without restricting material properties. Thus, for example, specific embodiments can determine the optical system requirements to best suit the acquisition of images related to tactile sensing with a CMOS pixel size of 1.1 μm. Parameters can be selected for the converging spot size to improve spatial resolution, intentionally allow chromatic aberration, introduce a shallow depth of field to allow defocus proportional to the object indentation depth, and remove the anti-reflection coating to allow the acquisition and interpretation of reflections and scattering within the fingertip. However, such parameters may require non-standard lenses. Thus, specific embodiments can utilize a custom solid immersion hyperfisheye lens to handle the unique environment of visual tactile sensing, rather than using off-the-shelf lenses catering to general imaging, thereby enabling full control over the lens geometry and optical parameters.

[0060] The embodiments disclosed herein describe two metric parameters for illumination performance within the volume, background uniformity, the degree of uniformity of the light distribution, the uniformity contrast of the image with the background, and the prominence of the imprint on the fingertip surface compared to the background, as follows. A common approach may be an embodiment of an internal structure that serves as a hemispherical light pipe and provides fingertip rigidity. However, the internal light pipe structure may produce illumination artifacts in the form of flicker and hot spots from the convex geometry, which results in the degradation of the image metrics. To reduce these artifacts, traditional methods can utilize a textured surface to induce Lambertian scattering of the incident light. Specific embodiments can simulate the reflective layer surface properties with a controlled degree of scattering from polished to Lambertian, where the entire hemispherical surface can act as an integrating sphere, showing that a Lambertian scattering surface may not be the best approach to achieve high performance. Specific embodiments can combine controlled reflection surface scattering parameters away from the Lambertian surface, using a rigid solid volume instead of the more common hollow volume or using an internal support structure.

[0061] The embodiments disclosed herein disclose a new platform and demonstrate that these advancements are far superior to traditional visual tactile technologies in terms of spatial sensitivity and force sensitivity. While visual information can provide insights into the environment and object contact (such as texture and surface deformation), this may only provide a subset of the fingertip-to-object-environment understanding.

[0062] Figure 3A Shows the ability to determine the volume of water in a container by tapping an opaque container with a finger and recording the response of the finger in contact with the container. The embodiments disclosed herein also show how this modality can be deconstructed into peak frequency analysis, which can be independent of finger position, while the decay time can depend on finger position. Figure 3B Shows an example spectrogram of surface audio textures recorded for different objects. Figure 3C Shows an example sensitivity of an artificial fingertip to a thermal gradient. Using a variable heat source as a control, the artificial fingertip can be sensitive to a thermal gradient. Figure 3D Shows an example accuracy of object recognition through local gas sensing. The artificial fingertip can provide object status and object recognition through local gas sensing, achieving an accuracy of 91%. Figure 3E Shows an example accuracy of object classification by odor. The accuracy of classifying an object by odor can depend on the integration time when the artificial fingertip starts approaching the object. Figure 2E Shows an accuracy of 61% achieved within 6 seconds. Figure 3F Shows an example positioning of a finger on an object during a moving transient in the presence of a liquid with an empty volume and a full volume. The embodiments disclosed herein measure the transient effect during a pulse, the resonance during a static hold, and the static hold without movement.

[0063] Certain embodiments can further develop the capabilities of the platform to include sensitivity to non-visual-based modalities. By way of example and not limitation, when in contact with the environment, dynamic forces and signals may be experienced; sliding a fingertip on a surface, or the instant when a contact transient or slide may occur. Certain embodiments can acquire this information through an audio microphone within the fingertip and a pressure-based MEMS sensor, and show the ability to determine the liquid level inside an opaque bottle (see Figure 2A ) and understand the subtle differences in surface textures between different objects (see Figure 2B ) at frequencies much higher than visual acquisition (e.g., 240 Hz), up to, for example, 10 kHz. Certain embodiments can also include modalities for understanding the state of an object, which is not necessarily a function of contact, but can provide a priori knowledge of touch through heat and odor. Such a priori knowledge can estimate whether an object is slippery due to the presence of water, soap, or butter, and whether the object is dangerous to human contact due to its temperature.

[0064] For the arc reflection of humans, a rapid response to input stimuli on the fingertips can benefit from the central nervous system rather than traveling back and forth to the brain. Specific embodiments can utilize a similar local processing response on artificial fingertips. By way of example and not limitation, specific embodiments can include a neural network accelerator within the shape elements of the fingertips to process sensing readings and allow direct control to provide actions to a robotic end effector to control the phalanges of a robotic finger. While this may be a new era of on-device fingertip processing, the embodiments disclosed herein disclose two main effects that contribute to faster response, namely latency and jitter. Latency may be caused by the average time required to process the signal of interest, and jitter may be the variation of the average time based on system overhead, which may occur due to host processing or bandwidth limitations. Compared with the traditional method of using an artificial fingertip with an external host that requires a 2x round-trip latency to perform an action, the embodiments disclosed herein show that, in an example embodiment, by directly processing data on the device through an on-board neural network accelerator, the latency and jitter of performing an action can be reduced to one-half.

[0065] In a specific embodiment, a 3D finite element method (FEM) model using Comsol Multiphysics can be utilized to analyze and characterize the fingertip material stack. The 3D FEM model can identify the sensitivity and resolution of the sensor. First, the FEM model can be used to determine the key parameters that exhibit the greatest changes in terms of sensitivity and resolution. Since the fingertip can be isotropic and can rotate around the origin, using a multi-layer based model can model only one-quarter of the sensor to achieve faster calculations. The multi-layer model can include a base gel layer, a polymer layer, and a coating.

[0066] Specific embodiments can use a specific system to generate a nanomechanical characterization of the Young's modulus E of the fingertip polymer. The system can have in-situ high-resolution imaging, dynamic nanoindentation, and a high-precision motion stage with a high-resolution force sensing tip. By way of example and not limitation, to perform the characterization, a force of 30 μN can be applied with a probe tip of 10 μm. The corresponding force-displacement curve can be measured, thereby generating a Young's modulus of, for example, E = 2.86 MPa. Using the experimental E value, the FEM model can be updated to correct the simulation. In addition, the same applied force value and the maximum displacement Dmax can be measured to verify the simulation, taking Dmax = 22.1 μm as an example and the experimental Dmax = 2.2 μm measurement result as an example, with an example error ≤ 5%. In addition, multiple measurements can be made on different samples of the fingertip. For example, the average value of E can be measured at E = 2.6 ± 0.74 MPa. Both Emean and Estd can be used for a detailed analysis of the total E range in the FEM model.

[0067] Certain embodiments may also employ design of experiments techniques for identifying key parameters that affect sensor sensitivity. By way of example and not limitation, six different parameters may be used: Rgel (gel radius), Tc (coating thickness), Tg (gel layer thickness), h (height), Ec (coating Young's modulus), Eg (gel Young's modulus). In certain embodiments, the material associated with the silicone hemispherical dome may be determined based on multiple material parameters, including one or more of the gel radius, coating thickness, gel layer thickness, height, coating Young's modulus, or gel Young's modulus. A full factorial design of the 6 parameters results in 64 models, and thus, a 1 / 4 factorial design method may be used to reduce the design to 16 models. Analysis of variance and predictive analysis may result in Ec and Eg as the main effects and interacting with the coating and gel thicknesses. Accordingly, the parameters height h and gel radius Rgel may be removed from the model. To analyze the effect of gel thickness Tg and coating thickness Tc on sensor performance, a design of experiments method may be used by sweeping the Young's modulus parameters Ec, Eg and the thickness parameters Tg and Tc. For example, for the Young's modulus of the protective fingertip layer, values of Ec, Eg = 0.5 MPa, 1.0 MPa, 3.0 MPa, 5.0 MPa may be used. As another example, for values of the coating thickness, Tc = 0.1 mm, 0.5 mm, 1.0 mm, 2.0 mm, 3.0 mm may be used, and for the gel thickness, Tg = 0.5 mm, 1.0 mm, 5.0 mm, 10 mm, 15 mm may be used. Figure 4 An example touch sensitivity surface mapping plot for Ec = 5.0 MPa and Eg = 3.0 MPa is shown. Figure 4 It is shown that the sensitivity increases as the gel thickness and the coating thickness decrease. Figure 5 An example visualization of the FEM analysis results using a surface mapping plot is shown. Figure 5 It shows the selection of regions of interest for the background and indentation (in the case of multi-contact indentation on the fingertip surface), where for each combination of the coating Young's modulus Ec and the gel Young's modulus Eg, the x-axis shows the gel thickness Tg values and the y-axis shows the coating thickness Tc. Figure 5It is shown that the Young's modulus of the material may significantly contribute to the minimum Tc and Tg required to produce high performance. For the exemplary conventional tactile sensor, a similar FEM analysis was performed. A multi-layer based model can be used. The embodiments disclosed herein analyzed Tg (gel thickness), Tc coating thickness, Ec (coating Young's modulus), and Eg (gel Young's modulus). For example, a parametric study was conducted where Tg was from 0.5 mm to 15 mm in steps of 0.25, Tc was from 0.1 mm to 3 mm, the coating Young's modulus Ec was 5 MPa, and the gel Young's modulus Eg was 1 MPa. The results of this study showed that the gel thickness and the coating thickness interacted, and the combination of the coating thickness and the gel thickness may affect the overall sensitivity of the sensor.

[0068] Certain embodiments can determine the average force and the maximum force that can be applied to the fingertip before the soft hemispherical surface is damaged. Certain embodiments can fix the sensor as a force-torque sensor and apply an increasing force until the fingertip separates from the body, which may occur, for example, when the normal force is 40 N and the shear force is 20 N. Figure 6 Shown are example maximum normal and shear forces applied to the surface of the artificial fingertip before separation from the body. Certain embodiments can apply incremental normal and shear forces to the body of the soft artificial fingertip and record the ground truth force data from the fixed force-torque sensor. When a sudden change occurs, this may indicate that the maximum force has been reached.

[0069] FEM simulations show that the Young's modulus may be an important parameter for the performance of the sensor in response to input stimuli and precise controller measurements may be required. In addition to nanoindentation measurements (point-based measurements), a set of dynamic mechanical thermal analysis (DMTA) measurements can be performed to obtain the overall Young's modulus of the gel. Using this method, certain embodiments can measure the viscoelastic properties of the polymer. During DMTA measurements, an oscillatory force can be applied to the material and its response recorded to calculate the viscosity and hardness of the material. Oscillatory stress and strain measurements can be important in determining the viscoelastic properties of the material.

[0070] When an oscillating force is applied, the values of sinusoidal stress and strain can be measured. The phase difference between the sinusoidal stress and strain can provide information about the viscous and elastic properties of the material. The phase angle of an ideal elastic system can be 0 °C, while that of a viscous system can be 90 °C. In addition, the elastic response of the material can be analogous to energy storage and can be obtained by the storage modulus, while the viscous response can be considered as energy loss and is obtained by the loss modulus. Therefore, the total modulus of a viscoelastic material can be a combination of the elastic and viscous components, in other words, the sum of the storage modulus and the loss modulus. Another value, tanδ, can be used to compare the viscous modulus and the elastic modulus. DMTA can measure the changes of the elastic modulus, loss modulus, and tanδ with respect to temperature. Since the viscosity of the material is affected by temperature and time, DMTA experiments can generally be carried out at different temperatures and frequencies. Specific embodiments can use common methods to select the operating conditions of the material. Therefore, for ideal sensitivity, it may be necessary to use fingertips at room temperature and low frequencies. Thus, in an exemplary embodiment, DMTA measurements can be carried out at 25 °C and a frequency of 5 Hz.

[0071] During the fabrication of the fingertips, different combinations of polymers with different shore values can be evaluated. For example, to determine the overall Young's modulus and influence of different gel mixtures, DMTA measurements can be carried out at 25 °C and a frequency of 5 Hz, as shown in Table 1. To optimize for higher sensitivity, fingertip materials with a lower Young's modulus can be preferred. Specific embodiments can choose to use, for example, the fabrication of a specific silicone encapsulating rubber with a ratio of 0.8:1 (Part A to Part B) of a gel fingertip base material, and a thin protective piece of another specific cured silicone rubber as the film layer.

[0072]

[0073]

[0074] Table 1: DMTA measurements on typical polymers used in fingertip gel fabrication.

[0075] In a particular embodiment, a silicone hemispherical dome can be generated by the following steps: fabricating a mold out of aluminum, finishing the mold with a machine polishing pass, preparing the mold for gel casting by a salinization process in a dryer, preparing a gel material using a curable silicone rubber compound, combining the gel material in a rapid mixer under vacuum, pouring the gel material into the mold, curing the poured gel material at a first temperature for a first amount of time, and once the poured gel material is cured, removing the gel hemispherical dome from the mold. By way of example and not limitation, a particular embodiment can fabricate a fingertip mold out of 6061 aluminum and finish the fingertip mold with a tool of 3 mm diameter and a machine polishing pass with a 50 μm step. Then, the mold for gel casting can be prepared by a silanization process using 50 μL of silane in a dryer under vacuum for 30 minutes. Subsequently, the gel material can be prepared using the above-mentioned silicone encapsulation rubber in a 1:1 ratio, and the gel material can be combined in a rapid mixer under vacuum for 3 minutes to release any trapped air in the sample. Then the gel material is poured into the mold and allowed to cure at 23 °C for 12 hours. Once cured, the gel fingertip can be removed from the mold using tweezers for transfer to a glass slide.

[0076] In certain embodiments, a reflective silver film layer can be generated based on the following steps: prepare a glucose solution by dissolving a first amount of glucose in a second amount of H2O and adding a third amount of KOH, prepare an AgNO3 solution by dissolving a fourth amount of AgNO3 in a fifth amount of H2O and adding a sixth amount of NH3, prepare a plating solution by mixing the glucose solution and the AgNO3 solution, clean the gel hemispherical dome with oxygen plasma for a second amount of time, activate the gel hemispherical dome in a solution of a seventh amount of SnCl2 in an eighth amount of H2O for a third amount of time, suspend the gel hemispherical dome in the plating solution for a fourth amount of time, rinse the gel hemispherical dome with H2O, and air-dry the gel hemispherical dome. By way of example and not limitation, the steps for preparing a thin film metal reflective layer on a gel fingertip by silver plating can be as follows. First, a glucose solution can be prepared by dissolving 2.035 g of glucose in 160 mL of H2O and then adding 0.224 g of KOH. The glucose solution can be set aside. An AgNO3 solution can be prepared by dissolving 1.02 g of AgNO3 in 120 mL of H2O and then adding 1.2 g of 25% NH3. Then, 2 parts of the glucose solution (80 mL total) can be mixed with 1 part of the AgNO3 solution (40 mL total) to prepare a plating solution for silver plating the gel fingertip. The silver plating solution can then be set to gently stir. Before coating with silver, the gel fingertip can be cleaned with oxygen plasma for 3 minutes. Then, the gel fingertip can be activated in a solution of 6.181 g of SnCl2 in 98 mL of H2O for 10 seconds. Once the gel fingertip is activated, the gel fingertip can be suspended in the silver plating solution for 3 minutes, then rinsed with H2O and air-dried. This process can form a silver-plated reflective layer with a thickness of 6 μm. For robotic applications and to increase resilience to ambient light intrusion, certain embodiments can coat the silver-plated layer in a white layer or a black layer. This layer can be produced by using the above-cured silicone rubber (mixing ratio of part A to part B is 1:1), and then adding 3% silicone pigment to part A of the above-cured silicone rubber. Then, part B of the cured silicone rubber is mixed by weight according to the previously specified mixing ratio and then mixed in a rapid mixer under vacuum for 3 minutes. Then, the silver-plated gel fingertip is immersed in the pigmented cured silicone rubber and set to cure for 6 hours.

[0077] Common vision-based tactile sensors can use static illumination configurations, while some other sensors can use a single light color and colored acrylic to simulate multiple colors. Static illumination is not ideal for promoting modular systems. Instead, the illumination system should be adapted to the need to extract information from the touch surface. Some traditional tactile sensors use gels coated with a Lambertian scattering layer, where volume illumination can produce an image by light scattering from the surface into the vision system. In the case of using a monolithic hemispherical gel dome in the artificial fingertips disclosed herein, the embodiments disclosed herein determine that Lambertian scattering may not be ideal for generating and optimizing force and spatial sensitivity. Additionally, the embodiments disclosed herein introduce a dynamic illumination system that provides volume illumination with configurable wavelength, intensity, and positioning. By way of example and not limitation, the illumination system can include 8 fully controllable RGB LEDs that emit Lambertian diffused light, which are equally spaced around a circle with a radius of 9 mm.

[0078] Figure 7 An example cross-section of the fingertip gel coating is shown, showing 3 layers: an outer layer, a silver layer, and a base gel layer. Figures 8A to 8H An example hyper-fisheye lens designed for collecting omnidirectional tactile sensing images is shown. Figure 8A An example fingertip area of the hyper-fisheye lens is shown. For the optical system performance, three regions of interest can be selected: the tip, the prominent contact surface, and the base. The fingertip gel in the system disclosed herein can include three components, as Figure 7 shown in FIGS. 7 to 8. The outer surface of the base gel can have a reflective silver thin film coating, which can be coated with a protective colored diffusive material. To produce an image, these two layers can scatter the internal incident light into the vision system according to the surface interaction. Specific embodiments can use an overmolding process to place the light-emitting diodes in optical contact with the gel. The fingertip gel can initially be manufactured to have a smooth and polished surface, and by texturing with a mold, how light scatters at the interface can be controlled and determined. Figure 8B An example lens of the hyper-fisheye lens is shown. Figure 8C An example lens cross-section view of the hyper-fisheye lens is shown.

[0079] Figure 8DAn example system layout of a solid immersion super fisheye lens system is shown. For example, the design configuration includes a silicone super hemispherical dome with a diameter of 25 mm and is equipped with a 5P lens system. This combination of silicone lens systems images object points located at or near the dome surface onto the sensor plane. For example, the EFL of the imaging system is 1.57 mm, the FOV is 194.5 degrees, and the focal ratio under the conjugates used is F / 3.68. The following slides show the image quality evaluation performance of these designs at l = 587.6 nm in terms of MTF, spot size, relative illumination, and wavefront analysis. Figure 8D The size / dimension of the lens system is also shown. Figure 8E An example lens layout and the plastic materials of the lens elements are shown. The lens system can use EP10000 plastic material for the L1 lens element, while the lens elements L2, L3, L4, and L5 are designed using APEL 5014 plastic material.

[0080] Figure 8F An example lens layout and the minimum clear aperture radius are shown. In a particular embodiment, the first lens element can be configured to have a maximum clear aperture diameter of 10 mm. As Figure 8F shown, the lens element L1 can have a maximum clear aperture diameter of 10 mm and is in contact with the super hemispherical silicone dome element (no air space). Figure 8F Additional example clear aperture data (based on the radius in millimeters) is also shown as follows: CIR S4 4.748376, CIR S5 1.707403, CIR S6 1.341748, CIR S7 0.538351, CIR S8 0.151529, CIR S9 0.574821, CIR S10 0.794950, CIR S11 0.862950, CIR S12 0.888102, CIR S13 0.897415, CIR S14 0.866286, CIR S15 0.842312, and CIR S16 0.808351. Figure 8G An example solid model of a silicone plus lens is shown. Figure 8H An example solid model of a lens is shown.

[0081] Using a Gaussian scattering distribution, a particular embodiment can model a range of scattering. The scattering parameter σ can be selected to achieve the half-width at half maximum angle α (e.g., α = 1° to 25°) of the bi-directional scatter distribution (BSDF) function at normal incidence and a Lambertian scatter model. Figure 9Shows example inference latency measurements comparing a device that combines a haptic data preprocessing stage and a transmission stage with a host. Figure 9 Indicates that using a CNN accelerator on the device improves the inference time and reduces the total latency to generate an action for an MLP deep neural network. Minimizing the scattered surface texture may produce little background illumination. It also does not produce large shadows caused by surface indentations. Additionally, minimizing scattering can produce specular glint reflections that do not uniformly illuminate all indentations and saturate the vision system.

[0082] Under a full Lambertian scattering model, the hemispherical surface of a gel fingertip can act as an integrating sphere. While Lambertian scattering provides uniform background illumination, the high-scattering illumination from nearby interactions may reduce the overall indentation contrast. The embodiments disclosed herein are optimized for high image contrast while maintaining uniform background illumination, thereby better imaging the imprints created by gel indentations and minimizing the amount of glint that would saturate the image sensor.

[0083] Two non-uniformity metrics are evaluated on the hemispherical surface of the fingertip: standard deviation / mean and (maximum - minimum) / mean. The embodiments disclosed herein show that low scattering produces images with a large variation in the image signal, requiring the camera to process a high dynamic range. If image saturation is allowed to address the variations due to indentations, the saturated regions of the image may be lost. Therefore, in these cases, stray light is more likely to cause unpleasant artifacts. Due to bright glint reflections, the contrast in the image caused by spherical indentations may be high in some regions; but regions with large gradients in the background may make the indentations difficult to detect. High scattering may produce images with a smaller variation in the image signal and does not lose any regions due to saturation. In the image caused by spherical indentations, the contrast in some regions may be low, but the uniform background may make these regions easier to detect. Figure 10 Shows an example evaluation of the non-uniformity metric. Minimizing the non-uniformity of the background can improve the detection of surface indentations as the scattering angle increases.

[0084] The embodiments disclosed herein define a contrast-to-noise (CNR) metric and study the background uniformity, noise, and indentation contrast of three regions of interest on the hemispherical surface. Plotting the calculated CNR on the hemispherical surface for different scattering angles, it can be observed that generally, the smaller the scattering, the higher the CNR, but the larger the scattering, the more uniform the CNR is across the FOV. Therefore, the embodiments disclosed herein determine that for a hemispherical fingertip surface, the desired texture scattering profile can be limited to, for example, between a half-width at half-maximum angle of 20° and 25°.

[0085] The embodiments disclosed herein include a platform based on modular principles to provide an omnidirectional vision-based multimodal sensing system and on-device AI capabilities. Specific embodiments can achieve modularity by isolating each part of the system into electronic components with small surface areas. Specifically, these modules can include a sensing subsystem, an optical subsystem, a vision acquisition subsystem, a processing subsystem, and a communication subsystem as shown in FIG. 2. By facilitating the reuse of each subsystem, the embodiments disclosed herein present a novel approach to advancing the technology in tactile sensor design - reducing the required design, iteration, and manufacturing processes. This can enable those skilled in the art to modify the entire system by adding, removing, and changing subsystems without having to design a bunch of new hardware to support changes to the system design. Additionally, this modular system architecture can allow for the selection of combinations of subsystems to facilitate the introduction of updated technologies. Feature removal can reduce the cost of adapting the system to new mechanical form factors. Additionally, this possibility of replacing the sensing fingertip can allow for adaptation to different environments and tasks, for example, by using different sensitivities and hardnesses in the fingertip material or by using fingertips with markers to introduce more prominent optical flow features to adapt to different environments and tasks. Based on simulations for achieving optimal spatial and force resolution, the embodiments disclosed herein introduce a novel process for designing and manufacturing sensing fingertips.

[0086] In a specific embodiment, the system disclosed herein can provide communication with a host device through a USB 3.0 standard interface. Three separate streams can be provided for data transfer, supporting video, audio, and multimodal data. For example, depending on the configuration sent to the device, these streams can jointly output at a maximum rate of 148 MBps or lower. In a specific embodiment, the tactile sensor can be used for open-loop control, providing information to the host device for processing and additional actions to the manipulator. In a specific embodiment, edge AI can be added to the fingertip for the following reasons. First, edge AI can help create a latent representation of the data and reduce the total bandwidth sent to the host device. Second, edge AI can help achieve fast local decision-making to convey actions to the manipulator. Third, edge AI can help improve the overall latency of the system while reducing the variation in jitter, i.e., the variation in latency. Specific embodiments can use a manipulator to model the tactile fingertip system, where both systems can be connected to the host device in a star configuration. This can lead to decisions - and actions resulting from tactile information - being processed by the host device and propagated to the manipulator. For more data-intensive designs where data is collected from multiple fingers at once, this arrangement can lead to an unstable control scheme where information and action latency cannot be guaranteed. To adapt to and expand the field of tactile sensing research, specific embodiments can integrate a specific neural network accelerator (e.g., a 9-core RISC-V computing cluster with AI acceleration) for on-device processing of selected data streams.

[0087] Following the high-level abstraction of the human reflex arc, the embodiments disclosed herein develop a fast reflex-like control loop that uses edge AI for local processing. The traditional paradigm of sending sensory inputs to a central control computer for processing and then sending back control signals may require high bandwidth while introducing communication latency. In contrast, the paradigm disclosed herein can locally process the sensory inputs inside the fingertip using edge AI. This can allow for a significant reduction in the required bandwidth while greatly minimizing communication latency and jitter. The embodiments disclosed herein perform an experimental comparison of these two paradigms by using a PCI-e based precise time measurement tool to measure the end-to-end latency of the system, as shown in Table 2. The embodiments disclosed herein evaluated this experiment on a Linux machine with 64GB of memory and a GPU. First, to ensure the granularity of the measurements, each section of the system was isolated and samples were collected in repeated trials. Second, the embodiments disclosed herein verified these results by taking repeated measurements of the entire system and comparing the timing results with the sum of the isolated components. This identified the regions that produce deterministic timing results and highlighted the regions where latency and jitter continuously increase. Additionally, these results indicate areas for performance improvement and design of future tactile sensors. The results of the entire control loop show how the edge AI paradigm reduces the latency from 4ms to 1ms with a desirable smaller variance. In a particular embodiment, appropriate edge AI processing can be extended to further utilize the sequential nature of the camera FIFO memory to parallelize data acquisition and processing, resulting in an even lower latency. In this case, instead of processing the entire image, selected horizontal lines can be sent for processing in the configured region of interest. This may be applicable when touch interactions are most likely to occur in certain regions on the omnidirectional fingertip disclosed herein. The system disclosed herein can support this region of interest data output selection to obtain higher resolution and image acquisition frequency.

[0088]

[0089] Table 2: Normal force prediction error (median) by surface type and region.

[0090] Figure 11Shows an example data acquisition pipeline 1100 for the vision system of the disclosed artificial fingertip and touch information from external stimuli. The external stimulus 1102 can be input to the exposure module 1104 of the image sensor 1106. The output of the exposure module 1104 can be input to the image buffer 1108 of the image sensor 1106. The output of the image buffer 1108 can be provided to the subsampling module 1110 of the artificial fingertip (on-device) 1112. Although this disclosure describes an image sensor for processing external stimuli as an example, this disclosure contemplates any suitable modality or any suitable subset of modalities that replace the image sensor.

[0091] The output of the subsampling module 1110 can be input to the SPI transfer module 1114. The output of the SPI transfer module 1114 can be input to the on-device inference module 1116 of the host inference module 1118. The output of the on-device inference module 1116 can be input to the finger motion transfer module 1120 based on I2C-to-finger transfer. The output of the finger motion transfer module 1120 can be used to generate finger motions 1122 for the fast hand 1124a to execute.

[0092] In a particular embodiment, the output of the image buffer 1108 can also be provided to the USB transfer module 1126 of the artificial fingertip (host) 1128. The output of the USB transfer module 1126 can be input to the subsampling module 1130 of the host inference module 1132. The output of the subsampling module 1130 can be input to the inference module 1134. The output of the inference module 1134 can be input to the finger motion transfer module 1136 based on USB-CAN. The output of the finger motion transfer module 1136 can pass through the palm processing module 1138 of the fast hand 1124b. The output of the palm processing module 1138 can be input to the I2C palm-to-finger transfer module 1140, and the output of the I2C palm-to-finger transfer module can be used to generate finger motions 1142 for the fast hand 1124b to execute.

[0093] Figure 12 Shows example data collection and prediction. The solid lines show the ground truth shear force and normal force trajectories during one indentation. The dashed scatter plots show the model-predicted shear force and normal force. The embodiments disclosed herein establish two pipelines for data transfer and processing: an on-device pipeline and a host pipeline, as Figure 12As shown. In addition, the hybrid mode can support data transfer to the host and data processing on the device. It can be observed that this affects the visual system because it may involve a much larger amount of data than a multimodal data system. The limiting parameter for dynamic operations in the task is the latency between the input data, processing, and operation commands. To elicit a quick response to environmental changes detected through changes in grasping stability, specific embodiments can be developed in the direction of faster processing and minimization of the latency between input and action.

[0094] The embodiments disclosed herein study the impact on the visual system because the visual system can be the most common modality used in touch sensors and has an impact on the overall system latency configured on the host and the device. The limiting factor for the system latency between the real-world input and the data processing input is the acquisition rate of the visual system. This may be limited by the frames per second (a latency of 1 / fps may be imposed) and the internal processing of the image signal processor. Thus, for example, a specific embodiment can incorporate a CMOS sensor with 240 fps and a pixel size of 1.1 μm in the system disclosed herein. This CMOS sensor can produce a shorter latency of 4.17 ms compared to a conventional sensor that can operate at 60 Hz and thus has a latency of no less than 16.7 ms.

[0095] The embodiments disclosed herein use two deep neural networks, MLP and MobileNetV2, to evaluate the inference latency in two scenarios: inference on the device and inference on the host. The two largest sources of latency are the transfer of tactile data from the device to the host and the transfer of action data from the host to the robotic end effector. For example, within the available margin between the differences in pipeline latency, Tlatency, an upper limit of Tlatency ≤ 2.463 ms is established. For the MLP-based network, the embodiments disclosed herein increase the layer depth and observe the latency cost in both scenarios. Table 3 shows that the operation latency for dynamic tasks involving high-speed movement is shortened to approximately one-quarter, thereby reducing the total action time to less than 1 ms.

[0096]

[0097] Table 3: Average data pipeline timing for processing on the host and the device.

[0098] Figure 13 An example simulation performance of the visual system of the artificial fingertip disclosed herein from on-axis contact to far-field contact is shown to increase the line pairs per millimeter for converting to sagittal response and tangential response. As Figure 13 shown, it becomes apparent that using the on-device accelerator without enabling the hardware engine will quickly exceed Tlatency at a depth of 10 layers. However, enabling hardware acceleration may allow the use of an MLP model with 60 layers.

[0099] Observing use cases more suitable for the field of haptic research, the embodiments disclosed herein deploy the MobileNetV2 model and determine the total system pipeline latency. Figure 14 Shows example average pipeline latencies of the MobileNetV2 network from acquiring haptic data, transmitting data, preprocessing, inference to providing actions for different image and channel width sizes. Figure 15 Shows an example comparison of touch information from the artificial fingertips disclosed herein, where the artificial fingertips tap water bottles with different volumes from empty, half-full, and full. Although the visual outputs of the sensors look almost the same except for the touch position changes, multimodal data can provide greater insights into object properties other than texture. Figure 15 Shows that using the MobileNetV2 architecture with an input size of 64×64 helps reduce Tlatency and is suitable for common underlying touch tasks such as touch detection and classification. Additionally, the upper limit of Tlatency can be determined by the output data rate, the size of data transmission, and the host system performance.

[0100] However, in a real robot environment, the host system may be running too many control and processing applications, where additional communication overheads between other sensors and devices introduce an overhead of, for example, 1.2 ms to Tlatency. Comparing this with traditional artificial fingertips, the embodiments disclosed herein observe an overhead of, for example, 4.7 ms. These differences can be attributed to the frame rate (e.g., 240 fps) of the artificial fingertips disclosed herein and data transmission via USB 3.0, while example conventional haptic sensors may be limited to 60 fps using USB 2.0. The systems disclosed herein can enable on-device inference for haptic devices with low-latency control, have the ability to perform class reflection control on the devices the system is connected to, and provide an abstraction of lower-level touch signals to the host. An example can be training a model on the device to regress force from multimodal data to introduce touch and manipulation force limits to an object. Another example can be using the on-device AI capabilities to identify slippage and provide actions to the robot end effector with low latency to reconfigure the grasp.

[0101] The embodiments disclosed herein design a controllable robotic indentor that can apply measured 3-axis forces with high precision at any spatial position of the sensor. Figure 16 Shows an example 6-DoF robotic indentor for testing the force resolution of haptic sensors. The robotic arm and stage setup can precisely apply the measured forces to the target device at a controlled contact spatial position and orientation. Figures 17A to 17B Shows example image snapshots obtained from the shear force data collection of the artificial fingertips disclosed herein at two critical moments. The timestamp t corresponds toFigure 12 The trajectories in. The superimposed arrow region shows the optical flow relative to the image without applying any force. As Figures 17A to 17B shown, for example, a tactile sensor is mounted on the robotic arm to direct the desired test surface downward for a probe with a precision of 5 μm. As another example, a probe with a hemispherical tip having a diameter of 4 mm is mounted on a force sensor that measures the ground truth contact force with a precision of 1 mN. As yet another example, the probe and force sensor assembly are then mounted on a hexapod that can be precisely controlled to translate by 0.1 μm and rotate in increments of 0.05°. Due to the rotational symmetry of the artificial fingertip, the embodiments disclosed herein decompose the full-surface force characterization into three representative approximate planar regions. For each region, the embodiments disclosed herein similarly repeat the collection process for the normal force and the shear force.

[0102] The embodiments disclosed herein begin with normal force collection. For example, for high precision, a particular embodiment can use a uniaxial force sensor that can measure up to 250 mm. As another example, for each region, the robotic indentor can spatially sample grid points at 0.5 mm intervals in the tangential plane. As yet another example, for each point, the probe can be moved perpendicular to the plane and pressed into the sensor until the normal force reaches 200 mm. During the contact of the probe with the gel (defined as the normal force, Fnorm > 0.2 mN), the sensor image and the measured normal force can be synchronously acquired. The embodiments disclosed herein collect approximately 550 image-force pairs at each spatial point. For a 7 mm × 6 mm region, the embodiments disclosed herein obtain approximately 12,000 points. This point data is randomly divided into a training set (70%) and a test set (30%).

[0103] For shear force data collection, a particular embodiment can select a 3-axis force sensor to simultaneously measure the normal force and the shear force. A particular embodiment can apply sufficient frictional force while changing the shear force. Figure 18 Shows an example normal force prediction error distribution by surface type and region. The dots and error bars show the median and 95th percentile of the error, respectively. Figure 18 Shows how each shear force indentation trajectory can be controlled. For example, first, the probe can be moved perpendicular to the contact surface to apply a normal force of up to 600 mm. Next, the probe can be moved tangentially to the surface to load a shear force of up to 100 mm. Finally, the probe can be moved back to the previous position to unload the shear force. If the remaining shear force after unloading is non-zero, slip may have occurred - in which case the data may be discarded.

[0104] An image-force regression model can be used to achieve contact force prediction for vision-based tactile sensors (e.g., the systems disclosed herein). The model can be calibrated based on reference data. Once calibrated, a particular embodiment can evaluate the sensor model as a force sensing performance system on a test dataset. Embodiments disclosed herein have collected datasets for training and evaluating the model to benchmark normal force and shear force sensing performance. A particular embodiment can use a modified ResNet50 deep neural network for the image-force regression model. For example, the network can take an input image of 224×224×3 and output 1024-way object classification probabilities. A particular embodiment can replace the classification head with a scalar output linear layer for predicting force. A particular embodiment can use mean squared error as the loss and then optimize using ADAM with an initial learning rate search. In one example embodiment, the original image from the sensor is 640×480, which can be downsampled to 224×224 with 20 pixel spatial jitter to improve spatial invariance. A particular embodiment can pool the training data from all three regions to train a single model and obtain the prediction performance (median error) broken down by region, as Figure 2A described.

[0105] Embodiments disclosed herein have evaluated the additional normal force resolution performance of two gel surface finishes (specular and Lambert). Lambertian surface scattering is generally considered to be preferred for vision-based tactile sensors, but it performs worse than specular reflection. This may be because specular reflection enhances the surface texture contrast, which helps the imaging system track gel deformation. Figure 19 An example acquisition in a vision-based touch sensor is shown. Objects shown from top to bottom: sandpaper, cloth, oyster, and configured cone. Figure 19 A central crop of an image acquired by the system disclosed herein is shown, where the texture contrast is more evident in the region with strong specular reflection from the LED. Embodiments disclosed herein have obtained clear optical flow (as shown by the arrows in the figure) within these texture-rich regions, which corresponds to fingertip deformation caused by shear and normal forces applied by the probe.

[0106] Generally, it is thought that some tracking pattern (e.g., dots) may be required to measure shear force. However, the optical flow results reported above suggest that this requirement may be relaxed due to the increasing resolution and quality of the images. These improvements can facilitate the use of natural fingertip surface texture to observe gel deformation and thereby estimate shear force.

[0107] Certain embodiments can establish modalities within an artificial fingertip to determine object states and obtain clues regarding object classification. Certain embodiments can identify two key performance metrics for fingertip gas sensing: accuracy and signal acquisition time. Embodiments disclosed herein observed 6 different materials from liquid to solid that are common in a home environment. These materials are coffee powder, liquid coffee, an unremarkable rubber material, cheese, and soap and butter applied to a surface. All materials were sampled at room temperature using a robotic arm, and the disclosed artificial fingertip approached the sample within 1 cm to make contact nearby for 90 seconds. Embodiments disclosed herein record multimodal data and isolate humidity, temperature, pressure, gas antioxidant data points at the maximum output frequency for each sampling modality. Over a 3-hour sampling period, over 100 methods for each material were collected. Between methods, embodiments disclosed herein sampled air from the local environment. The raw data with the modalities of interest listed above was provided as input to a multi-layer perceptron network with a single 64-node hidden layer. Embodiments disclosed herein used an ADAM optimizer with a learning rate of 0.1 to train the network with cross-entropy loss. Embodiments disclosed herein show that the final accuracy of the model is insensitive to the size of the hidden layer or the learning rate. For example, embodiments disclosed herein show a classification accuracy of 91% across these 6 materials. Additionally, as another example, embodiments disclosed herein show an accuracy rate of 66% for the signal acquisition time.

[0108] Embodiments disclosed herein evaluated the performance of the disclosed artificial fingertip in terms of spatial resolution, shear and normal forces, illumination, vibration, heat, and local gas sensitivity.

[0109] Figure 20A An example simulation of the human reflex arc is shown that rapidly processes sensing inputs within the fingertip and directly controls the actuators of a robotic hand to retract in response to touching an object. Figure 20B An example tactile processing and control paradigm is shown that transmits sensing data to a remote computer for processing. This requires sufficient bandwidth and introduces communication latency. Figure 20C An example local processing for simulating the reflex arc with a system fingertip is shown. The systems disclosed herein can use an on-device AI neural network accelerator for local processing to reduce the total latency between an event and an action. Figure 20D An example mean and standard deviation of the event-action latency are shown. For example, on-device local processing and control loops may take 1.2 ms compared to traditional paradigms that may take 2.5 ms on the disclosed artificial fingertip and over 6 ms on an example conventional tactile sensor.

[0110] The embodiments disclosed herein model the fingertip surface as a two-layer stack formed by an external diffusive material attached to an internal reflective film, which is grown on a non-rigid solid silicone body of the fingertip. In other words, the silicone hemispherical dome may also include a protective diffusive layer coated onto the reflective silver film layer. Thus, the embodiments disclosed herein explore the effects of the mechanical properties, texture, and degree of controlled light scattering of the non-rigid solid silicone surface to find the optimal performance metrics between background uniformity and image contrast. The embodiments disclosed herein show that increasing the controlled surface texture scattering from 1-degree scattering to Lambertian scattering results in an increase in background illumination uniformity, thereby increasing the image imprint contrast. However, at low scattering degrees, strong hot spot artifacts may dominate the background, and when the scattering degree approaches Lambertian scattering, these artifacts may decrease with a decrease in image contrast, which can directly lead to a reduced sensitivity to imprint stimuli. In the case of little or no scattering on a polished surface, there may be minimal background illumination, which stimulates the generation of shadows by the indentations on the fingertip surface. Additionally, the specular reflections on the generated indentations may be few or non-existent and may not produce a consistent appearance on the surface. In contrast, for the conventional approach of vision-tactile sensors using Lambertian scattering surfaces, the embodiments disclosed herein show that the hemispherical sensing surface can act as an integrating sphere, where the shadows projected by direct illumination onto the indentations can be eliminated by the scattered illumination from other regions, and their contrast may be low even if imaging occurs at a far off-axis angle. The embodiments disclosed herein introduce a controlled degree of scattering, where optimized uniform background illumination can be achieved, which helps in the contrast between the indentations and the surrounding surface, and in addition, all indentations are imaged (see Figure 2).

[0111] The embodiments disclosed herein evaluate the normal force sensitivity and first collect the tuples of the normal forces applied by a micro-indentation instrument and the corresponding outputs from the sensor, and then train a deep learning model based on this data set. For example, the trained model (see Figure 2A ) can predict the applied normal force with a median error of 1.01 mN (Region 1). Similarly, to measure the shear force sensitivity, the embodiments disclosed herein first collect the tuples of the shear forces applied by a micro-indentation instrument and the corresponding outputs from the sensor, and then train a deep learning model based on this data set. For example, the model (see Figure 2A ) is capable of predicting the applied shear force with a median error of 1.27 mN (Region 1). Compared with conventional sensors that require the presence of explicit markings, this result shows that, with a sufficiently high optical resolution, the embodiments disclosed herein can directly use the internal texture of the elastomer to measure the shear force.

[0112] To perform spatial resolution assessment, the embodiments disclosed herein define the spatial resolution of the artificial fingertip sensor as the minimum feature size that can be resolved when the MTF ≥ 0.5; this can be determined by the degree of contrast retention quantified in line pairs per millimeter. The embodiments disclosed herein first simulate the imaging system in design, which produces features with a size of ≥ 6 μm for region 1, ≥ 8 μm for region 2, and ≥ 22 μm for region 3, and the on-axis contact is resolvable. The embodiments disclosed herein then collect data to verify these results by pressing a double-fork micro-indentation instrument on the fingertip, changing the distance between the two forks, and observing the contact pixel intensity line profile; the visual verification and inspection of the contact pixel profile intensity confirm that the embodiments disclosed herein can clearly distinguish features as small as ≥ 7 μm in region 1 (see Figure 2C ).

[0113] Multimodal information (e.g., vibrations above 10 kHz, auditory cues, sensitivity to heat and smell) may play an important role in human touch. However, typical vision-based tactile sensors may not incorporate a wide range of multimodal capabilities to acquire this information or operate at lower sensing frequencies (e.g., 60 Hz). Even using a fast camera for the disclosed artificial fingertip operating at, for example, 240 Hz, highly dynamic motions cannot be fully captured. The embodiments disclosed herein evaluate the acquisition of vibrations up to 10 kHz, which may be sufficient to distinguish different materials during a simple slight slide of the finger. In addition, the embodiments disclosed herein show that these multimodal features can be used to detect the amount of liquid in a bottle by simply tapping the bottle with the fingertip (see Figure 2A ), such as the audio and vibration cues consumed by humans during object interaction. Based on the audio and vibration cues, humans evaluate touch interactions based on changes in local thermal gradients. The embodiments disclosed herein can detect thermal gradient changes that reflect the object state, indoor temperature, warmth, heat, danger (see Figure 2C ). Regarding the object state, there may be a limited amount of information collected during visual-tactile contact. The embodiments disclosed herein employ local gas sensing at each fingertip to understand the nuances of the object state, e.g., to determine whether the object is slippery or wet. Using this modality, the embodiments disclosed herein can sense these parameters during approach and contact (see Figure 2E ). Specifically, the embodiments disclosed herein evaluate contacting different samples that provide not only gas characteristics but also local environmental information such as humidity and temperature gradients to distinguish between two seemingly similar liquids and coffee or coffee grounds (see Figure 2D ). The use of multimodal sensing can complement the primary vision-based sensing modality and enable future work to study the importance of different touch patterns for specific task applications.

[0114] Inspired by the human reflex arc, the embodiments disclosed herein demonstrate a fast reflex-like control loop that uses an on-device AI neural network accelerator for local processing. Compared to traditional sensors that use an external computer for processing, on-device processing on the artificial fingertips disclosed herein can reduce latency, e.g., from 6 ms to 1.2 ms (see Figure 3D ). As the computing power of on-device accelerators increases, the surface area available for sensing touch grows, and AI models are increasingly used for touch processing, the ability to locally process data and transmit only high-level features may prove crucial for touch processing.

[0115] The embodiments disclosed herein can advance the state sensed by the artificial fingertips towards digitizing the fingertip interaction between the environment and the object. The embodiments disclosed herein disclose an artificial fingertip that may be more sensitive in terms of spatial and force sensitivity compared to traditional methods, with additional technical advantages of multimodal sensing features and local processing capabilities. Experimental results show that the ability to digitize touch is superior to that of human fingertips. The rich touch digitized by the disclosed modular platform can open up new and promising venues for studying the nature of human touch and investigating key issues surrounding the digitization and processing of touch as a sensor modality. In addition, the embodiments disclosed herein can open the door for wider adoption of touch sensors beyond traditional business opportunities: in robotics, improving sensing and manipulation capabilities, bringing benefits to applications in manufacturing and logistics, medical robots, agricultural robots, and consumer-grade robots; in artificial intelligence, studying the learning of appropriate tactile and multimodal representations, and corresponding computational models that can better utilize the active, spatial, and temporal characteristics of touch. Further potential applications can include virtual reality and telepresence, prosthetics, and e-commerce.

[0116] Other aspects

[0117] In this document, unless otherwise expressly stated or the context otherwise indicates, "or" is inclusive rather than exclusive. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A or B" means "A, B, or both". Further, unless otherwise expressly stated or the context otherwise indicates, "and" is both conjunctive and disjunctive. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A and B" means "A and B, jointly or severally".

[0118] The scope of the present disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or shown herein, which will be understood by those skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or shown herein. Additionally, although the present disclosure describes and shows the various embodiments herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any of the components, elements, features, functions, operations, or steps described or shown anywhere herein that would be understood by a person of ordinary skill in the art. Further, a reference in the appended claims to a device or system, or a component of a device or system, that is adapted to, arranged to, capable of, configured to, enabling, operable, or operable to perform a particular function includes that device, system, component, whether or not the particular function is activated, turned on, or unlocked, so long as the device, system, or component is so adapted to, arranged to, capable of, configured to, enabling, operable, or operable. Additionally, although the present disclosure describes or shows a particular embodiment as providing a particular advantage, a particular embodiment may not provide that advantage, or may provide some or all of that advantage.

Claims

1. A system for touch digitization, the system comprising: A silicone hemispherical dome, the silicone hemispherical dome including a surface containing a reflective silver film layer; An omnidirectional optical system, the omnidirectional optical system comprising: A lens, the lens including a plurality of lens elements, wherein a first lens element of the plurality of lens elements is in direct contact with the silicone hemispherical dome without an air gap, and wherein the lens is configured to collect scattering of internally incident light generated by the reflective silver film layer; and An image sensor, the image sensor being configured to generate image data based on data collected by the lens; One or more non-image sensors, the one or more non-image sensors being disposed below the omnidirectional optical system; One or more processors; and A non-transitory memory, the non-transitory memory being coupled to the processor, the non-transitory memory including instructions executable by the processor, the processor being operable when executing the instructions to: Access the image data from the omnidirectional optical system and the sensed data from the one or more non-image sensors; and Generate the touch digitization based on the accessed image data and sensed data through one or more machine learning models.

2. The system according to claim 1, wherein, The lens is a solid immersion lens.

3. The system according to claim 1, wherein, The lens is a hyperfisheye lens.

4. The system according to claim 1, wherein The one or more non-image sensors include one or more of an inertial measurement unit (IMU) sensor, a microphone, an environmental sensor, a gas sensor, a pressure sensor, or a temperature sensor.

5. The system according to claim 1, wherein, The first lens element is configured to have a maximum clear aperture diameter of 10 millimeters.

6. The system according to claim 1, wherein The material associated with the silicone hemispherical dome is determined based on a plurality of material parameters, the plurality of material parameters including one or more of a gel radius, a coating thickness, a gel layer thickness, a height, a coating Young's modulus, or a gel Young's modulus.

7. The system according to claim 1, wherein The silicone hemispherical dome is based on a polydimethylsiloxane (PDMS) material.

8. The system according to claim 1, wherein The silicone hemispherical dome further includes a protective diffusion layer coated onto the reflective silver film layer.

9. The system according to claim 1, the system further comprising: A housing based on the shape of a human thumb, wherein the silicone hemispherical dome, the omnidirectional optical system, the one or more non-image sensors, the one or more processors, and the non-transitory memory are disposed in the housing.

10. The system according to claim 1, the system further comprising: A stack including a plurality of printed circuit boards for the one or more processors, the omnidirectional optical system, and a data transfer system, wherein the plurality of printed circuit boards share a common electrical interface and a connector stack.

11. The system according to claim 1, wherein, The one or more processors include one or more of a microprocessor or an accelerator.

12. The system according to claim 1, wherein, The one or more machine learning models include one or more neural network models, wherein the one or more processors include one or more neural network accelerators, and wherein the one or more neural network accelerators are configured to accelerate real-time inference of the accessed image data and sensed data by the one or more neural network models.

13. The system according to claim 1, wherein, The processor is further operable when executing the instructions to: Provide one or more control signals to an auxiliary device associated with the system.

14. The system according to claim 13, wherein, The auxiliary device includes a robotic end effector.

15. The system according to claim 1, wherein The silicone hemispherical dome is produced based on: Manufacturing a mold with aluminum; Finishing the mold with a machine polishing pass; Preparing the mold for gel casting through a salinization process in a dryer; Preparing a gel material using a curable silicone rubber compound; Combining the gel material in a rapid mixer under vacuum; Pouring the gel material into the mold; Curing the poured gel material for a first amount of time at a first temperature; And Once the poured gel material is cured, removing the gel hemispherical dome from the mold.

16. The system according to claim 15, wherein, The reflective silver film layer is produced based on: Preparing a glucose solution by dissolving a first amount of glucose in a second amount of H2O and adding a third amount of KOH; Preparing an AgNO3 solution by dissolving a fourth amount of AgNO3 in a fifth amount of H2O and adding a sixth amount of NH3; Preparing a plating solution by mixing the glucose solution and the AgNO3 solution; Cleaning the gel hemispherical dome with oxygen plasma for a second amount of time; Activating the gel hemispherical dome in a solution of a seventh amount of SnCl2 in an eighth amount of H2O for a third amount of time; Suspending the gel hemispherical dome in the plating solution for a fourth amount of time; Rinsing the gel hemispherical dome with H2O; And Air-drying the gel hemispherical dome.

17. The system according to claim 1, the system further comprising: A lighting system, wherein the lighting system includes a plurality of controllable light-emitting diodes that emit Lambertian diffused light, and wherein the lighting system is configured to produce volumetric lighting having one or more configurable lighting parameters.

18. The system according to claim 17, wherein, The one or more configurable lighting parameters include one or more of wavelength, intensity, or positioning.

19. An artificial fingertip for touch digitization, the artificial fingertip comprising: A silicone hemispherical dome; An omnidirectional optical system, the omnidirectional optical system including: A lens, the lens including a plurality of lens elements, wherein a first lens element of the plurality of lens elements is in direct contact with the silicone hemispherical dome without an air gap; and An image sensor configured to generate image data based on data collected by the lens; and one or more non-image sensors disposed below the omnidirectional optical system.

20. A system for touch digitization, the system comprising: A silicone hemispherical dome, the silicone hemispherical dome including a surface containing a reflective silver film layer; An omnidirectional optical system, the omnidirectional optical system including: A lens configured to collect the scattering of internally incident light generated by the reflective silver film layer; and An image sensor configured to generate image data based on data collected by the lens; and one or more non-image sensors disposed below the omnidirectional optical system.

Citation Information

Patent Citations

  • Mastic, caulking, sealant and adhesive compositions containing photosensitive compounds and method of reducing their surface tack

    EP0010000A1