Event-driven visual and tactile sensing and learning for robots
The VT-SNN system addresses the inefficiencies of existing tactile and event-based sensing by integrating event-based visual and tactile sensors with spiking neural networks, achieving efficient object classification and slip detection in robots.
Patent Information
- Application Number
- JP2022576503
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-15
- Filing Date
- 2021-06-15
- Publication Date
- 2025-09-04
- Estimated Expiration
- 2041-06-15
AI Technical Summary
Existing tactile and event-based sensing technologies for robots are resource-intensive and underdeveloped, lacking efficient integration and scalability, especially in multimodal sensory applications such as visual-tactile learning, which are crucial for tasks like object classification and slip detection.
A visual-tactile spiking neural network (VT-SNN) system that combines event-based outputs from visual and tactile sensors using spiking neural networks (SNNs), integrated with a neuromorphic tactile sensor (NuTouch) for efficient, low-latency processing, and a scalable event-based tactile sensor design.
The system achieves competitive performance in object classification and slip detection, outperforming traditional deep learning methods with reduced power consumption and latency, demonstrating the potential for power-efficient intelligent robots.
Smart Images

Figure 0007734155000012 
Figure 0007734155000013 
Figure 0007734155000014
Abstract
Description
[Technical Field]
[0001] The present invention relates to a classification sensing system and method for large markers, and in particular to event-driven visual and tactile sensing and learning for robots. [Background technology]
[0002] Any mention and / or discussion of prior art throughout this specification should in no way be taken as an admission that this prior art is well known or forms part of the common general knowledge in the art.
[0003] Many everyday tasks require multiple sensory modalities to be successful. Consider, for example, removing a carton of soy milk from the refrigerator [1]; humans can use vision to locate the carton and infer from a simple grasp how much soy milk the carton contains. They can also use their vision and touch to lift objects without slipping them. These actions (and inferences) are performed robustly using power-efficient neural substrates, and the human brain requires much less energy compared to common deep learning approaches that use multiple sensor modalities in artificial systems [2], [3].
[0004] Below, we provide a brief overview of work on visual-tactile sensing for robotics and event-driven sensing and learning. In visual-tactile sensing for robots in general, there has been recognition of the importance of multimodal sensing for robotics, leading to innovations in both sensing methods and sensing techniques. Recently, there have been many papers combining vision and touch sensing, e.g., [8]–
[13] . However, research on visual-tactile learning of objects dates back to (at least) 1984, when vision and tactile data were used to create surface descriptions of primitive objects
[14] ; in this early work, tactile sensing played a supporting role to vision due to the low resolution of tactile sensors at that time.
[0005] Recent advances in haptic technology
[15] have encouraged the use of tactile sensing for more complex tasks, including object exploration
[16] and classification
[17] , shape completion
[18] , and slip detection
[19] ,
[20] . One popular sensor is BioTac, which uses textured skin similar to a human finger, allowing vibration signatures to be used for high-precision material and object identification as well as slip detection
[21] . BioTac has also been used in visuo-tactile learning, e.g., [9] combining tactile data with RGB images to recognize objects via deep learning. Other recent work has used Gelsight
[22] —an optically based tactile sensor—for visual-tactile slip detection
[10] ,
[23] , grasp stability, and texture recognition
[24] . Very recent work has used unsupervised learning to generate neural representations of visuo-tactile data (with proprioception) for reinforcement learning
[11] .
[0006] In event-based sensing, sensors and learning are primarily focused on vision (see
[25] for a comprehensive survey). The emphasis on vision can be attributed both to its applicability across many tasks, as well as the recent availability of event cameras such as those onboard DVS and Prophesee, as well as the fact that, unlike traditional optical sensors, event cameras change pixels asynchronously. Event-based sensors have been successfully used in combination with deep learning techniques
[25] . Binary events are first converted into real-valued tensors, which are processed downstream by deep artificial neural networks (ANNs). This approach generally yields good models (e.g., motion segmentation
[26] , optical flow estimation
[27] , and car steering prediction
[28] ), but is computationally expensive.
[0007] Neuromorphic learning, specifically spiking neural networks (SNNs) [4],
[29] , offers a competing approach for learning with event data. Like event-based sensors, SNNs operate directly on discrete spikes and therefore have similar properties: low latency, high temporal resolution, and low power consumption. Historically, SNNs have been hampered by the lack of good training procedures. Gradient-based methods such as backpropagation have been unavailable due to the non-differentiable nature of spikes. Recent developments in effective SNN training
[30] –
[32] , and the emerging availability of neuromorphic hardware (e.g., IBM TimeNorth
[33] and Intel Loihi [7]), have spurred renewed interest in neuromorphic learning for a variety of applications, including robotics. SNNs have yet to consistently outperform their deep ANN cousins on simulated event image datasets, and the research community is actively exploring better training methods for real-event data. Summary of the Invention
[0008] SUMMARY OF THE INVENTION Embodiments of the present invention seek to address at least one of the above problems.
[0009] According to a first aspect of the present invention, there is provided a classification sensing system comprising: a first spiking neural network (SNN) encoder configured to encode an event-based output of a visual sensor into individual visual modality spiking representations having a first output size; a second SNN encoder configured to encode an event-based output of a tactile sensor into individual tactile modality spiking representations having a second output size; a combination layer configured to merge the visual modality spiking representations and the tactile modality spiking representations; and a task SNN configured to receive the merged visual modality spiking representations and the tactile modality spiking representations and output a visual-tactile modality spiking representation having a third output size.
[0010] According to a second aspect of the present invention, there is provided a classification method performed using a sensing system, the classification method including: using a first spiking neural network (SNN) encoder to encode event-based outputs of a visual sensor into individual visual modality spiking representations having a first output size; using a second SNN encoder to encode event-based outputs of a tactile sensor into individual tactile modality spiking representations having a second output size; using a connection layer to merge the visual modality spiking representations and the tactile modality spiking representations; and using a task SNN to receive the merged visual modality spiking representations and the tactile modality spiking representations and output visual-tactile modality spiking representations having a third output size for classification.
[0011] According to a third aspect of the present invention, there is provided a tactile sensor comprising: a carrier structure; an electrode layer disposed on a surface of the carrier structure and including an array of a plurality of taxel electrodes; a plurality of electrode wires individually electrically connected to each one of the plurality of taxel electrodes; a protective layer disposed on the electrode layer and made of an elastically deformable material; and a pressure transducer layer disposed between the electrode layer and the protective layer, wherein detectable electrical signals in the plurality of electrode wires responsive to contact forces acting on the pressure transducer layer through the protective layer provide spatiotemporal data for neuromorphic tactile sensing applications.
[0012] According to a fourth aspect of the present invention, there is provided a method for manufacturing a tactile sensor, the method comprising: providing a carrier structure; providing an electrode layer disposed on a surface of the carrier structure and including an array of taxel electrodes; providing a plurality of electrode wires individually electrically connected to each one of the plurality of taxel electrodes; providing a protective layer disposed on the electrode layer and made of an elastically deformable material; and providing a pressure transducer layer disposed between the electrode layer and the protective layer, wherein detectable electrical signals in the plurality of electrode wires responsive to contact forces acting on the pressure transducer layer through the protective layer provide spatiotemporal data for neuromorphic tactile sensing applications.
[0013] Embodiments of the present invention will be better understood and readily apparent to those skilled in the art from the following description, given by way of example only, in conjunction with the drawings in which: [Brief explanation of the drawings]
[0014] [Figure 1a] FIG. 1a shows a photograph of a NeuTouch event-driven tactile sensor in accordance with one embodiment compared to a human finger. [Figure 1b] FIG. 1b shows a photograph of a partial cross section of a new touch-event-driven tactile sensor according to one embodiment. [Figure 1c]FIG. 1c shows a photograph of the spatial distribution of 39 taxels on a new touch event-driven tactile sensor according to one embodiment. [Figure 1d] Figure 1d shows the pressure response of the converter in a new touch-event-driven tactile sensor according to one embodiment. Low hysteresis can be observed from the loading and unloading curves. [Figure 1e] FIG. 1e shows a graph illustrating asynchronous (signature encoded) transmission of tactile information from a new touch-event driven tactile sensor according to an exemplary embodiment. [Figure 1f] FIG. 1f illustrates a graph showing decoded tactile information (ie, events) from a new touch-event-driven tactile sensor according to an exemplary embodiment. [Figure 2] FIG. 2 shows a schematic diagram of the architecture of a visual-tactile spiking neural network (VT-SNN) according to an exemplary embodiment, which first encodes the two modalities into individual latent (spiking) representations, which are then combined in a combination layer and further processed through additional layers to produce a task-specific output. [Figure 3] Figure 3a shows a photograph of a 7-DoF Franka Emika Panda arm equipped with a Robotiq 2F-140 gripper, a Prophesee Onboard even-based camera, and an RGB camera, all equipped with a new touch event-driven tactile sensor, according to one embodiment. Figure 3b shows a photograph of the 7-DoF Franka Emika Panda arm equipped with the Robotiq 2F-140 gripper and Optitrack motion capture system of Figure 3a. [Figure 4] FIG. 4 shows graphs with visual spike images and records illustrating tactile and visual data from the grasping, lifting, and holding phases for training and testing of a VT-SNN in accordance with an exemplary embodiment. [Figure 5] FIG. 5 shows photographs of containers used in a container classification task: a coffee can, a plastic soda bottle, a soy milk carton, and a metal tuna can, for classification using a VT-SNN according to an example embodiment. [Figure 6] FIG. 6 shows a graph illustrating output spikes for models trained on different modalities with correct and incorrect predictions, while a VT-SNN according to one embodiment, along with haptic (only) and visual (only) models for comparison, can grasp a coffee can with 100% weight in a classification task. [Figure 7] FIG. 7 shows a graph illustrating classification accuracy of containers and weights over time in a classification task using a VT-SNN according to an example embodiment, and tactile (only) and visual (only) models for comparison. [Figure 8] Figure 8a shows a photograph of an object for a check classification task with attached OptiTrack markers using a VT-SNN according to an exemplary embodiment, a tactile (only) model, and a visual (only) model for comparison. Figure 8b is a photograph of the subject in Figure 8a during a stable grasp. Figure 8c is a photograph of the subject in Figure 8a during an unstable grasp due to rotational slippage. [Figure 9] FIG. 9 shows a graph illustrating slip classification accuracy over time in a classification task using a VT-SNN according to an example embodiment, and a haptic (only) model and a visual (only) model for comparison. [Figure 10] FIG. 10 shows a photograph of a 3D printed main holder, according to one embodiment. [Figure 11] FIG. 11 shows a photograph of an enclosure used in an ACES encoder according to an example embodiment. [Figure 12] FIG. 12 shows a photograph of a coupler used in NuTouch according to one embodiment. [Figure 13] Figure 13a shows a graph of pz of the end effector over time in accordance with an exemplary embodiment, and Figure 13b shows a graph of Θt (the shortest angle in radians) calculated between qt and qo in accordance with an exemplary embodiment. [Figure 14] FIG. 14 is a schematic diagram illustrating a classification sensing system in accordance with an exemplary embodiment. [Figure 15]FIG. 15 shows a flowchart illustrating a classification method performed using a sensing system according to an example embodiment. [Figure 16] FIG. 16 shows a schematic diagram illustrating a tactile sensor according to one embodiment. [Figure 17] FIG. 17 shows a flow chart illustrating a method for manufacturing a tactile sensor, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0015] Embodiments of the present invention provide an important step toward efficient visual-tactile sensing for asynchronous and event-driven robotic systems. In contrast to resource-intensive deep learning methods, event-driven sensing forms an alternative approach that promises power efficiency and low latency, features ideal for real-time mobile robots. However, event-driven systems remain underdeveloped compared to standard synchronous perception methods [4], [5].
[0016] To enable richer tactile sensing, example embodiments provide a 39-taxel fingertip sensor, referred to herein as NuTouch. Compared to existing commercially available tactile sensors, NuTouch's neuromorphic design allows scaling to a larger number of taxels while maintaining low latency.
[0017] Multimodal learning with NuTouch and the Prophesee event camera was investigated in accordance with the illustrative embodiment. Specifically, we present a visual-tactile spiking neural network (VT-SNN) that incorporates both sensory modalities for a supervisory learning task. Unlike traditional deep artificial neural network (ANN) models [6], SNNs process discrete spikes asynchronously and are therefore potentially more suitable for the event data generated by neuromorphic sensors in accordance with the illustrative embodiment. Furthermore, SNNs can be utilized on efficient, low-power neuromorphic chips such as Intel Loihi [7].
[0018] It should be noted that in exemplary embodiments, other event-based tactile sensors may be used, and the tactile sensor may also include a converter for converting the intrinsic output of the tactile sensor to the event-based output of the tactile sensor.
[0019] Similarly, it should be noted that in exemplary embodiments, other event-based visual sensors may be used, and the visual sensor may include a converter for converting the intrinsic output of the visual sensor to the event-based output of the visual sensor.
[0020] Experiments performed in accordance with exemplary embodiments center on two robotic tasks: object classification and (rotational) slippage detection. In the former, the robot was tasked with determining the type of container being handled and the amount of liquid held therein. The containers were opaque with different stiffnesses, and therefore both visual and tactile sensing were relevant for accurate classification. Relatively small differences in weight (~30 g across 20 object weight classes) were shown to be distinguishable by the prototype sensor and spiking model according to exemplary embodiments. Similarly, slippage detection experiments showed that rotational slippage could be accurately detected within 0.08 s (visual-tactile spikes processed every ~1 ms). In both experiments, the SNN achieved competitive (and sometimes superior) performance compared to ANNs with similar architectures.
[0021] Considering a broader perspective, event-driven sensing in accordance with illustrative embodiments represents an interesting opportunity to enable power-efficient intelligent robots. In accordance with illustrative embodiments, an "end-to-end" event-driven sensing framework can be provided.
[0022] NuTouch, according to one embodiment, provides a scalable, event-based tactile sensor for robotic end effectors.
[0023] A visual-tactile spiking neural network according to an exemplary embodiment utilizes multiple event sensor modalities.
[0024] Systematic experiments demonstrate the effectiveness of the event-driven sensing system according to the exemplary embodiment for object classification and slip detection in comparison with conventional ANN methods.
[0025] Using the exemplary embodiment, a visual-tactile event sensor dataset was obtained that included over 50 different object classes, and also included RGB images and proprioceptive data.
[0026] NuTouch: An Event-Based Tactile Sensor According to an Embodiment While tactile sensors have numerous applications (e.g., minimally invasive surgery
[38] and smart prosthetic devices
[39] ), current tactile sensing technology lags behind vision. In particular, current tactile sensors remain difficult to scale and integrate with robotic platforms. This is due to two reasons. First, many tactile sensors are interfaced via time-division multiple access (TDMA), where individual taxel electrodes, hereafter referred to as "taxels," are sampled periodically and sequentially. The serial readout nature of TDMA inherently leads to increased readout latency as the number of taxels in the sensor increases. Second, high spatial localization accuracy is typically achieved by adding more taxels in the sensor; this inevitably leads to more wiring, which complicates integration of the skin onto robotic end-effectors and surfaces.
[0027] Motivated by the limitations of existing tactile sensing technology, a neuro-inspired tactile sensor 100 (NuTouch) is provided in accordance with an illustrative embodiment for use in robotic end effectors (see FIG. 1 ). The structure of NuTouch 100 resembles a human fingertip 102, including "skin" and "bone," and, according to an illustrative embodiment, has physical dimensions of 37 x 21 x 13 mm. This design facilitates integration with humanoid end effectors (for prosthetic or humanoid robots) and standard multi-finger grippers; in experiments, NuTouch 100 was used with a Robotiq 2F-140 gripper. Note that in addition to the fingertip design, alternative structures can be developed to suit different applications according to different illustrative embodiments.
[0028] Specifically, Figure 1a shows the NuTouch 100 compared to a human finger 102. Figure 1c shows the spatial distribution of 39 taxels, such as taxel 104, on the NuTouch 100. Figure 1b shows a partial cross-sectional view of the NuTouch 100 and its components. The NuTouch 100 performs tactile sensing using an electrode layer 106 having 39 taxels, such as taxel 104, and a graphene-based piezoresistive thin film 108 as a pressure transducer embedded under a protective Ecoflex "skin" 110, all of which is supported on a 3D printed part ("bone") 112.
[0029] Tactile sensing is achieved through an electrode layer 106 folded around the bone 112 such that an array of electrodes with 39 taxels, e.g., taxel 104, is on top of the bone 112 with a graphene-based piezoresistive thin film 108 covering the 39 taxels, e.g., taxel 104. The graphene-based piezoresistive thin film 108 functions as a pressure converter, forming an effective tactile sensor
[40] ,
[41] due to its high Young's modulus, which helps reduce the converter's hysteresis and response time. The radial placement of taxels, e.g., 106, on the NuTouch 100 sensor is designed so that the taxel concentration varies from high to low from the center to the periphery of the NuTouch 100 sensor's "top" touch surface. The initial contact point between the object and the sensor is located in the central region of the NuTouch 100, where the density of taxels, e.g., 106, is highest, allowing for the capture of rich spatiotemporal tactile data of the initial contact (between the object and the sensor). This rich tactile information can help algorithms accelerate inference (e.g., early classification, as described in more detail below).
[0030] FIG. 1d shows the pressure response of the transducer in the NuTouch 100, where low hysteresis can be observed from the loading and unloading curves.
[0031] 3D-printed bone components 112 were used to serve as fingertip bones, and Ecoflex 00-30 (Ecoflex) 110 was used to emulate the skin of the NuTouch 100. Ecoflex 110 provides protection for electrodes / taxels, such as taxel 104, for a longer service life and amplifies the stimuli exerted on the NuTouch 100. The latter allows for the collection of more tactile features, as the transient phase of contact (between the object and the sensor) encodes many of the physical descriptions of the grasped object, such as stiffness or surface roughness
[42] . The NuTouch 100 exhibits a slight delay of ~300 ms in recovering from deformation due to the soft nature of Ecoflex 110. Nevertheless, experiments described below showed that this effect does not interfere with the NuTouch 100's sensitivity to various tactile stimuli.
[0032] Compared to existing tactile sensors, the NuTouch 100 is event-based and scales well with the number of taxels. The NuTouch 100 can accommodate 240 taxels in a non-limiting embodiment while maintaining an exceptionally low constant readout latency of 1 ms for rapid tactile recognition
[43] . This is achieved in accordance with an exemplary embodiment by leveraging the Asynchronous Encoding Electronic Skin (ACES) platform
[43] —an event-based, neuromimetic architecture that enables asynchronous transmission of tactile information. In ACES, the NuTouch 100's taxels, such as taxel 104, mimic the function of fast-adapting (FA) mechanoreceptors in the human fingertip, which capture dynamic pressure (i.e., dynamic skin deformation)
[44] . FA responses are important for fine manipulation tasks that require rapid detection of object slippage, object stiffness, and local curvature.
[0033] A variety of suitable materials may be used to manufacture the NuTouch 100 according to exemplary embodiments, including, but not limited to: Skin layer: Ecoflex series (Smooth-On), polydimethylsiloxane (PDMS), Dragon Skin series (Smooth-On), silicone rubber.
[0034] Converter layer (piezoresistive): Velostat (3M), Linqstat series (Caplinq), conductive foam sheet (e.g., Faird Technologies EMI), conductive fabric / textile (e.g., 3M), any piezo-resistive material.
[0035] Electrode layer: Flexible printed circuit board (Flex PCB) of different thickness. Material: Polyimide Electrode wire: Metal layer such as copper or any conductive metal (e.g. silver) Taxel: Copper, conductive metal (silver, etc.) Asynchronous Transmission of Tactile Stimuli According to an Exemplary Embodiment Compared to existing tactile sensors, NuTouch 100 is event-based and scales well with the number of taxels, e.g., taxel 104, while maintaining exceptionally low, constant readout latency of 1 ms for rapid tactile sensations. This is achieved in accordance with exemplary embodiments by leveraging the Asynchronous Encoding Electronic Skin (ACES) platform
[50] —an event-based, neuromimetic architecture that enables asynchronous transmission of tactile information. It was developed to address the increasing complexity and need for transferring large arrays of skin-like transducer inputs while maintaining a high level of responsiveness (i.e., low latency).
[0036] In ACES, taxels in NuTouch 100, such as taxel 104, mimic the function of fast-adapting (FA) mechanoreceptors in human fingertips, capturing dynamic pressure (i.e., dynamic skin deformation). Similar to biological systems, tactile stimulus information is transmitted in the form of asynchronous spikes (i.e., electrical pulses), with data transmitted by individual taxels, such as taxel 104, only when necessary via a single common conductor for signaling. This is made possible by encoding taxels, such as taxel 104 in NuTouch 100, with unique electrical pulse signatures. These signatures are robust enough to overlap, allowing multiple taxels, such as taxel 104, to transmit data without specific time synchronization (see Figure 1e). Therefore, the stimulus information of all activated taxels, such as taxel 104, can be combined upstream and propagated to a decoder via a single conductor. This results in low readout latency and simplified wiring. The decoder correlates the received pulses (i.e., the combined pulse signature) with the known signatures of each taxel, e.g., taxel 104, to extract spatiotemporal tactile information (see FIG. 1e). Each "signature" is a sequence of spikes, i.e., when one taxel "fires," a time sequence of spikes is generated instead of a single spike that can be identified in the decoder, for output of a sequence of (single) spikes for each identified taxel that fired (see FIG. 1e).
[0037] In the illustrated embodiment, each taxel, e.g., taxel 104, connects to an encoder via an electrode wire, e.g., electrode wire 105 (e.g., if there are 39 taxels, there will be 39 encoders). The signal outputs of the encoders are combined into one "common" output conductor for data transmission to a decoder. The decoder then decodes the combined pulse (spike) signature to identify activated taxels.
[0038] Real-time decoding of haptic information (captured by NuTouch 100) is performed via a Field Programmable Gate Array (FPGA) according to an exemplary embodiment. Event-based haptic information is easily accessible to a PC via UART (Universal Asynchronous Receiver / Transmitter) readout according to an exemplary embodiment.
[0039] In an exemplary embodiment, for details of the asynchronous transmission of tactile stimuli for an event-based tactile sensor suitable for us, see WO 2019 / 112516.
[0040] Below we provide details on how the decoded haptic event data is used for training and classification according to an embodiment.
[0041] Visual-Tactile Spiking Neural Network (VT-SNN) with Examples As mentioned above, successful completion of many tasks relies on the use of multiple sensory modalities. In the exemplary embodiment, the focus is on touch and vision, i.e., tactile and visual data from the NuTouch 100 and event-based cameras, respectively, are fused via a spiking neural model. This visual-tactile spiking neural network (VT-SNN) enables learning and sensing using both of these modalities and can be easily extended to incorporate other event sensors according to different exemplary embodiments.
[0042] Model Architecture According to an Illustrative Embodiment From a bird's-eye view, the VT-SNN 200 according to the exemplary embodiment uses a simple architecture (see FIG. 2 ) that first encodes the two modalities into individual latent (spiking) representations, denoted by numerals 202, 204, which are then combined in a combination layer 211 and further processed through additional layers to produce a task-specific output 213.
[0043] While the exact network architecture used in one exemplary embodiment is detailed below, the VT-SNN can use alternative network architectures for the haptic, visual, and task SNNs according to different exemplary embodiments. The haptic SNN 208 uses a fully connected (FC) network consisting of two dense spike layers (note that in preliminary experiments, convolutional layers were also tested according to other exemplary embodiments, resulting in poor performance). It has an input size of 156 (two fingers, 39 taxels each with positive and negative polarity channels per taxel) and a hidden layer size of 32. Input to the haptic SNN 208 is obtained via the signature decoder described above with reference to Figures 1e and 1f; see specifically Figure 1f for an example decoder output. The visual SNN 210 uses three layers, the first of which is a pooling layer with a kernel size and stride length of 4. The pooled spike trains are passed as input to a two-layer FC architecture identical to the haptic SNN 208. The haptic and visual encoders have output sizes of 50 and 10, respectively (several different dimension sizes were tested according to the exemplary embodiment, with 50-10 encoding providing the best results). The encoded spike inputs for both modalities are merged in a combination layer 211 and passed to a dense spiking layer (i.e., task SNN 212), which generates output spikes 206. Note that the lower part of Figure 2 (SRM type) shows the operation of a single neuron in the combination layer 211. The SRM type is used in the exemplary embodiment in all layers in the neural network, including the haptic, visual, and task SNNs 208, 210, task SNN 212, and the combination layer. The output spikes 206 are input to task SNN 212. The lower part of Figure 2 shows only a subset of the various inputs to a single neuron for illustrative purposes; typically, many more such inputs exist, as will be understood by those skilled in the art. Note that the output dimensionality (output 213) of task SNN 212 depends on the task: 20 for container & weight classification and 2 for rotation slip classification.The model architecture is independent of the size of the input time dimension, and the same model architecture is used for both classification tasks.
[0044] Neuron Model According to an Exemplary Embodiment The Spike Response Model (SRM)
[30] ,
[45] was used in the exemplary embodiment. In SRM, a spike is generated whenever the internal state ("membrane potential") u(t) of a neuron exceeds a given threshold φ. The internal state of each neuron is affected by the incoming spike and its refractory response:
[0045]
number
[0046] where n is the synaptic weight, * denotes convolution, and s i where (t) is the incoming spike from input i, ε(-) is the response kernel, v(-) is the refractory kernel, and o(t) is the neuron's output spike train 206. In other words, the incoming spike s i t is convolved with the response kernel ε(-) to generate a spike response signal scaled by the synaptic weight. That is, referring again to FIG. 2, the visual-tactile spiking neural network (VT-SNN) 200 comprises two "spiking encoders" 208, 210 per modality. Spikes from these two encoders are combined via a fixed-width connection layer 210 and propagated to a task SNN 212, which outputs a task-specific output spike train 213. The VT-SNN 200 uses a spike-response model (SRM) neuron, denoted by numeral 214, that integrates incoming spikes and spikes when a threshold is violated.
[0047] Model Training According to an Exemplary Embodiment The spiking network was optimized using SLAYER
[30] in the exemplary embodiment. As mentioned above, the derivatives of spikes are undefined, which prohibits the direct application of backpropagation to SNNs. SLAYER overcomes this problem by using a stochastic spiking neuron approximation to derive approximate gradients and a temporal credit allocation policy to distribute the error. SLAYER trains the model "offline" on GPU hardware. Therefore, the spiking data needs to be binned into fixed-width intervals during the training process, but the resulting SNN model can be run on neuromorphic hardware. Each bin window V w The (binary) value for that window V w The total spike count within the threshold S min was 1 whenever exceeded. In the exemplary embodiment, a straight-forward binning process was used:
[0048]
number
[0049] Following
[30] , class prediction is determined by the number of spikes in the output layer spike train, with each output neuron associated with a particular class, and the neuron producing the most spikes representing the winning class. In an exemplary embodiment, the model was trained by minimizing the loss:
[0050]
number
[0051] A generalization of the spike count loss in equation (3) is introduced to incorporate temporal weighting:
[0052]
number
[0053] L ω is called the weighted spike counting loss. In the experiment, ω(t) is set to be monotonically decreasing, encouraging early classification by lowering the weight of later spikes. Specifically, ω(t)=βt 2 A simple quadratic function of +γ with β<0 is used, although other functions may be used in different exemplary embodiments. ω For , corresponding counts are specified for the correct and incorrect classes, which are task-specific hyperparameters. The hyperparameters were manually tuned, and it was found that setting the positive class count to 50% of the maximum spike count (across each input in the considered time interval) worked well. In initial testing, it was observed that training with only the above loss led to rapid overfitting and poor performance on the validation set. Several techniques to mitigate this issue (e.g., l1 regularization and shedding) were explored, and simple l2 regularization was found to lead to the best results.
[0054] Robot and Sensor Setup in Accordance with Illustrative Embodiments Figure 3 shows the robot hardware setup used throughout the experiments, according to an exemplary embodiment. It consists of a 7-DoF Franka Fmika Panda arm 300 with a Robotiq 2F-140 gripper 302 and collects data from four main sensor types: NuTouch 304, 306, Prophesee Onboard 308, RGB camera 310, and Optitrack motion capture system 314. The latter two are non-event sensors, and their data streams were not used by the VT-SNN.
[0055] New Touch Tactile Sensor according to the embodiment Two NuTouch sensors 304, 306 were mounted on a Robotiq 2F-140 gripper 302, and an ACES decoder 316 was mounted on a Panda arm 300 (Figure 3a). To ensure consistent data, a sensor warm-up was performed before each data collection session to obtain baseline results to check for sensor drift. Specifically, for 100 warm-up cycles, the gripper was closed on a flat, hard object (the '9-hole peg test' from the YCB dataset
[46] ), opened for 3 seconds, and then paused for 2 seconds. A benchmark data set was then collected: 20 repetitions of closing the gripper on the same '9-hole peg test' for 3 seconds. Throughout the experiment, periodic tests for sensor drift were performed by repeating the closure test with the '9-hole peg test' as described above and then examining the sensor data; no significant drift was found throughout the experiment.
[0056] Prophesee Event Camera, According to an Exemplary Embodiment Event-based vision data was captured using the Prophesee Onboard (https: / / www.prophesee.ai) 308. Similar to a tactile sensor, each camera pixel fires asynchronously, resulting in positive (negative) spikes when there is an increase (decrease) in luminance. The Prophesee Onboard 308 was attached to an arm 300 and aimed at a gripper 302 to capture information about the object of interest (Figure 3a). The camera 308 has a maximum resolution of 640 x 480, but to minimize noise from irrelevant areas, according to one embodiment, spikes were captured from a cropped 200 x 250 rectangular window. The bias parameters of the event camera 308 were tuned according to recommended guidelines (https: / / support.prophesee.ai / portaFkb / articles / bias-tuning), and the same parameters were used throughout all testing. Table 1 shows the key biases selected using Prophesee's rules. Note that the parameter values are unitless. During preliminary experiments, we found that the Prophesee Onboard 308 was sensitive to high-frequency (≥100 Hz) light intensity changes; in other words, flickering light bulbs triggered unwanted spikes. To counter this effect, we used six Philips 12W LED white light bulbs mounted around the experimental setup to provide consistent, non-flickering illumination.
[0057] [Table 1]
[0058] RGB Camera According to an Exemplary Embodiment Two Intel RealSense D435s RGB cameras 310, 312 were used to provide additional non-event image data (the infrared emitter was disabled as it increased noise on the event camera and therefore depth data was not recorded). The first camera 310 was mounted on the end effector, with the camera 310 pointed towards the gripper 302 (providing a view of the grasped object) and the second camera 312 positioned to provide a view of the scene. The RGB images were used for visualization and validation but were not used as input to the model; integration of these standard sensors to provide even better model performance may be provided according to various exemplary embodiments. OptiTrack according to an exemplary embodiment An OptiTrack motion capture system 314 was used to collect object movement data for the slip detection experiment. Six reflective markers were attached to the rigid portion of the end effector, and 14 markers were attached to the object of interest. Eleven OptiTrack Prime 13 cameras were strategically placed around the experimental area to minimize tracking error (e.g., see 316 and 318 in Figure 3b). In all cases, each marker was visible to most, if not all, cameras. This ensured continuous and reliable tracking. Motive Body v1.10.0 was used for marker tracking, and detected markers were manually annotated. Initial testing showed that the OptiTrack system 314 provided reliable position estimates with an error of <1 mm at 120 Hz.
[0059] 3D-Printed Parts for Use in Exemplary Embodiments In one embodiment, the visual-tactile sensor components are attached to the robot via 3D-printed parts. In this example embodiment, there are three main 3D-printed parts: a main holder (Figure 10) for attaching the Intel Realsense D435, Prophesee Onboard, and ACES encoder to the Franka Fmika Panda arm; an enclosure for the ACES encoder (Figure 11); and a coupler (Figure 12) for attaching the NuTouch finger to the Robotiq 2F-140. All of the 3D-printed parts were printed using acrylonitrile butadiene styrene (ABS) with a layer thickness of 0.2 mm. By maximizing the inclusion of only a select few components, the overall weight was minimized while maintaining structural integrity.
[0060] Specifically, in Figure 10, the 3D printed main holder 1000 has four parts: a) a semicircular arc (infill 99%) for securing the main holder to the seventh link of the panda arm, b) a connector (infill 99%) for attaching the sensor to the panda, c) a base (infill 80%) for attaching the ACES encoder enclosure, and d) a holder for the Intel RealSense D435 and Prophesee Onboard (infill 80%).
[0061] Referring to Figures 11 and 12, the enclosure 1200 for the ACES encoder is designed to have a 65% infill, and the coupler for the NuTouch is designed to have a 99% infill.
[0062] Further Details According to Exemplary Embodiments In addition to the above sensors, proprioceptive data was also collected for the Panda arm 300 and Robotiq gripper 302; these are not currently used in the model, but may be included in different exemplary embodiments.
[0063] Minimizing phase shifts is important to enable machine learning models to learn meaningful interactions between different modalities. The setup according to the exemplary embodiment spanned multiple machines, each with its own real-time clock (RTC). Chronyd was used to synchronize the various clocks to Google Public NTP Pool time servers. During data collection, for each machine, the recording start time was recorded according to its own RTC, so that differences between different RTCs could be detected and synchronized accordingly during data preprocessing.
[0064] During the data collection procedure, rotational slips typically occurred midway through the recording. To extract the relevant location when a slip occurred, the slip onset was first detected and annotated. OptiTrack markers were attached to the Panda's end effector and the object so that OptiTrack could determine their pose. Figure 13 shows a visualization of the OptiTrack data for a typical slipping data point. The OptiTrack frame fua was annotated when the robot first picked up the object using the following heuristics:
[0065]
number
[0066] When the robot arm is stationary, p z is f 1;:::;120 The empirical noise distribution within the noise distribution was checked for deviations.
[0067] For object orientation, θ t =cos -1 (2 <q0,q t > 2 -1) The change in angle from rest calculated using where q0 is the quaternion orientation at rest. Similarly, the frame f when the object first rotates slipwas annotated using the following heuristics:
[0068]
number
[0069] The time it took for the object to rotate during lift was found to be an average of 0.03 seconds across all slip data points.
[0070] Figure 13a shows the time course of the end effector's p z When the robot arm lifts the object, p z increases. Figure 13b shows that q t Θ calculated between and q0 t (shortest angle in radians), which increases as the object slips. In Figure 13a, the vertical line moves from rest to p z In Fig. 13b, the vertical line indicates the point at which Θ increases significantly from the resting point. t This indicates the point where there is a significant increase. The difference between these data points is 0.03 seconds.
[0071] I. Containers and Weight Classifications According to Exemplary Embodiments The first experiment applied an event-driven sensing framework, including a NewTouch, an onboard camera, and a VT-SNN according to an exemplary embodiment, to classify containers with varying amounts of liquid. The primary objective was to determine whether the multimodal system according to the exemplary embodiment was effective in detecting differences between objects that would have been difficult to separate using a single sensor. Note that the objective was not to derive the best possible classifier; in fact, the experiment did not include proprioceptive data, which would likely have improved results
[11] , and we did not conduct an exhaustive (and computationally expensive) search for the best architecture. Rather, the experiment was designed to study the potential benefits of using both visual and tactile spiking data in a reasonable setting, according to an exemplary embodiment.
[0072] I.1. Methods and Procedures According to Exemplary Embodiments I.1.1. Objects Used According to the Exemplary Embodiment Four different containers were used: an aluminum coffee can, a plastic Pepsi bottle, a cardboard soy milk carton, and a metal tuna can (see Figure 5). These objects had different degrees of hardness, with the soy milk container being the softest and the tuna can being the hardest. Due to size differences, the four containers held a maximum of 250g, 400g, 300g, and 140g, respectively (the tuna can had no lid and was filled with rice to prevent spills and liquid damage. The tuna can was placed open-side down, so the rice was not visible). For each object, data were collected for 0%, 25%, 50%, 75%, and 100%g of the maximum amount. This resulted in 20 object classes, each containing four containers with five different weight levels.
[0073] I.1.2. Robot Motion According to Example Embodiments The robot grasped and lifted each object class 15 times, generating 15 samples per class. Trajectories for each portion of the motion were calculated using a simple Movelt Cartesian Pose Controller
[47] , and the robot gripper was initialized 10 cm above the designated grasp point for each object. The end effector was then moved to the grasp position (2 seconds) and the gripper closed using the Robotiq grasp controller with a force setting of 1 (4 seconds). The gripper then lifted the object by 5 cm (2 seconds) and held it for 0.5 seconds.
[0074] I.1.3. Data Preprocessing According to an Exemplary Embodiment For both modalities, we selected data from the grasping, lifting, and holding phases (corresponding to the window from 2.0 to 8.5 seconds in Figure 4), setting a bin duration of 0.02 seconds (325 bins) and a binning threshold Smin = 2. We used stratified K-folds to create 5 folds, each containing 240 training and 60 test examples of the graded distribution.
[0075] I.1.4. Classification Model Including VT-SNN According to an Exemplary Embodiment The SNN was compared with traditional deep learning methods, specifically a multilayer perceptron (MLP) with gated recurrent units (GRU)
[48] and a 3D convolutional neural network (CNN-3D)
[51] . Each model was trained using (i) haptic data only, (ii) visual data only, and (iii) combined visual-haptic data. Note that the SNN model on the combined data corresponds to the VT-SNN model in the exemplary embodiment. When training with a single modality, either the visual or haptic SNN was used as appropriate. All models were implemented using PyTorch. The SNN was trained with SLAYER to minimize spike count differences
[30] , and the ANN was trained to minimize cross-entropy loss using RMSPROP. All models were trained for 500 epochs.
[0076] I.2. Results and Analysis I.2.1. Model Comparison Including VT-SNN According to an Exemplary Embodiment The test accuracy of the models is summarized in Table 2. The tactile-only modality SNN gives 12% higher accuracy than the visual-only modality. The multimodal VT-SNN variant according to the exemplary embodiment achieves the highest score of 81%, achieving an improvement of over 11% compared to the tactile modality variant. Note that closer inspection of the visual-only modality data showed that (i) the Pepsi bottle was not completely opaque and the water level was observable by Onboard in some trials, and (ii) Onboard was able to see deformation of the object when the gripper was closed, revealing the "fullness" of the softer container. Thus, the visual-only modality results were better than expected.
[0077] [Table 2]
[0078] FIG. 6 provides an instructive example illustrating the benefit of fusing both modalities according to an exemplary embodiment, showing output spikes from different SNN models while grasping a 100% weight coffee can. Weight categories are arranged from 0% to 100% (bottom to top) for each container class. Models trained on tactile and visual data, respectively, in graphs 600 and 602, are agnostic about the container and weight category, respectively. Specifically, it can be seen that tactile model 600 is unable to distinguish between a tuna can and a coffee can. Meanwhile, visual model 602 accurately predicts the container (i.e., coffee can) but is agnostic about the weight category. The combined visual-tactile model according to an exemplary embodiment, graph 604, incorporates information from both modalities and is able to predict the correct class (both container and weight category, i.e., 100% weight coffee can) with high certainty.
[0079] Referring again to Table I, the SNN model performed much better than the ANN (MLP-GRU) model, especially for the combined visual-tactile data. The poor performance was likely due to the relatively long sample duration (325 time steps) and the large number of parameters in the ANN model relative to the size of the dataset.
[0080] I.2.2. Early Classification with VT-SNN According to an Exemplary Embodiment Instead of waiting for all output spikes to accumulate, early classification can be performed based on the number of spikes seen by time t. Figure 7 shows the accuracy over time for each model. While both combined visual-tactile models 700a,b achieve the highest overall accuracy, between 0.5 and 3.0 seconds, both visual models 702a,b were already able to distinguish specific objects. This is likely due to small movements (of the onboard cameras) occurring as the gripper closes, resulting in changes recognized by Onboard. As expected, tactile spikes do not appear until contact with the object at ~2 seconds for both models 704a,b.
[0081] The line in Figure 7 shows the mean test accuracy, and the shaded area shows the standard deviation. Two losses L and L ω have similar "final" accuracy, but from Figure 7, L ω It can be seen that 700b, 702b, and 704b have a significant effect on test accuracy over time compared to 700a, 702a, and 704a. This effect is most clearly seen for the coupled visual-tactile model, where L ω Variant 700b has a similar initial accuracy profile to visual 702a,b, but achieves better performance because tactile information is accumulated over a period of more than 2 seconds.
[0082] II. Rotational Slip Classification by Embodiment In this second experiment, a sensing system according to an exemplary embodiment was used to classify rotational slippage, which is important for stable grasping; stable grasp points can be incorrectly predicted for objects with centers of mass that are not easily determined by vision, such as hammers and other irregularly shaped items. Accurately detecting rotational slippage allows the controller to re-grasp the object and address poor initial grasp positions. However, to be effective, slippage detection must be performed accurately and quickly.
[0083] II.1. Methods and Procedures According to Exemplary Embodiments II.1.1. Objects Used by the Exemplary Embodiment Test objects were constructed using Lego Duplo blocks (see Figure 8) with a hidden 10 g mass in each leg. The "control" object was designed to be balanced at the grasp point. To induce rotational slippage, the object was modified by moving the hidden mass from the right leg to the left. Thus, the stable and unstable objects were visually identical and had the same total weight.
[0084] II.1.2. Robot Operation According to Example Embodiments The robot grasped and lifted both object variants 50 times, generating 50 samples per class. As in previous experiments, the motion trajectories were computed using the Movelt Cartesian Pose Controller
[47] . The robot was instructed to close the object, lift it 10 cm (0.75 seconds) from the table, and hold it for an additional 4.25 seconds. The gripper's gripping force was adjusted to allow for lifting the object and rotational slip for off-center objects (see Figure 8, right).
[0085] II.1.3. Data Preprocessing According to an Exemplary Embodiment Instead of training the model over the entire movement period, we extracted a short period during the lifting phase. The exact onset time was obtained by analyzing the OptiTrack data. Specifically, a baseline orientation distribution (for 1 s or 120 frames) was obtained, and a rotational slip was defined as an orientation greater than (or less than) 98% of the baseline frames lasting four or more consecutive OptiTrack frames. The slip occurred almost instantly during the lifting phase. Because we were interested in fast detection, we extracted a 0.15 s window near the onset of the lift and set a bin duration of 0.001 s (150 bins) with a binning threshold Smin = 1. Again, we used stratified K-folds to obtain five folds, each containing 80 training examples and 20 test examples.
[0086] II.1.4. Classification Model Including VT-SNN According to an Exemplary Embodiment The model setup and optimization procedure was the same as the previous task / experiment, with three slight modifications. First, the binary label output size was reduced to 2. Second, the sequence length of the ANN GRU was set to 150, the number of time bins. Third, the true and false spike counts of the desired SNN were set to 80 and 5, respectively. Again, the SNN and ANN models were compared using (i) tactile data only, (ii) visual data only, and (iii) combined visual-tactile data, including the VT-SNN according to the exemplary embodiment. II.2. Results and Analysis II.2.1. Model Comparison Including VT-SNN According to an Exemplary Embodiment The test accuracies of the models are summarized in Table 3. For both the SNN and the ANN, both the visual and multimodal models achieved 100% accuracy. This suggests that the visual data is highly indicative of slippage, which is not surprising since rotational slippage produces a visually distinguishable signature. The SNN and MLP-GRU using only haptic events achieved 100% accuracy (L w (with) achieves 91% and 87% accuracy.
[0087] [Table 3]
[0088] II.2.2. Early Slip Detection Including VT-SNN According to an Exemplary Embodiment Similar to the previous analysis of initial container classification, slippage test accuracy at different time points is summarized in Figure 9. It can be seen that the object begins to lift at approximately 0.01 seconds, and that at 0.1 seconds, a multimodal VT-SNN 900a,b according to one embodiment is able to perfectly classify the slippage. Again, it can be seen that vision and touch have different accuracy profiles, with the tactile-only classification 902a,b being more accurate than the VT-SNN with spike counts 900a (between 0.01 and 0.05 seconds), and the vision-based classification 904a,b being better than the tactile-based classification 902a,b after ~0.6 seconds.
[0089] For all SNNs, models trained with weighted spike count losses 900b, 902b, and 904b achieve better early classification compared to spike count losses 900a, 902a, and 904a, and note that the early classification accuracy of the VT-SNN with weighted spike count loss 900b achieves essentially the same early classification accuracy as the tactile-based classification with weighted spike count loss 902b. III. Speed and Power Efficiency According to Exemplary Embodiments We compared the inference speed and energy usage of a classification model (using VT-SNN with spike count loss according to the example implementation) on both a GPU (Nvidia GeForce RTX 2080 Ti) and Intel Loihi, keeping in mind that weighted spike count loss should not affect power consumption.
[0090] Specifically, a multimodal VT-SNN was trained using the SLAYER framework, which was run identically on Loihi and via simulation on a GPU. This model is identical to the one described in the previous section, except for two changes: 1) the Loihi neuron model is used instead of the SRM neuron model, and 2) the polarity of the visual output is discarded to reduce the visual input size to a single core on Loihi.
[0091] Both models achieved 100% test accuracy and produced identical results on Loihi and the GPU. All benchmarks were obtained on Loihi using NxSDK version 0.9.5 on a Nahuku 32 board and an Nvidia RTX 2080Ti GPU, respectively.
[0092] The model is tasked to run 1000 forward passes on a GPU with a batch size of 1. A dataset of 1000 samples is obtained by iterating over samples from our test set, with each sample consisting of 0.15 seconds of spike data, binned into 150 time steps of 1 millisecond each.
[0093] Latency measurement: The GPU uses the CPU's system clock to measure the start time of model inference (t start ) and the end time (t end ) and in Loihi, the system clock of the superhost was used. The latency per time step is defined as (t end -t start ) / (1000x150), splitting into 1000 samples with 150 time steps each.
[0094] Power Utilization Measurement: To obtain power utilization on the GPU, we used the approach of
[52] , logging (timestamp, power_draw) pairs with utilities at 200-millisecond intervals using the NVIDIA System Management Interface. Power consumption over the time spent was extracted and averaged to obtain average power consumption under load. To obtain idle power consumption of the GPU, power usage on the GPU was recorded for 15 minutes with no processes running on the GPU, and the power consumption was averaged over the period. We obtained power utilization of the VT-SNN on Loihi using the performance profiling tools available within NxSDK 0.9.5. The model according to the exemplary embodiment is small, occupying less than one chip on a 32-chip Nahuku 32 board. To obtain more accurate power measurements, we replicated the workload 32 times and reported results for each copy. The replicated workload used 594 neuromorphic cores and 5 x86 cores, with 624 neuromorphic cores corresponding to barrier synchronization. To simulate a real-world setting (where data arrives in online order), 1) the x86 cores are artificially slowed down to match the duration of a 1 ms timestep of data, and 2) an artificial delay of 0.15 seconds is introduced in the GPU's dataset fetch, simulating waiting for an entire window of data before being able to perform inference.
[0095] The benchmark results are shown in Table 4, where latency is the time it takes to process one time step. Latency on Loihi is observed to be slightly lower because it can perform inference as spiking data arrives. Power consumption on Loihi is significantly (1900x) lower than that of a GPU.
[0096] [Table 4]
[0097] 14 shows a schematic diagram illustrating a classification sensing system 1400 according to an example embodiment. The system 1400 includes a first spiking neural network (SNN) encoder 1402 configured to encode the event-based output of a visual sensor 1404 into individual visual-modality spiking representations having a first output size, a second SNN encoder 1406 configured to encode the event-based output of a tactile sensor 1408 into individual tactile-modality spiking representations having a second output size, a combining layer 1410 configured to merge the visual and tactile modality spiking representations, and a task SNN 1412 configured to receive the merged visual and tactile modality spiking representations and encode them into visual-tactile modality spiking representations having a third output size for classification.
[0098] Task SNN1412 may be configured for classification based on spike count loss in the respective output visual / tactile modality representations compared to a desired spike count indexed by output size. Preferably, task SNN1412 is configured to classify based on weighted spike count loss in the respective output visual / tactile modality representations compared to a desired weighted spike count indexed by output size.
[0099] Neurons in each of the first SNN encoder 1402, the second SNN encoder 1406, and the task SNN 1412 may be configured to apply a spike response model (SRM).
[0100] The sensor system 1400 may include a tactile sensor 1404. Preferably, the tactile sensor 1404 includes an event-based tactile sensor. Alternatively, the tactile sensor 1404 includes a converter for converting the intrinsic output of the tactile sensor 1404 into the event-based output of the tactile sensor 1404.
[0101] The sensor system 1400 may include a visual sensor 1408. Preferably, the visual sensor 1408 includes an event-based visual sensor. Alternatively, the vision sensor 1408 includes a converter for converting the intrinsic output of the vision sensor to the event-based output of the vision sensor 1408.
[0102] The sensor system 1400 may include a robotic arm and an end effector. The end effector may include a gripper. Preferably, the tactile sensor 1406 may include one tactile element on each finger of the gripper.
[0103] The visual sensor 1408 may be mounted on the robot arm or on the end effector.
[0104] FIG. 15 shows a flowchart 1500 illustrating a classification method performed using a sensing system, according to an exemplary embodiment. In step 1502, the event-based output of a visual sensor is encoded into individual visual modality spiking representations having a first output size using a first spiking neural network (SNN) encoder. In step 1504, the event-based output of a tactile sensor is encoded into individual tactile modality spiking representations having a second output size using a second SNN encoder. In step 1506, the visual modality spiking representations and the tactile modality spiking representations are merged using a combination layer. In step 1508, a task SNN is used to receive combined visual and tactile modality spiking representations, a task SNN is used to receive combined visual and tactile modality spiking representations, and a task SNN is used to output visual-tactile modality spiking representations with a third output size for classification.
[0105] The task SNN may be configured for classification based on spike count loss in the respective output visual / tactile modality representations compared to a desired spike count indexed by output size. Preferably, the task SNN is configured for classification based on weighted spike count loss in the respective output visual / tactile modality representations compared to a desired weighted spike count indexed by output size.
[0106] Each of the first SNN encoder, the second SNN encoder, and the task SNN may be configured to apply a spike response model (SRM).
[0107] Preferably, the tactile sensor comprises an event-based tactile sensor. Alternatively, the tactile sensor comprises a converter for converting the intrinsic output of the tactile sensor into the event-based output of the tactile sensor.
[0108] Preferably, the visual sensor comprises an event-based visual sensor. Alternatively, the visual sensor comprises a converter for converting an intrinsic output of the visual sensor into an event-based output of the visual sensor.
[0109] The method can include placing one tactile element of the tactile sensor on each finger of a gripper of the robotic arm.
[0110] The method may include mounting a visual sensor on the robot arm or on the end effector.
[0111] FIG. 16 shows a schematic diagram illustrating a tactile sensor 1600 comprising a carrier structure 1602 and an electrode layer 1604 disposed on a surface of the carrier structure 1602, the electrode array 1604 comprising an array of taxel electrodes, e.g., 1606, and a plurality of electrode wires, e.g., 1608, individually electrically connected to each one of the taxel electrodes, e.g., 1602; and a protective layer 1610 disposed on the electrode layer 1604, the protective layer 1610 being made of an elastically deformable material, and a pressure transducer layer 1612 disposed between the electrode layer 1604 and the protective layer 1610, wherein electrical signals detectable in the electrode wires, e.g., 1608, in response to a contact force applied to the pressure transducer layer 1612 through the protective layer 1610 provide spatiotemporal data for neuromorphic tactile sensing applications.
[0112] The taxel electrodes of the electrode array, e.g., 1606, may be arranged in a density that varies radially around the center of the electrode array. The density of the taxel electrodes, e.g., 1606, may decrease with radial distance from the center.
[0113] The tactile sensor may include multiple encoder elements, e.g., 1614, connected to respective ones of the electrode lines, e.g., 1608, and the decoder elements, e.g., 1614, are configured to asynchronously transmit tactile information based on the electrical signals in the electrode lines, e.g., 1608, via a common output conductor 1616.
[0114] The carrier structure 1602 can be configured to be connectable to a robotic gripper.
[0115] The electrode layer 1604 and / or the electrode lines, eg, 1608, may be flexible.
[0116] FIG. 17 shows a flowchart 1700 illustrating a method for manufacturing a tactile sensor, according to an exemplary embodiment. In step 1702, a carrier structure is provided. In step 1704, an electrode layer is disposed on a surface of the carrier structure, the electrode array including an array of taxel electrodes. In step 1706, a plurality of electrode wires are provided, each individually electrically connected to each of the taxel electrodes. In step 1708, a protective layer is disposed over the electrode layer, the protective layer being made of an elastically deformable material. In step 1710, a pressure converter layer is disposed between the electrode layer and the protective layer, where detectable electrical signals in the electrode wires responsive to contact forces applied to the pressure converter layer through the protective layer provide spatiotemporal data for neuromorphic tactile sensing applications.
[0117] The taxel electrodes of the electrode array may be arranged at a density that varies radially around the center of the electrode array, and the density of the taxel electrodes may decrease with radial distance from the center.
[0118] The method can include providing a plurality of encoder elements connected to respective electrode lines, and configuring the decoder elements to asynchronously transmit tactile information based on electrical signals in the electrode lines via a common output conductor.
[0119] The method may include configuring the carrier structure to be connectable to a robotic gripper.
[0120] The electrode layer and / or the electrode lines may be flexible.
[0121] As described above, an event-based sensing framework is provided by exemplary embodiments that combine vision and touch to achieve better performance in two robotic tasks. In contrast to traditional synchronous systems, the event-driven framework according to exemplary embodiments can process discrete events asynchronously, thus achieving higher time resolution and lower latency with lower power consumption.
[0122] We have described NuTouch, a neuromorphic event tactile sensor according to an exemplary embodiment, and VT-SNN, a multimodal spiking neural network that learns from raw unstructured event data according to an exemplary embodiment. Experimental results on container & weight classification and rotational slippage detection showed that combining both modalities according to an exemplary embodiment is key to achieving high accuracy.
[0123] Embodiments of the invention may have one or more of the following features and associated benefits / advantages.
[0124] [Table 5]
[0125] The various functions or processes disclosed herein may be described in terms of their behavior, register transfers, logical components, transistors, layout geometry, and / or other characteristics as data and / or instructions embodied in various computer-readable media. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media, or any combination thereof. Examples of transfer of such formatted data and / or instructions via a carrier wave include, but are not limited to, transfer (upload, download, email, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.).
[0126] Aspects of the systems and methods described herein can be implemented as functionality programmed into any of a variety of circuits, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices, and standard cell-based devices, as well as application specific integrated circuits (ASICs). Some other possibilities for implementing aspects of the system include microcontrollers with memory (such as electronically erasable programmable read-only memory (EEPROM)), embedded microprocessors, firmware, software, etc. Additionally, aspects of the system can be embodied in microprocessors with software-based circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and hybrids of any of the above device types. Of course, the underlying device technology can be provided in a variety of component types, e.g., metal-oxide field effect transistor (MOSFET) technologies such as complementary metal-oxide semiconductor (CMOS), bipolar technologies such as emitter coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal-conjugated polymer-metal structures), mixed analog and digital, etc.
[0127] The above description of illustrated embodiments of the systems and methods is not intended to be exhaustive or to limit the systems and methods to the precise forms disclosed. While specific embodiments and examples of system components and methods are described herein for illustrative purposes, those skilled in the art will recognize that various equivalent modifications are possible within the scope of the systems, components, and methods. The teachings of the systems and methods provided herein may be applied to other processing systems and methods, as well as the systems and methods described above.
[0128] Those skilled in the art will appreciate that numerous variations and / or modifications can be made to the present invention as illustrated in the specific embodiments without departing from the spirit or scope of the invention as broadly described. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive. The present invention also includes any combination of features described in the different embodiments, including in the summary section, even if such feature or combination of features is not explicitly specified in the claims or the detailed description of the present embodiments.
[0129] In general, in the following claims, the terms used should not be construed to limit the systems and methods to the specific embodiments disclosed in the specification and claims, but should be construed to include all processing systems that operate within the scope of the claims. Accordingly, the systems and methods are not limited by this disclosure; instead, the scope of the systems and methods should be determined entirely by the claims.
[0130] Unless the context clearly requires otherwise, throughout this specification and claims, the words "comprise," "comprising," and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; i.e., words using the singular or plural number in the sense of "including, but not limited to," also include the plural or singular, respectively. Furthermore, the terms "herein," "below," "above," "below," and words of similar import refer to this application as a whole and not to any particular portion of this application. When the word "or" is used in connection with a list of two or more items, the word includes all of the following interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list.
[0131] (References) [1] A. Billard and D. Kragic, “Trends and challenges in robot manipulation,” Science, vol. 364, no. 6446, p. eaat8414, 2019. [2] D. Li, X. Chen, M. Becchi, and Z. Zong, “Evaluating the energy efficiency of deep convolutional neural networks on cpus and gpus,” 102016, pp. 477-484. [3] E. Strubell, A. Ganesh, and A. McCallum, “Energy and policy considerations for deep learning in NLP,” in Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, 2019, pp. 3645-3650. [Online]. Available: https: / / doi.org / 10.18653 / vl / pl9-1355 [4] M. Pfeiffer and T. Pfeil, “Deep Learning With Spiking Neurons: Opportunities and Challenges,” Frontiers in Neuroscience, vol. 12, no. October, 2018. [5] S.-C. Liu, B. Rueckauer, E. Ceolini, A. Huber, and T. Delbruck, “Eventdriven sensing for efficient perception: Vision and audition algorithms,” IEEE Signal Processing Magazine, vol. 36, no. 6, pp. 29-37, 2019. [6] Y. A. LeCun, Y. Bengio, and G. E. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436-444, 2015. [7] M. Davies, N. Srinivasa, T. Lin, G. Chinya, Y. Cao, S. H. Choday, G. Dimou, P. Joshi, N. Imam, S. Jain, Y. Liao, C. Lin, A. Lines, R. Liu, D. Mathaikutty, S. McCoy, A. Paul, J. Tse, G. Venkataramanan, Y.Weng, A. Wild, Y. Yang, and H. Wang, “Loihi: A neuromorphic manycore processor with on-chip learning,” IEEE Micro, vol. 38, no. 1, pp. 82- 99, January 2018. [8] J. Sinapov, C. Schenck, and A. Stoytchev, “Learning relational object categories using behavioral exploration and multimodal perception,” in 2014 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2014, pp. 5691-5698. [9] Y. Gao, L. A. Hendricks, K. J. Kuchenbecker, and T. Darrell, “Deep learning for tactile understanding from visual and haptic data,” in 2016 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2016, pp. 536-543.
[10] J. Li, S. Dong, and E. Adelson, “Slip detection with combined tactile and visual information,” in 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 7772-7777.
[11] M. A. Lee, Y. Zhu, K. Srinivasan, P. Shah, S. Savarese, L. Fei- Fei, A. Garg, and J. Bohg, “Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 8943-8950.
[12] J. Lin, R. Calandra, and S. Levine, “Learning to identify object instances by touch: Tactile recognition via multimodal matching,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 3644-3650.
[13] H. Liu, F. Sun et ah, “Robotic tactile perception and understanding,” 2018.
[14] P. Allen, “Surface descriptions from vision and touch,” in Proceedings. 1984 IEEE International Conference on Robotics and Automation, vol. 1. IEEE, 1984, pp. 394-397.
[15] S. Luo, J. Bimbo, R. Dahiya, and H. Liu, “Robotic tactile perception of object properties: A review,” Mechatronics, vol. 48, pp. 54-67, 2017.
[16] H. Liu, Y. Yu, F. Sun, and J. Gu, “Visual-tactile fusion for object recognition,” IEEE Transactions on Automation Science and Engineering, vol. 14, no. 2, pp. 996-1008, 2016.
[17] H. Soh, Y. Su, and Y. Demiris, “Online spatio-temporal Gaussian process experts with application to tactile classification,” in Intelligent Robots and Systems (IROS), 2012 IEEE / RSJ International Conference on. IEEE, 2012, pp. 4489-4496.
[18] J. Varley, D. Watkins, and P. Allen, “Visual-tactile geometric reasoning,” in RSS Workshop, 2017.
[19] J. Reinecke, A. Dietrich, F. Schmidt, and M. Chalon, “Experimental comparison of slip detection strategies by tactile sensing with the biotac(R) on the dir hand arm system,” in 2014 IEEE international Conference on Robotics and Automation (ICRA). IEEE, 2014, pp. 2742-2748.
[20] Y. Bekiroglu, R. Detry, and D. Kragic, “Learning tactile characterizations of object-and pose-specific grasps,” in 2011 IEEE / RSJ international conference on Intelligent Robots and Systems. IEEE, 2011, pp. 1554- 1560.
[21] Z. Su, K. Hausman, Y. Chebotar, A. Molchanov, G. E. Loeb, G. S. Sukhatme, and S. Schaal, “Force estimation and slip detection / classification for grip control using a biomimetic tactile sensor,” in 2015 IEEE-RAS 15th International Conference on Humanoid Robots (Humanoids). IEEE, 2015, pp. 297-303.
[22] W. Yuan, S. Dong, and E. H. Adelson, “Gelsight: High-resolution robot tactile sensors for estimating geometry and force,” Sensors, vol. 17, no. 12, p. 2762, 2017.
[23] R. Calandra, A. Owens, D. Jayaraman, J. Lin, W. Yuan, J. Malik, E. H. Adelson, and S. Levine, “More than a feeling: Learning to grasp and regrasp using vision and touch,” IEEE Robotics and Automation Letters, vol. 3, no. 4, pp. 3300-3307, 2018.
[24] S. Luo, W. Yuan, E. Adelson, A. G. Cohn, and R. Fuentes, “Vitae: Feature sharing between vision and tactile sensing for cloth texture recognition,” in 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 2722-2727.
[25] G. Gallego, T. Delbr, G. Orchard, C. Bartolozzi, B. Taba, A. Censi, K. Daniilidis, D. Scaramuzza, S. Leutenegger, and A. Davison, “Eventbased Vision : A Survey,” Tech. Rep., 2018.
[26] A. Mitrokhin, C. Ye, C. Fermuller, Y. Aloimonos, and T. Delbruck, “EVIMO: Motion Segmentation Dataset and Learning Pipeline for Event Cameras,” in 2019 IEEE / RSI International Conference on Intelligent Robots and Systems (IROS), 2019.
[27] A. Z. Zhu and L. Yuan, “EV-FlowNet: Self-Supervised Optical Flow Estimation for Event-based Cameras,” in Robotics: Science and Systems, 2018.
[28] A. I. Maqueda, A. Loquercio, G. Gallego, N. Garcn’nia, and D. Scaramuzza, “Event-based vision meets deep learning on steering prediction for self-driving cars,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5419-5427.
[29] A. Tavanaei, M. Ghodrati, S. R. Kheradpisheh, T. Masquelier, and A. Maida, “Deep learning in spiking neural networks,” Neural Networks, vol. I l l, pp. 47-63, 2019. [Online]. Available: https: / / doi.org / 10.1016 / j.neunet.2018.12.002
[30] S. B. Shrestha and G. Orchard, “Slayer: Spike layer error reassignment in time,” in Advances in Neural Information Processing Systems, 2018, pp. 1412-1421.
[31] G. Bellec, F. Scherr, E. Hajek, D. Salaj, R. Legenstein, and W. Maass, “Biologically inspired alternatives to backpropagation through time for learning in recurrent neural nets,” arXiv preprint arXiv: 1901.09049, 2019.
[32] M. Akrout, C. Wilson, P. Humphreys, T. Lillicrap, and D. B. Tweed, “Deep learning without weight transport,” in Advances in Neural Information Processing Systems, 2019, pp. 974-982.
[33] P. A. Merolla, J. V. Arthur, R. Alvarez-Icaza, A. S. Cassidy, J. Sawada, F. Akopyan, B. L. Jackson, N. Imam, C. Guo, Y. Nakamura, B. Brezzo, I. Vo, S. K. Esser, R. Appuswamy, B. Taba, A. Amir, M. D. Flickner, W. P. Risk, R. Manohar, and D. S. Modha, “A million spiking- 980 neuron integrated circuit with a scalable communication network and interface,” Science, vol. 345, no. 6197, pp. 668-673, 2014. [Online]. Available: https: / / science.sciencemag.org / content / 345 / 6197 / 668
[34] S. Chevallier, H. Paugam-Moisy, and F. Lem ah t re, “Distributed processing for modelling real-time multimodal perception in a virtual robot.” in Parallel and Distributed Computing and 985 Networks, 2005, pp. 393-398.
[35] N. Rathi and K. Roy, “Stdp-based unsupervised multimodal learning with cross-modal processing in spiking neural network,” IEEE Transactions on Emerging Topics in Computational Intelligence, pp. 1-11, 2018.
[36] E. Mansouri-Benssassi and I. Ye, “Speech emotion recognition with early visual cross- 990 modal enhancement using spiking neural networks,” in 2019 International loint Conference on Neural Networks (IJCNN). IEEE, 2019, pp. 1-8.
[37] T. Zhou and I. P. Wachs, “Spiking neural networks for early prediction in human-robot collaboration,” The International Journal of Robotics Research, vol. 38, no. 14, pp. 1619-1643, 2019. [Online] Available: https: / / doi.org / 10.1177 / 0278364919872252 995
[38] J. Konstantinova, A. Jiang, K. Althoefer, P. Dasgupta, and T. Nanayakkara, “Implementation of tactile sensing for palpation in robot-assisted minimally invasive surgery: A review,” IEEE Sensors Journal, vol. 14, no. 8, pp. 2490-2501, 2014.
[39] Y.Wu, Y. Liu, Y. Zhou, Q. Man, C. Hu,W. Asghar, F. Li, Z. Yu, J. Shang, G. Liu et ah, “A skin-inspired tactile sensor for smart prosthetics,” Science Robotics, vol. 3, no. 22, p. 1000 eaat0429, 2018.
[40] Q.-J. Sun, X.-H. Zhao, Y. Zhou, C.-C. Yeung, W. Wu, S. Venkatesh, Z.-X. Xu, J. J. Wylie, W.-J. Li, and V. A. Roy, “Fingertip-skin-inspired highly sensitive and multifunctional sensor with hierarchically structured conductive graphite / polydimethylsiloxane foams,” Advanced Functional Materials, vol. 29, no. 18, p. 1808829, 2019. 1005
[41] J. He, P. Xiao, W. Lu, J. Shi, L. Zhang, Y. Liang, C. Pan, S.-W. Kuo, and T. Chen, “A universal high accuracy wearable pulse monitoring system via high sensitivity and large linearity graphene pressure sensor,” Nano Energy, vol. 59, pp. 422-433, 2019.
[42] T. Callier, A. K. Suresh, and S. J. Bensmaia, “Neural coding of contact events in somatosensory cortex,” Cerebral Cortex, vol. 29, no. 11, pp. 4613-4627, 2019. 1010
[43] W. W. Lee, Y. J. Tan, H. Yao, S. Li, H. H. See, M. Hon, K. A. Ng, B. Xiong, J. S. Ho, and B. C. Tee, “A neuro-inspired artificial peripheral nervous system for scalable electronic skins,” Science Robotics, vol. 4, no. 32, p. eaax2198, 2019.
[44] R. S. Johansson and J. R. Flanagan, “Coding and use of tactile signals from the fingertips in object manipulation tasks,” Nature Reviews Neuroscience, vol. 10, no. 5, pp. 345-359, 2009. 1015
[45] W. Gerstner, “Time structure of the activity in neural network models,” Physical review E, vol. 51, no. 1, p. 738, 1995.
[46] B. Calli, A. Walsman, A. Singh, S. Srinivasa, P. Abbeel, and A. M. Dollar, “Benchmarking in manipulation research: Using the yale-cmuberkeley object and model set,” IEEE Robotics Automation Magazine, vol. 22, no. 3, pp. 36-52, Sep. 2015. 1020
[47] D. Coleman, I. Sucan, S. Chitta, and N. Correll, “Reducing the barrier to entry of complex robotic software: a moveit! case study,” arXiv preprint arXiv: 1404.3785, 2014.
[48] K. Cho, B. van Mem 'enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural 1025 Language Processing (EMNLP), 2014, pp. 1724-1734.
[49] P. Blouw, X. Choo, E. Hunsberger, and C. Eliasmith, “Benchmarking keyword spotting efficiency on neuromorphic hardware,” 2018, arXiv: 1812.01739.
[50] Lee, Wang Wei, et al. "A neuro-inspired artificial peripheral nervous system for scalable electronic skins." Science Robotics 4.32 (2019): eaax2198. 1030
[51] J. M. Gandarias, F. Pastor, A. J. Garc ia-Cerezo, and J. M. G'omezde Gabriel, “Active tactile recognition of deformable objects with 3d convolutional neural networks,” in 2019 IEEE World Haptics Conference (WHC). IEEE, 2019, pp. 551-555.
[52] P. Blouw, X. Choo, E. Hunsberger, and C. Eliasmith, “Benchmark-ing keyword spotting efficiency on neuromorphic hardware,” 2018, arXiv: 1812.01739]
Claims
1. a first spiking neural network (SNN) encoder configured to encode the event-based output of the visual sensor into individual visual modality spiking representations having a first output size; a second SNN encoder configured to encode the event-based output of the tactile sensor into an individual tactile modality spiking representation having a second output size; a combination layer configured to merge the visual modality spiking representation and the haptic modality spiking representation; a task SNN configured to receive the merged visual modality spiking representation and the haptic modality spiking representation and to output a visual-haptic modality spiking representation having a third output size; Equipped with the first output size is smaller than the input size of the first SNN encoder; The second output size is smaller than the input size of the second SNN encoder. Classification sensing system.
2. The task SNN is configured for classification based on spike count loss in each visual modality representation and each tactile modality representation compared to a desired spike count indexed by the output size. The system of claim 1 .
3. The task SNN is configured for classification based on weighted spike count loss in the respective visual modality representations and the respective tactile modality representations compared to desired weighted spike counts indexed by the output size. The system of claim 1 .
4. Neurons in each of the first SNN encoder, the second SNN encoder, and the task SNN are configured to apply a spike response model (SRM). A system according to any one of claims 1 to 3.
5. The tactile sensor A system according to any one of claims 1 to 4.
6. The tactile sensor comprises an event-based tactile sensor. The system of claim 5.
7. The tactile sensor includes a converter for converting an intrinsic output of the tactile sensor to an event-based output of the tactile sensor. The system of claim 5.
8. The visual sensor A system according to any one of claims 1 to 7.
9. The visual sensor comprises an event-based visual sensor. The system of claim 8.
10. The visual sensor includes a converter for converting an intrinsic output of the visual sensor to an event-based output of the visual sensor. The system of claim 8.
11. Equipped with a robot arm and end effector 11. A system according to any one of claims 1 to 10.
12. The end effector comprises a gripper The system of claim 11.
13. The tactile sensor comprises one tactile element for each finger of the gripper. The system of claim 12.
14. The visual sensor is mounted on the robot arm or on the end effector.
14. A system according to any one of claims 11 to 13.
15. 1. A classification method performed using a sensing system, comprising: encoding, using a first spiking neural network (SNN) encoder, the event-based output of the visual sensor into individual visual modality spiking representations having a first output size; encoding the event-based output of the tactile sensor into an individual tactile modality spiking representation having a second output size using a second SNN encoder; merging the visual modality spiking representation and the tactile modality spiking representation using a combination layer; receiving, using a task SNN, the merged visual modality spiking representation and the haptic modality spiking representation, and outputting a visual-haptic modality spiking representation having a third output size for classification; Including, the first output size is smaller than the input size of the first SNN encoder; The second output size is smaller than the input size of the second SNN encoder. Classification method.
16. The task SNN is configured for classification based on spike count loss in each visual modality representation and each tactile modality representation compared to a desired spike count indexed by the output size.
16. The method of claim 15.
17. The task SNN is configured to classify based on a weighted spike count loss in each visual modality representation and each tactile modality representation compared to a desired weighted spike count indexed by the output size.
17. The method of claim 16.
18. Each of the first SNN encoder, the second SNN encoder, and the task SNN is configured to apply a spike response model (SRM).
18. The method of any one of claims 15 to 17.
19. The tactile sensor includes an event-based tactile sensor.
19. The method of any one of claims 15 to 18.
20. converting the intrinsic output of the tactile sensor into an event-based output of the tactile sensor.
19. The method of any one of claims 15 to 18.
21. The visual sensor includes an event-based visual sensor.
21. The method of any one of claims 15 to 20.
22. converting the intrinsic output of the visual sensor into an event-based output of the visual sensor.
21. The method of any one of claims 15 to 20.
23. and disposing one tactile element of said tactile sensor on each finger of a gripper of a robotic arm.
23. The method of any one of claims 15 to 22.
24. and attaching the visual sensor to the robot arm or end effector.
24. The method of any one of claims 15 to 23.