Manipulator multi-mode tactile perception and object recognition system and method
By using a multi-channel flexible tactile sensor array and a multi-task timing model, the problem of the robotic arm's inability to distinguish objects and accurately predict weight under complex surface conditions was solved, achieving synchronization and robustness improvement of texture perception and weight prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI UNIV OF ENG SCI
- Filing Date
- 2026-03-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing robotic gripping systems struggle to reliably distinguish different objects and accurately predict their weight under complex surface conditions, and traditional texture sensing schemes increase system complexity.
A multi-channel flexible tactile sensor array is employed, including an encapsulation layer, an electrode layer, a conductive sensitive layer, and an elastomer layer. The microstructure array enables multimodal fusion of pressure, strain, and texture-related signals, and is combined with a multi-task temporal model for object category recognition and weight prediction.
Texture perception and differentiation can be achieved without additional texture sensors during normal grasping, which improves the robustness of recognition and the accuracy of weight prediction under complex surface conditions, reduces hardware and control complexity, and realizes multi-task synchronous output in a single grasping operation.
Smart Images

Figure CN121973235A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tactile perception and flexible sensing technology for intelligent robots, and in particular to a multimodal tactile perception and object recognition system and method for a robotic arm. Background Technology
[0002] In grasping and manipulating tasks, robotic arms need to sense information such as contact pressure, joint deformation, and friction / micro-slippage generated by contact with the object surface in order to achieve stable grasping, grasping status assessment, and object property perception. Existing grasping sensing systems mostly use pressure sensors or bending / strain sensors to collect single or a small number of modal signals, and then combine them with machine learning for object recognition or weight estimation.
[0003] When it comes to distinguishing surface textures, traditional methods often rely on additional texture-specific sensors or add active sweeping / scanning actions to excite texture signals, and then perform frequency domain / time-frequency domain analysis on high-frequency components to complete the identification.
[0004] However, the surface texture of actual objects to be grasped varies, including smoothness, roughness, and fabric texture. These texture differences lead to significant variations in contact friction, micro-slippage, and signal fluctuation characteristics. Relying solely on pressure magnitude or joint posture signals often fails to reliably distinguish different objects under complex surface conditions, and weight prediction is easily affected by changes in the grasping contact state.
[0005] Traditional texture-aware solutions typically rely on additional texture-specific sensors or active sweeping actions to elicit texture signals, which increases system complexity and makes them unsuitable for deployment in regular grasping operations. Summary of the Invention
[0006] The purpose of this invention is to provide a multimodal tactile perception and object recognition system and method for robotic arms, enabling the sensor to simultaneously obtain pressure response, deformation strain response, and texture-related fluctuation signals caused by natural micro-friction / micro-slippage during conventional grasping processes without the need for an additional texture sensor; and through multimodal feature fusion and a multi-task temporal model, to achieve synchronous output of object category recognition results and weight prediction values under a single grasping temporal input, thereby improving the robustness of recognition under complex surface conditions and reducing weight prediction errors.
[0007] The objective of this invention can be achieved through the following technical solutions: A multimodal tactile sensing and object recognition system for a robotic arm includes a robotic arm actuator, a multi-channel flexible tactile sensor array, a signal acquisition module, and a processing terminal. The flexible tactile sensor in the multi-channel flexible tactile sensor array includes an encapsulation layer, an electrode layer, a conductive sensitive layer, and an elastomer layer. The multi-channel flexible tactile sensor array contacts the target object. The encapsulation layer and / or the elastomer layer are provided with a microstructure array on the side facing the conductive sensitive layer, and the microstructure array and the conductive sensitive layer form an interlocking / interlocking interface.
[0008] Furthermore, the multi-channel flexible tactile sensor array also includes a strain channel group and a pressure channel group, with the flexible tactile sensor connected to the strain channel group and the pressure channel group.
[0009] Furthermore, the strain channel group and the pressure channel group are connected together to the flexible printed circuit board (FPC), the FPC is connected to the signal acquisition module, the signal acquisition module is connected to the processing terminal, and a multi-task timing model is set in the processing terminal.
[0010] Furthermore, the microstructure array is a periodic array structure.
[0011] Furthermore, the microstructure array can be any one of the following: micro pyramid array, micro needle array, micro pillar array, hemispherical array, or biomimetic texture array.
[0012] Furthermore, the conductive sensitive layer is a conductive hydrogel sensitive layer.
[0013] A multimodal tactile perception and object recognition method for a robotic arm, employing the aforementioned system, includes the following steps: S1 controls the robotic arm actuator to execute a sequence of grasping actions including initial homing, closed gripping, stable holding, and release reset, and acquires multi-channel raw resistance timing signals during the grasping process based on the signal acquisition module; S2 preprocesses the multi-channel raw resistance timing signal in the processing terminal to obtain the preprocessed resistance timing signal. S3 extracts pressure features, strain features, and texture-related features from the preprocessed resistance time-series signal and fuses them to form a fused feature set; The S4 fusion feature set is input into a pre-trained multi-task temporal model, which outputs object category recognition results and weight prediction values.
[0014] Furthermore, texture-related features include one or more of the following: high-frequency fluctuation features of the pressure signal, frequency domain energy features, and time-frequency domain features.
[0015] Furthermore, the specific steps for fusing features to form a fused feature set are as follows: The pressure features, strain features, and texture-related features are aligned and fused according to a unified time window to obtain a fused feature set.
[0016] Furthermore, the preprocessing includes synchronous alignment, obtaining the baseline of the static phase before the start of the grabbing and performing baseline correction, normalization and filtering for noise reduction, outlier removal and drift correction.
[0017] Compared with the prior art, the present invention has the following beneficial effects: (1) Texture information can be obtained without additional texture sensors: the surface texture difference is mapped to the texture-related dynamic fluctuation component in the resistance timing signal through the microstructure interlock / hook interface, and texture perception and differentiation can be achieved in the normal grasping process, reducing hardware and control complexity.
[0018] (2) Multimodal fusion improves robustness: The fusion of pressure, strain and texture-related features can reduce the influence of different friction conditions and contact stability changes on recognition and regression, improve the stability of category recognition under complex surface conditions and reduce weight prediction error.
[0019] (3) Single-time grasping synchronous output of multi-task results: The multi-task temporal model can synchronously output the object category and weight prediction results under the same grasping temporal input, and can extend the output of surface morphology recognition results to improve system integration and inference efficiency. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the overall architecture of the bionic hand grasping perception and recognition system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the microstructure interlocking flexible resistance sensor structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the integrated arrangement of the sensor array on the bionic hand according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the connection of the pressure / tactile channel group through FPC in an embodiment of the present invention; Figure 5 This is a schematic diagram of the connection of the strain channel group led out by FPC according to an embodiment of the present invention; Figure 6 This is a schematic diagram showing the correspondence between the acquisition channels and the sensor unit numbers in an embodiment of the present invention; Figure 7 This is a schematic diagram of the standardized grasping experimental process and grasping action sequence according to an embodiment of the present invention; Figure 8 This is a comparison diagram of the texture-related resistance fluctuation response of different surface texture materials in the embodiments of the present invention.
[0021] Figure 9 This is a schematic diagram of the overall process of data acquisition, signal processing, and recognition prediction of the system of the present invention. In the figure: 1 Encapsulation layer; 2 Electrode layer; 3 Conductive sensitive layer; 4 Elastomer layer; 5 Target object; 6 Robotic actuator; 7 Pressure channel group; 8 Strain channel group; 9 FPC; 10 Signal acquisition module; 11 Processing terminal; 12 Multi-task timing model; 13 Microstructure array. Detailed Implementation
[0022] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0023] The system of this invention includes a robotic actuator, a flexible tactile sensor, a signal acquisition and processing unit, and a machine learning recognition and prediction unit. If necessary, it may also include a control unit for stimulating texture fluctuations, as detailed below: (1) Flexible tactile sensor structure The flexible tactile sensor is a resistive flexible sensor, comprising an encapsulation layer 1, an electrode layer 2, a conductive sensitive layer 3, and an elastomer layer 4. A microstructure array 13 is disposed on the side of the encapsulation layer 1 and / or the elastomer layer 4 facing the conductive sensitive layer 3, and the microstructure array is bonded to the conductive sensitive layer to form an interlocking / fitting interface. The electrode layer 2 is electrically connected to the conductive sensitive layer 3 to form a resistive acquisition channel.
[0024] The microstructure array can be a periodic array structure, preferably a microcone array, or at least one of a micro pyramid array, microneedle array, microcolumn array, hemispherical array, or biomimetic texture array; in a preferred embodiment, the height of the microstructure can be 100–800 μm, and the array spacing can be 150–1200 μm.
[0025] The conductive sensitive layer is preferably a conductive hydrogel sensitive layer, but other flexible conductive sensitive material layers that satisfy the resistance change output mechanism can also be used.
[0026] (2) Multimodal response mechanism During the grasping and contact process, the output resistance timing signal of the resistive flexible tactile sensor can simultaneously reflect: a) Stress response caused by normal loading; b) Strain response caused by bending or tensile deformation; c) Texture-related dynamic fluctuation response caused by micro-friction / micro-slippage at the contact interface.
[0027] Among them, texture-related dynamic fluctuations originate from the rapid micro-deformation and contact state changes generated by the microstructure interlocking / interlocking interface under the action of micro-slippage and micro-friction, which in turn form distinguishable dynamic fluctuation components in the resistance timing signal.
[0028] (3) Sensor array arrangement and channel division The flexible tactile sensor can be attached to the gripping surface and / or the dorsal region of the joint of the robotic hand. In one embodiment, the pressure channel group (7) is arranged in the contact areas such as the fingertips, finger pads, and / or palm, and the strain channel group (8) is arranged in the dorsal region of the joint or the area with significant deformation, so as to simultaneously obtain pressure-related signals and strain-related signals during a single gripping process, and the pressure-related signals carry texture-related dynamic fluctuation information. The number of channels can be flexibly configured according to the degrees of freedom and coverage area of the robotic hand; for example, it can be 23 channels in one embodiment, but is not limited to this.
[0029] (4) Electrical connection and data acquisition interface Each channel can be converged and led out through a flexible printed circuit board (FPC) (9) and connected to the signal acquisition module (10) to achieve multi-channel synchronous sampling and stable electrical connection. The processing terminal (11) can establish and store the correspondence between the acquisition channel and the sensor unit number / installation location to realize data traceability management and label management required for subsequent model training.
[0030] (5) Signal acquisition and processing unit The signal acquisition and processing unit synchronously samples the multi-channel resistance timing signal and performs alignment, baseline correction, and normalization processing; normalization can be based on the baseline resistance during the static phase before the capture begins. The processing unit can also selectively perform at least one of filtering and noise reduction, outlier removal, and drift correction. The sampling frequency can be set according to application requirements, for example, approximately 50 Hz in one embodiment, but is not limited thereto.
[0031] (6) Machine learning recognition and prediction unit The machine learning recognition and prediction unit includes a feature extraction and fusion module and a temporal learning model inference module. The temporal learning model is preferably a multi-task temporal model, including a shared temporal encoder and classification and regression branches; the classification branch outputs object category recognition results and can further output surface morphology / texture category recognition results; the regression branch outputs weight (mass) prediction values. The shared temporal encoder can preferably be an LSTM, but can also be replaced by any one or a combination of GRU, TCN, and Transformer.
[0032] (7) Optional texture excitation control In an optional implementation, the system further includes a control unit for controlling the robotic arm to generate micro-lateral displacement or micro-vibration during the gripping and holding phase to stimulate or enhance texture-related dynamic fluctuation responses; and / or to generate relative lateral sliding at a preset speed to acquire texture sample data during the data acquisition phase. This control method is optional and does not constitute a limitation on the basic gripping process.
[0033] A multimodal tactile sensing and object recognition system for a robotic arm includes: A flexible tactile sensor is attached to the gripping surface and / or the dorsal region of the joint of the robotic hand. The flexible tactile sensor is a resistive flexible sensor, including an encapsulation layer, an electrode layer, a conductive hydrogel sensitive layer, and an elastomer layer. The encapsulation layer and / or the elastomer layer are provided with a microcone array on the side facing the conductive hydrogel sensitive layer and form an interlocking / interlocking interface with the conductive hydrogel sensitive layer. This allows the resistance timing signal output by the flexible tactile sensor during the gripping process to simultaneously reflect the pressure response under normal pressure, the strain response under bending or tensile deformation, and the texture dynamic fluctuation response related to contact micro-friction / micro-slippage. The signal processing unit, electrically connected to the electrode, is configured to perform synchronous alignment, baseline correction and normalization processing on the resistance timing signal, extract pressure features, strain features and texture features, and fuse the features to form a multimodal input. The machine learning recognition and prediction unit is communicatively connected to the signal processing unit and integrates a trained temporal learning model. The temporal learning model is configured to receive the multimodal input and output recognition and prediction results, which are used to identify the object category, quality and surface morphology of the grasped object.
[0034] The microcone array is disposed on the contact surface of the encapsulation layer and / or the elastomer layer, and the microstructure is a periodic array structure, which further includes at least one of the following: micro pyramid array, micro needle array, micro column array, hemispherical array or biomimetic texture array.
[0035] The height of the microstructure array is 100–800 mm. The array spacing is 150–1200. .
[0036] The conductive hydrogel sensitive layer is an ion-conductive hydrogel or a composite conductive hydrogel. The resistive flexible tactile sensor generates a pressure response under normal pressure, a strain response under bending or tensile deformation, and a texture-related dynamic fluctuation response when micro-friction / micro-slippage occurs with the surface of the object.
[0037] The texture features are extracted from the dynamic fluctuation components related to micro-friction / micro-slip in the resistance time-series signal, and the texture features include at least one of high-frequency fluctuation amplitude, fluctuation statistics, and frequency domain energy distribution features.
[0038] The texture features also include time-frequency domain features obtained through short-time Fourier transform and / or wavelet transform.
[0039] The signal processing unit is configured to synchronously sample the multi-channel resistance timing signal and perform alignment, baseline correction and normalization processing; the normalization is performed with reference to the static baseline resistance before the start of the capture.
[0040] The temporal learning model in the machine learning recognition and prediction unit is a multi-task temporal model, including a shared temporal encoder and a classification branch and a regression branch. The classification branch outputs the object category, and the regression branch outputs the quality prediction value. The shared temporal encoder is preferably an LSTM, and can be replaced by any one or a combination of GRU, TCN, or Transformer.
[0041] The system also includes a control unit configured to control the robotic arm to generate micro-lateral displacement or micro-vibration during the grasping and holding phase to stimulate or enhance the texture-related dynamic fluctuation response; and / or to generate relative lateral sliding at a preset speed during the data acquisition phase to acquire texture sample data.
[0042] The flexible tactile sensor is a resistive flexible sensor, attached to the gripping surface and / or the dorsal region of the joint of the robotic arm. The resistive flexible sensor includes an encapsulation layer, an electrode layer, a conductive sensitive layer, and an elastomer layer. The encapsulation layer and / or the elastomer layer has a microstructured surface on the side facing the conductive sensitive layer and forms an interlocking / fitting interface with the conductive sensitive layer, so that the resistance timing signal output during the gripping process simultaneously reflects: the pressure response under normal pressure, the strain response under bending or tensile deformation, and the dynamic texture fluctuation response related to contact micro-friction / micro-slippage; wherein the conductive sensitive layer is preferably a conductive hydrogel sensitive layer.
[0043] The signal acquisition and processing unit is electrically connected to the electrodes and is used to synchronously acquire the original resistance timing signal and perform alignment, baseline correction and normalization processing, and extract pressure features, strain features and texture features from the signal.
[0044] The machine learning recognition and prediction unit is communicatively connected to the signal acquisition and processing unit. The machine learning recognition and prediction unit includes a feature fusion module and a machine learning recognition and prediction module. The feature fusion module is used to fuse pressure features, strain features, and texture features to form a multimodal input. The machine learning recognition and prediction module integrates a trained temporal learning model to receive the multimodal input and output the recognition and prediction results of the grasped object. The recognition and prediction results include at least the object category recognition result and the weight prediction value, and may further include the surface morphology recognition result; wherein the surface morphology recognition result is used to characterize the category of surface roughness, smoothness, or texture type of the grasped object.
[0045] The methods include: Multi-channel resistance timing signals are acquired during the process of a robotic arm grasping an object; the resistance timing signals are synchronized, baseline corrected, and normalized; pressure features, strain features, and texture features are extracted and fused to form a multimodal input; the multimodal input is input into a pre-trained temporal learning model, and the output is used to identify the object category, mass, and surface morphology of the grasped object and predict its characteristics.
[0046] A method for object recognition, surface morphology, and weight prediction based on multimodal tactile temporal signals, applied to the aforementioned system, includes at least the following steps: S1 Grasping Timing Signal Acquisition: Control the robotic arm to execute a grasping action sequence including initial homing, closed grasping, stable holding, and release reset, and acquire multi-channel raw resistance timing signals during the grasping process. The sampling frequency and acquisition duration can be set according to the task; for example, in one embodiment, the sampling frequency is about 50 Hz and the single grasping timing duration is about 8 s, but it is not limited to this.
[0047] S2 Signal Preprocessing: Synchronizes and aligns the raw signals from multiple channels; acquires the baseline of the stationary phase before capture begins and performs baseline correction; normalizes the signal; and can selectively perform filtering and noise reduction, outlier removal, and drift correction.
[0048] S3 Feature Extraction and Construction: Extract pressure features, strain features, and texture-related features from the preprocessed signal. Texture-related features are obtained by extracting dynamic wave components related to micro-friction / micro-slip, and may include at least one of high-frequency wave amplitude, wave statistics, and frequency domain energy distribution features, and may further include time-frequency domain features obtained through short-time Fourier transform (STFT) and / or wavelet transform.
[0049] S4 Multimodal Fusion: Align and fuse pressure features, strain features, and texture-related features according to a unified time window to form multimodal input samples; or, use the aligned multi-channel time series segments as end-to-end inputs, and let the time series learning model automatically learn the multimodal time series representation.
[0050] S5 Multi-Task Model Recognition and Weight Prediction: Input multimodal input samples into a pre-trained multi-task temporal model, output object category recognition results and weight prediction values, and can further output surface morphology / texture category recognition results.
[0051] S6 Model Training: The constructed samples are divided into training and test sets proportionally, for example, the training set accounts for 70%–90% and the test set accounts for 10%–30%; a joint loss function is used for training, where the classification branch can use cross-entropy loss, and the regression branch can use at least one of mean squared error, mean absolute error, or Huber loss; a staged training strategy can be adopted, that is, pre-training for the category recognition task is performed first, followed by weighted regression training and joint fine-tuning, in order to improve convergence stability and generalization performance.
[0052] The above data acquisition, signal processing, and recognition prediction process can be found in [reference needed]. Figure 9 The overall process diagram shown is shown below.
[0053] Compared with the prior art, the present invention has at least the following beneficial effects: (1) Texture information can be obtained without additional texture sensors: the surface texture difference is mapped to the texture-related dynamic fluctuation component in the resistance timing signal through the microstructure interlock / hook interface, and texture perception and differentiation can be achieved in the normal grasping process, reducing hardware and control complexity.
[0054] (2) Multimodal fusion improves robustness: The fusion of pressure, strain and texture-related features can reduce the influence of different friction conditions and contact stability changes on recognition and regression, improve the stability of category recognition under complex surface conditions and reduce weight prediction error.
[0055] (3) Single-time grasping synchronous output of multi-task results: The multi-task temporal model based on the shared temporal encoder can synchronously output the object category and weight prediction results under the same grasping temporal input, and can extend the output of surface morphology recognition results to improve system integration and inference efficiency.
[0056] (4) Implementation verification: In one embodiment, the system can use a multi-channel array to sample at about 50 Hz and collect about 8 s of grasping time sequence to construct samples; multiple types of object data can be used for training and verification, and good results are obtained in classification accuracy and weight regression index. The above parameters and results are only used to illustrate the feasibility and effect of the present invention and do not constitute a limitation on the scope of protection of the present invention.
[0057] In one embodiment, the texture-related features include high-frequency fluctuation features, frequency domain energy features, and / or time-frequency domain features of the pressure signal, wherein the time-frequency domain features are obtained by short-time Fourier transform or wavelet transform.
[0058] In one embodiment, the multi-task temporal model includes a shared temporal encoder and a classification branch and a regression branch; the shared temporal encoder is any one or a combination of LSTM, GRU, TCN, and Transformer; the classification branch outputs the object category; and the regression branch outputs the weight prediction value.
[0059] In one embodiment, the normalization process employs... or ,in To capture the average baseline resistance when the sensor is stationary before the capture begins.
[0060] In one embodiment, the multimodal input can be either a fused feature extracted from the signal or an aligned multi-channel temporal segment, with the model learning textures and capturing dynamic features end-to-end.
[0061] Model training can use a joint loss function. ,in For classifying losses, The regression loss can be used; and a phased training strategy can be adopted, which involves classification pre-training, regression training, and joint fine-tuning.
[0062] In one embodiment, the model output results are statistically evaluated, and the classification task uses at least one of the following metrics: accuracy, precision, recall, and F1 score; in the embodiment where the surface morphology recognition results are output, the surface morphology recognition is also evaluated using the above-mentioned classification metrics.
[0063] The overall methodology when combining systems and steps is as follows: Step 1: Fabrication of a microstructure interlocking resistive flexible tactile sensor (1) Using a mold with a microstructure array pattern, an elastomer layer with microstructures (4) is prepared by casting / curing to form a microstructure array (13). The morphology and size parameters of the microstructure array can be selected and adjusted according to the application requirements. For example, a microcone array, or a micro pyramid, microneedle, microcolumn, hemisphere or biomimetic texture array can be used.
[0064] (2) An electrode layer (2) is set in a preset area of the elastomer layer to ensure a stable electrical connection with the conductive sensitive layer (3).
[0065] (3) After the conductive sensitive layer (3) is shaped, it is attached to the microstructure surface of the elastomer layer (4) so that the microstructure array and the conductive sensitive layer form an interlocking / interlocking interface.
[0066] (4) The encapsulation layer (1) forms a sealed protective structure to reduce moisture evaporation and improve wear resistance and anti-pollution ability; in an optional embodiment, the encapsulation layer (1) is also provided with a microstructure array on the side facing the conductive sensitive layer (3) to form an interlocking interface on the upper and lower sides of the conductive sensitive layer.
[0067] (5) The above preparation process is used to illustrate the feasibility of the present invention. The material formulation and process parameters can be adjusted according to the application requirements and do not constitute a limitation.
[0068] Step 2: Array Integration and Acquisition Channel Configuration (1) The pressure channel group (7) is arranged in the fingertip, fingertip and / or palm contact area to obtain pressure response and carry texture-related dynamic fluctuation information.
[0069] (2) Arrange the strain channel group (8) on the back side of the joint or in the area of significant deformation to obtain the strain response caused by bending / tension.
[0070] (3) Each channel is converged and led out to the signal acquisition module (10) via FPC (9), and the processing terminal (11) saves the mapping relationship between the channel number and the installation location.
[0071] (4) The number of channels can be configured according to the structure of the robot and the coverage area. For example, in one embodiment, a total of 23 acquisition channels are set, but it is not limited to this.
[0072] Step 3: Standard Grabbing Actions and Data Acquisition (1) The sequence of actions performed by the robotic arm may include: initial homing → closing gripping → stable holding → release and reset.
[0073] (2) In the initial positioning stage, a static baseline is acquired; in the closing gripping stage, contact is established and a pressure-related resistance response is generated, and the strain channel responds synchronously; in the stable holding stage, due to the natural micro-friction / micro-slippage of the contact interface, texture-related dynamic fluctuations appear in the resistance timing signal; in the release and reset stage, the signal returns to the vicinity of the baseline.
[0074] (3) The sampling frequency and the duration of a single sample collection can be set according to the task. For example, in one embodiment, the sampling frequency is about 50 Hz and the duration of a single sample collection is about 8 s; and the samples can be constructed in segments according to a preset time window, but it is not limited to this.
[0075] Step 4: Preprocessing, Feature Construction, and Multi-Task Model Training / Inference (1) Preprocessing: Synchronization alignment, baseline correction and normalization are performed on the multi-channel signals, and selective processing such as filtering and noise reduction, outlier removal and drift correction can be performed. For example, in one embodiment, the above preprocessing is performed on a 23-channel signal, but it is not limited to this.
[0076] (2) Texture features: Extract dynamic fluctuation features related to micro-friction / micro-slip from the normalized resistance time-series signal of the pressure-related channel. These features may include high-frequency fluctuation features, frequency domain energy features, and / or time-frequency domain features. The time-frequency domain features may be obtained by short-time Fourier transform (STFT) or wavelet transform.
[0077] (3) Multimodal fusion: The pressure features, strain features and texture-related features are fused to form the input sample; or, the aligned multi-channel time series segments are used as end-to-end inputs and the model automatically learns the multimodal time series representation.
[0078] (4) Model: A multi-task temporal model with a shared temporal encoder, a classification branch, and a regression branch is adopted; the shared encoder can be any one of LSTM, GRU, TCN, Transformer, or a combination thereof. The classification branch outputs the object category recognition result and can further output the surface morphology / texture category recognition result; the regression branch outputs the weight prediction value.
[0079] (5) Training data: The data scale can be constructed according to application requirements. For example, in one embodiment, it includes, for example, 30 categories of objects, with, for example, 2000 samples per category; each sample can be, for example, time-series data obtained by sampling at, for example, 23 channels, about 8 s, and about 50 Hz; the weight range covers, for example, from 1 g to 300 g, but is not limited thereto.
[0080] (6) Training method: The joint loss function is used for optimization; the classification branch can use cross-entropy loss; the regression branch can use at least one of MSE, MAE or Huber loss; and a phased training strategy can be adopted to improve training stability and generalization ability.
[0081] (7) Performance Verification: Based on the data from the above embodiments, a high object classification accuracy and a low weight prediction error can be obtained; for example, in one embodiment, the classification accuracy can reach 98.89%, and the weight prediction can reach R 2 =0.9960, RMSE=0.0281, MAE=0.0131. The above results are only used to illustrate the feasibility and effects of the present invention and do not constitute a limitation on the scope of protection.
[0082] The application scenarios of this invention include: Tennis ball grabbing and recognition, surface morphology judgment and weight prediction.
[0083] like Figure 2 As shown, a microstructure interlocked resistive flexible tactile sensor is used as the tactile acquisition unit. The sensor includes an encapsulation layer 1, an electrode layer 2, a conductive sensitive layer 3, and an elastomer layer 4. The elastomer layer 4 is located on the upper and / or lower side of the sensor, and a microstructure array 13 is provided on the side facing the conductive sensitive layer 3. The conductive sensitive layer 3 and the microstructure array 13 are bonded together to form an interlocking / fitting interface, so that the surface texture difference caused by contact micro-friction / micro-slippage is mapped into the texture-related dynamic fluctuation response in the resistance timing signal. The conductive sensitive layer 3 is preferably a conductive hydrogel sensitive layer.
[0084] The sensor size and thickness can be set according to the robot arm positioning and mounting space; for example, the effective size can be approximately 10 mm × 5 mm. The thickness can be adjusted according to the packaging method and installation space. Preferably, the thickness of the packaging layer 1 and the elastomer layer 4 can each be approximately 1 mm.
[0085] The sensor fabrication and packaging process may include: (1) Using a mold with a microstructure array pattern, an elastomer layer 4 with microstructures is prepared by casting / curing. (2) Fix the electrode layer 2 in the preset area to form an electrical connection interface. The electrode layer 2 can be a copper foil electrode. (3) Shape the conductive sensitive layer 3 and attach it to the electrode area, and attach it to the microstructure surface to form an interlocking interface. (4) A sealed protective structure is formed by covering the encapsulation layer 1; preferably, a microstructure array 13 is also provided on the side of the encapsulation layer 1 facing the conductive sensitive layer 3 to form an interlocking interface on the upper and lower sides of the conductive sensitive layer 3.
[0086] like Figure 1 As shown, the system in this embodiment includes a robotic actuator 6, a multi-channel flexible tactile sensor array, a signal acquisition module 10, and a processing terminal 11. The sensors are attached to the gripping surface and / or the dorsal region of the joints to acquire pressure-strain-texture multimodal temporal signals during the gripping process; such as... Figure 4 and Figure 5 As shown, each channel can be converged and led out via FPC 9 and connected to the signal acquisition module 10; as Figure 6 As shown, the processing terminal 11 establishes and stores the correspondence between the acquisition channel and the sensor unit number / installation location, such as... Figure 9 As shown, the data acquisition, signal processing, and recognition prediction of the system of the present invention form an integrated process.
[0087] like Figure 3 As shown, the sensor can be placed in the contact area of the fingertip, fingertip and / or palm, as well as the area of significant deformation on the dorsal side of the joint, to simultaneously obtain pressure-related response, strain-related response and texture-related dynamic fluctuation response during a single grasping process; among them, the resistance timing signal of the contact area is more likely to reflect pressure response and texture fluctuation information, and the signal of the dorsal side of the joint is more likely to reflect strain response.
[0088] Using a tennis ball as the object to be grasped. The robotic arm presses... Figure 7The sequence of grasping actions shown includes initial homing, closed gripping, stable holding, and release reset: the initial homing phase acquires a static baseline; the closed gripping phase establishes contact and generates a resistance response related to normal loading; when the robot arm bends or stretches, the corresponding channels generate strain-related responses; during the stable holding phase, due to the natural micro-friction / micro-slippage at the contact interface, a texture dynamic fluctuation response related to the differences in the object's surface morphology appears in the resistance timing signal; during the release reset phase, the signal returns to the vicinity of the baseline.
[0089] The sampling frequency is preferably 50 Hz; the preferred duration of a single sample acquisition is approximately 8 seconds; the processing terminal can construct samples in segments or windows according to a preset time window. The processing terminal can perform at least one of the following on the multi-channel resistance timing signal: filtering and noise reduction, outlier removal, and drift correction, and align the multi-channel signals. The normalized form can be... or , To capture the average baseline resistance when the sensor is stationary before the capture begins.
[0090] Multimodal features are extracted from the aligned and normalized multichannel resistance time-series signal. These features include at least pressure features, strain features, and texture features. Texture features are obtained by extracting dynamic fluctuation components related to micro-friction / micro-slip, and may include high-frequency fluctuation amplitudes and statistics, frequency domain energy distribution characteristics, and / or time-frequency energy characteristics. Due to the rough surface structure of tennis balls, the equivalent roughness is high, and interfacial friction is more pronounced. During the stable holding phase, alternation between micro-friction and micro-slip is more likely to occur, resulting in texture-related dynamic fluctuation responses with larger amplitudes and richer detail ripples. In the frequency / time-frequency domain, this can manifest as a relatively higher proportion of high-frequency energy or a more dispersed distribution. This difference can serve as an important basis for surface morphology recognition and object category recognition. The pressure features, strain features, and texture features are then fused to form a multimodal input sample.
[0091] The processing terminal inputs multimodal input samples into the multi-task temporal model 12 for training or inference. The multi-task temporal model includes a shared temporal encoder and classification and regression branches. The classification branch is used to output object category recognition results, and the regression branch is used to output weight prediction values. Optionally, the model can also output surface morphology recognition results to characterize roughness, smoothness, or different texture types.
[0092] Baseball grabbing and recognition, surface shape judgment and weight prediction.
[0093] like Figure 2As shown, this embodiment also uses a microstructure interlocking resistive flexible tactile sensor. The sensor includes an encapsulation layer 1, an electrode layer 2, a conductive sensitive layer 3, and an elastomer layer 4. The encapsulation layer and / or the elastomer layer are provided with a microstructure array 13 on the side facing the conductive sensitive layer and are bonded to the conductive sensitive layer 3 to form an interlocking / embedding interface. The conductive sensitive layer 3 is preferably a conductive hydrogel sensitive layer.
[0094] The sensor size and thickness can be set according to the mounting space, for example, the effective size is about 10 mm × 5 mm; the thickness of the encapsulation layer 1 and the elastomer layer 4 can be about 1 mm each, or adjusted according to the flexibility requirements.
[0095] The sensor fabrication and packaging process may include: (1) Mold casting / curing to form an elastomer layer with microstructure 4; (2) The fixed electrode layer 2 forms an electrical connection interface; (3) After the conductive sensitive layer 3 is shaped, it is attached to the electrode area and interlocked with the microstructure surface; (4) A sealed protective structure is formed by covering the encapsulation layer 1, and a microstructure array 13 can also be set on the side of the encapsulation layer 1 facing the conductive sensitive layer 3 to enhance the interlocking effect.
[0096] like Figure 1 As shown, sensors are attached in an array to the gripping surface and / or the dorsal region of the joints of the robotic arm, forming a multi-channel flexible tactile sensor array; each channel is connected to the signal acquisition module 10 via FPC 9 and transmitted to the processing terminal 11; the processing terminal 11 stores the correspondence between the channel number and the installation position to achieve data management (e.g., Figure 6 ).
[0097] like Figure 3 As shown, the channels in the contact area are used to acquire pressure-related responses and contain texture dynamic fluctuation information, while the back side region of the joint is used to acquire strain-related responses.
[0098] Using a baseball as the object to be grasped, the robotic arm performs the operation. Figure 7 The grasping action sequence shown is used to acquire multi-channel resistance timing signals: initial positioning acquires the baseline; closing the grasp establishes contact and generates a pressure-related resistance response; during the stable holding phase, a texture dynamic fluctuation response related to micro-friction / micro-slippage appears; during the release and reset phase, the signal returns to the vicinity of the baseline.
[0099] The sampling frequency and single acquisition duration can be the same as in Example 1; alignment, baseline correction, and normalization of multi-channel signals can be performed as follows: or , (Based on the average value of the baseline resistance), and can selectively perform filtering and noise reduction, outlier removal and drift correction.
[0100] Pressure features, strain features, and texture features are extracted and fused to form a multimodal input sample. Texture features are obtained by extracting dynamic fluctuation components related to the stable holding phase and natural microslip, and may include high-frequency fluctuation statistics, frequency domain energy distribution features, and optional time-frequency domain features. Compared to the textured tennis ball in Example 1, the leather surface of a baseball is relatively smooth, and the interface friction and microslip morphology are closer to stable slip or low-amplitude microslip. Therefore, the texture-related dynamic fluctuation response usually shows distinguishable differences in amplitude, waveform regularity, and spectral energy distribution. This difference can be used to characterize the smooth surface morphology category and, together with pressure / strain features, improve the stability of ball object category recognition and weight prediction.
[0101] The multimodal input samples are input into the multi-task temporal model 12 for training or inference, and the output results are the object category recognition result and weight prediction value corresponding to the baseball. The surface morphology recognition result can be output optionally.
[0102] Grasping and recognition of plastic bottles / cylindrical containers, surface morphology judgment and weight prediction.
[0103] like Figure 2 As shown, this embodiment employs a microstructure interlocking resistive flexible tactile sensor. The sensor includes an encapsulation layer 1, an electrode layer 2, a conductive sensitive layer 3, and an elastomer layer 4. A microstructure array 13 is disposed on the side of the encapsulation layer and / or the elastomer layer facing the conductive sensitive layer, forming an interlocking / interlocking interface with the conductive sensitive layer. The conductive sensitive layer 3 is preferably a conductive hydrogel sensitive layer.
[0104] The sensor size and thickness can be set according to the robot arm mounting space, for example, the effective size is about 10 mm × 5 mm; the thickness of the encapsulation layer 1 and the elastomer layer 4 can be about 1 mm or adjusted as needed.
[0105] The sensor fabrication and packaging process may include: (1) The microstructured elastomer layer 4 is formed by casting / curing through a mold; (2) The fixed electrode layer 2 forms an electrical connection interface; (3) The conductive sensitive layer 3 is attached to the electrode area and interlocked with the microstructure surface to form an interlocking interface; (4) The encapsulation layer 1 is covered to form a sealed protection structure, and a microstructure array 13 can be set on the side of the encapsulation layer 1 facing the conductive sensitive layer 3 to enhance the interlocking effect.
[0106] like Figure 1 , Figure 4 , Figure 5 As shown, the sensor array is attached to the gripping surface and / or the dorsal region of the joint of the robotic arm, and transmitted to the processing terminal 11 via the signal acquisition module 10 connected to the FPC 9; the processing terminal establishes and stores the correspondence between the acquisition channel and the installation position (e.g., Figure 6 ).
[0107] like Figure 3 As shown, the gripping contact area is used to collect pressure-related responses and includes texture dynamic fluctuation information, while the dorsal joint area is used to collect strain-related responses caused by bending / tension.
[0108] The robotic arm performs the grasping action, using plastic bottles or rigid cylindrical containers as the objects to be picked up. Figure 7 The grasping action sequence is shown, and multi-channel resistance timing signals are acquired. Because these objects typically exhibit regular cylindrical curvature, high material stiffness, and small indentation deformation, the pressure-related resistance response during the contact establishment phase differs from that of spherical objects in terms of rise slope, steady-state level, and multi-channel distribution statistical characteristics. The strain-related response of the dorsal channel of the joint shows different amplitudes and rates of change depending on the degree of closure and gripping posture. During the stable holding phase, cylindrical objects are more likely to exhibit micro-rolling or micro-rotation tendencies under clamping action, resulting in a different interface micro-slip pattern compared to spherical objects. Consequently, the texture-related dynamic fluctuation response exhibits another type of characteristic in waveform details and spectral energy distribution (e.g., superimposed low-frequency slow fluctuations and local high-frequency jitter, or the appearance of characteristic frequency bands related to the micro-texture of the cylindrical surface).
[0109] The sampling frequency and acquisition duration can be preferably set to 50 Hz and approximately 8 s, respectively; alignment, baseline correction, and normalization of multi-channel signals can be performed as follows: or , It is the baseline resistance average value, and can selectively perform filtering and noise reduction, outlier removal and drift correction.
[0110] Pressure, strain, and texture features are extracted from the preprocessed signal and fused to form a multimodal input sample. Texture features are obtained by extracting dynamic fluctuation components related to the stable holding phase and natural micro-slip, and may include high-frequency fluctuation features, frequency domain energy features, and / or time-frequency energy features. Compared to spherical objects, in this embodiment, pressure and strain features more prominently characterize the "mechanical differences in contact between regular cylinders," while texture features provide supplementary criteria for distinguishing "surface morphology differences," thereby improving overall recognition robustness and weight prediction stability.
[0111] The multimodal input samples are input into the multi-task temporal model 12 for training or inference, and the output results are the object category recognition results and weight prediction values corresponding to the plastic bottle / cylindrical container. The surface morphology recognition results can also be output.
[0112] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A multimodal tactile sensing and object recognition system for a robotic arm, characterized in that, The system includes a robotic arm actuator (6), a multi-channel flexible tactile sensor array, a signal acquisition module (10), and a processing terminal (11). The flexible tactile sensor in the multi-channel flexible tactile sensor array includes an encapsulation layer (1), an electrode layer (2), a conductive sensitive layer (3), and an elastomer layer (4). The multi-channel flexible tactile sensor array contacts the target object (5). The encapsulation layer (1) and / or the elastomer layer (4) are provided with a microstructure array (13) on the side facing the conductive sensitive layer (3), and the microstructure array (13) and the conductive sensitive layer (3) form an interlocking / interlocking interface.
2. The multimodal tactile sensing and object recognition system for a robotic arm according to claim 1, characterized in that, The multi-channel flexible tactile sensor array also includes a strain channel group (8) and a pressure channel group (7), and the flexible tactile sensor is connected to the strain channel group (8) and the pressure channel group (7).
3. The multimodal tactile sensing and object recognition system for a robotic arm according to claim 2, characterized in that, The strain channel group (8) and the pressure channel group (7) are connected to the flexible printed circuit board (FPC) (9). The flexible printed circuit board (FPC) (9) is connected to the signal acquisition module (10). The signal acquisition module (10) is connected to the processing terminal (11). The processing terminal (11) is set with a multi-task timing model (12).
4. The multimodal tactile sensing and object recognition system for a robotic arm according to claim 1, characterized in that, The microstructure array (13) is a periodic array structure.
5. A multimodal tactile sensing and object recognition system for a robotic arm according to claim 4, characterized in that, The microstructure array (13) is any one of the following: micro pyramid array, micro needle array, micro column array, hemispherical array or biomimetic texture array.
6. The multimodal tactile sensing and object recognition system for a robotic arm according to claim 4, characterized in that, The conductive sensitive layer (3) is a conductive hydrogel sensitive layer.
7. A method for multimodal tactile perception and object recognition of a robotic arm, characterized in that, The method using the system according to any one of claims 1 to 6 includes the following steps: S1 controls the robotic arm actuator (6) to execute a grasping action sequence including initial homing, closed grasping, stable holding and release reset, and collects multi-channel raw resistance timing signals during the grasping process based on the signal acquisition module (10); S2 preprocesses the multi-channel original resistance timing signal in the processing terminal (11) to obtain the preprocessed resistance timing signal; S3 extracts pressure features, strain features, and texture-related features from the preprocessed resistance time-series signal and fuses them to form a fused feature set; The S4 fusion feature set is input into the pre-trained multi-task temporal model (12), which outputs the object category recognition result and weight prediction value.
8. The method for multimodal tactile perception and object recognition of a robotic arm according to claim 7, characterized in that, Texture-related features include one or more of the following: high-frequency fluctuation features of pressure signals, frequency domain energy features, and time-frequency domain features.
9. A method for multimodal tactile perception and object recognition of a robotic arm according to claim 7, characterized in that, The specific steps for fusing features to form a fused feature set are as follows: The pressure features, strain features, and texture-related features are aligned and fused according to a unified time window to obtain a fused feature set.
10. A method for multimodal tactile perception and object recognition of a robotic arm according to claim 7, characterized in that, Preprocessing includes synchronization alignment, obtaining the baseline of the static phase before the start of the grabbing and performing baseline correction, normalization and filtering for noise reduction, outlier removal and drift correction.