Magnetic haptic perception method and system based on double-code dendritic modulation
Patent Information
- Application Number
- CN202610961451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]针对现有技术中磁触觉感知方法难以同时兼顾阵列空间分布、时序动态变化和任务上下文解释的问题,本发明提供一种基于双码树突调制的磁触觉感知方法、磁触觉网络训练方法、磁触觉感知系统、机器人、电子设备及计算机可读存储介质
[0045]1)本发明以机器人末端磁触觉阵列传感器采集的三轴磁触觉时序阵列为基础,构造能够反映原始磁场分布、磁场强度、动态变化、高频扰动和慢变趋势的磁触觉扩展特征;再通过空间码分支和时间码分支分别提取磁触觉空间表征和磁触觉时间表征,并经双码门控融合得到触觉主特征;同时,获取视觉上下文、语言上下文、音频上下文和本体感觉上下文中的至少两类候选上下文数据,并根据模态可用性标记对上下文候选特征进行路由评分和加权聚合,得到上下文聚合特征;最后,由上下文聚合特征生成增益向量和偏置向量,以特征维度级调制方式对触觉主特征进行增强、抑制或修正,从而获得上下文增强触觉特征,并输出接触位置、法向形变、切向滑移量和接触模式等机器人接触状态感知结果。
Smart Images

Figure CN122807878A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot tactile perception and multimodal intelligent perception technology, and in particular to a magnetic tactile perception method and system based on dual-code dendritic modulation. Background Technology
[0002] When performing tasks such as grasping, pressing, sliding, insertion and assembly, edge tracking, and human-robot interaction, robots typically rely on end-effector sensors to perceive the contact state of objects. Compared to perception methods that rely solely on vision, tactile perception can directly reflect the actual contact between the robot's end effector and the target object, such as the contact position, contact area, normal indentation degree, tangential slippage trend, and changes in contact pattern. Therefore, tactile perception is a crucial foundation for robots to achieve compliant operation, stable grasping, anti-slip control, and safe interaction.
[0003] Magnetotactic array sensors are a commonly used type of robotic tactile sensor. They typically obtain triaxial magnetic field information related to the contact state by detecting changes in the magnetic field caused by compression, shearing, or sliding of a magnetic body or magnetoelastic structure. For array-type magnetotactic sensors, the output data includes not only the triaxial magnetic field components of each array node, but also the spatial distribution relationship between multiple array nodes and the dynamic changes between consecutive time frames. In other words, the magnetotactic signal itself has both spatial and temporal characteristics: spatial characteristics reflect the contact center, the diffusion of the compressed area, and the array topology; temporal characteristics reflect the contact establishment process, sliding direction, velocity changes, vibration disturbances, and contact state transitions.
[0004] In existing magnetotactile sensing methods, one type primarily processes single-frame tactile array data, such as inputting the magnetic field distribution at a specific moment into a convolutional network or a fully connected network to identify contact location or contact pattern. While this type of method can utilize the spatial distribution information of the magnetotactile array, it is insufficient in representing temporal dynamic processes such as sliding, vibration, contact establishment, and contact release. Another type of method inputs continuous tactile data into recurrent neural networks, temporal convolutional networks, or attention networks to enhance the ability to recognize dynamic changes in tactile sensation. However, these methods often fail to adequately distinguish between spatial and temporal code information in the magnetotactile signal, easily mixing different types of tactile information, such as contact area distribution, array node topology, sliding motion, and high-frequency perturbations, in the same representation, resulting in unclear physical meaning of the tactile features.
[0005] Furthermore, in real-world robotic tasks, tactile signals are not isolated. When the robot's end effector contacts an object, visual information can provide the target's location, edges, object type, and expected contact area; verbal information can provide task instructions, action intentions, and contact strength constraints; audio information can provide cues such as collision sounds, friction sounds, and knocking sounds; and proprioceptive information can provide the robot's end effector speed, joint posture, motion stage, and force control state. The same normal deformation may represent different meanings in touch confirmation, stable grasping, and collision detection tasks; similarly, the same tangential slip may correspond to different control strategies in edge tracking, grasping anti-slip, and abnormal contact detection. Therefore, the robot's understanding of magnetotactic signals needs to be interpreted within the context of the current task.
[0006] Existing multimodal tactile perception methods typically employ direct splicing or simple fusion, concatenating tactile features with visual, linguistic, audio, or proprioceptive features before inputting them into a subsequent recognition network. While this approach incorporates multi-source information, it still has several shortcomings. First, direct splicing tends to place tactile perception on an entirely parallel footing with other modalities, potentially causing the model to over-rely on high-dimensional visual or linguistic features, thus weakening tactile perception's role as the primary pathway for physical contact. Second, direct splicing lacks clear modulation relationships, making it difficult to explain which tactile feature dimensions are enhanced, suppressed, or modified by vision, language, audio, or proprioception. Third, in real-world robotic applications, the contextual modality is not always complete and reliable. For example, vision may be occluded, linguistic commands may be missing, audio may be interfered with by environmental noise, and proprioception may experience sampling delays or state asynchrony. Simple splicing methods are prone to unstable perception results when context is missing or noise levels are high.
[0007] Furthermore, the results of robot tactile perception typically need to serve subsequent motion control. For example, during grasping, the grasping force needs to be adjusted based on normal deformation and slippage trend; during sliding contact, the end effector speed needs to be adjusted based on tangential slippage; and in the event of a collision or abnormal contact, timely obstacle avoidance or shutdown control is required. Therefore, magnetic tactile perception methods not only need to output contact pattern classification results, but also continuous quantities such as contact position, normal deformation, and tangential slippage that can be used for robot control. Ideally, they should also provide some intermediate explanatory information so that the system can determine the degree of influence of different contexts on the tactile perception results.
[0008] Therefore, how to fully utilize the spatial and temporal code information in the magnetic tactile timing array while preserving the physical meaning of the main magnetic tactile pathway, and further introduce contextual information such as vision, language, audio, and proprioception to conditionally modulate the main tactile features, while simultaneously addressing the robust perception problem when contextual modalities are missing, is a pressing technical challenge in the field of robotic magnetic tactile perception. Summary of the Invention
[0009] To address the problem that existing magnetic tactile sensing methods struggle to simultaneously consider array spatial distribution, temporal dynamic changes, and task context interpretation, this invention provides a magnetic tactile sensing method based on dual-code dendritic modulation, a magnetic tactile network training method, a magnetic tactile sensing system, a robot, an electronic device, and a computer-readable storage medium.
[0010] To achieve the above objectives, the present invention provides a magnetic tactile sensing method based on dual-code dendritic modulation, comprising the following steps:
[0011] S1. Acquire the perception data generated by the robot end effector when performing a contact task. The perception data includes a three-axis magnetic tactile temporal array collected by a magnetic tactile array sensor installed on the robot end effector, as well as contextual source data and modal availability markers related to the current contact task. The three-axis magnetic tactile temporal array includes multiple time frames, multiple array nodes, and a three-axis magnetic field component corresponding to each array node. Before or after constructing the three-axis magnetic tactile temporal array, the three-axis magnetic field components are subjected to contactless baseline subtraction, zero-point drift compensation, and normalization to obtain a baseline-corrected three-axis magnetic tactile temporal array. The contextual source data includes at least two types of candidate contextual data from visual context, linguistic context, audio context, and proprioceptive context. The modal availability markers are used to characterize whether each type of candidate contextual data is available in the current inference sample.
[0012] S2. Construct magnetic tactile extended features based on the triaxial magnetic tactile timing array, wherein the magnetic tactile extended features include original triaxial magnetic field features, magnetic field modulus features, time difference features, high-frequency residual features, and exponential smoothing features.
[0013] S3. Input the magnetic tactile extended features into the tactile dual-code main path, extract spatial code tactile features through the spatial code branch, extract temporal code tactile features through the temporal code branch, and perform dual-code gating fusion on the spatial code tactile features and temporal code tactile features to obtain the tactile main features.
[0014] S4. Input the context source data into the corresponding context encoder, generate corresponding context candidate features for candidate context data in an available state, generate preset placeholder features or default features for candidate context data in an unavailable state, and map candidate context data of different categories to the same context feature dimension through the corresponding context encoder or linear mapping layer to obtain multiple context candidate features.
[0015] S5. Perform context routing scoring and weight normalization on the multiple context candidate features according to the modal availability label to obtain context weights, and perform weighted aggregation on the multiple context candidate features according to the context weights to obtain context aggregate features.
[0016] S6. Input the context aggregation feature into the dendritic modulator, which generates a gain vector and a bias vector corresponding to the tactile main feature dimension, and uses the gain vector and bias vector to perform feature dimension-level modulation on the tactile main feature to obtain context-enhanced tactile features.
[0017] S7. Input the context-enhanced tactile features into the task output head and output the robot contact state perception results. The robot contact state perception results include one or more of the following: contact position, normal deformation, tangential slip, and contact mode.
[0018] In one optional embodiment, the triaxial magnetic tactile temporal array is represented as a tensor of B×T×H×W×C, where B represents the batch size, T represents the time window length, H and W represent the number of rows and columns of the magnetic tactile array, respectively, and C represents the number of magnetic field channels, with C=3.
[0019] In one optional implementation, the magnetic field modulus feature is obtained by taking the square root of the sum of the squares of the three-axis magnetic field components of the same array node in the same time frame; the time difference feature is obtained by the difference between the three-axis magnetic field components of the same array node in adjacent time frames; the exponential smoothing feature is obtained by performing an exponential moving average of the magnetic field modulus feature along the time dimension; and the high-frequency residual feature is obtained by the difference between the magnetic field modulus feature and the exponential smoothing feature. Thus, the static distribution, dynamic changes, high-frequency disturbances, and slow-changing trends in the original magnetic tactile signal can be explicitly expressed.
[0020] In one optional implementation, the spatial code branch takes node features including original triaxial magnetic field characteristics, magnetic field magnitude characteristics, and exponential smoothing characteristics as input, and performs spatial topology modeling based on the local adjacency relationships between array nodes in the magnetic tactile array to obtain spatial code tactile features reflecting the contact center location, contact area diffusion, and array spatial distribution. The temporal code branch takes temporal features including temporal difference characteristics, high-frequency residual characteristics, and magnetic field magnitude variation characteristics as input, and performs dynamic modeling along the time dimension to obtain temporal code tactile features reflecting the sliding direction, velocity changes, vibration disturbances, and contact establishment process. The magnetic field magnitude variation characteristics are obtained based on the difference in magnetic field magnitude characteristics of the same array node in adjacent time frames.
[0021] In one optional implementation, the spatial code branch includes a node embedding module, a local graph attention module, a node attention pooling module, and a temporal attention pooling module connected in sequence; wherein, the local graph attention module only interacts with array nodes that satisfy a preset adjacency relationship. The temporal code branch includes a single-frame projection module, a temporal convolution module, a recurrent dynamic coding module, and a temporal attention pooling module connected in sequence. Through the coordinated setting of the spatial code branch and the temporal code branch, the tactile main feature simultaneously includes contact spatial distribution information and contact temporal change information. The preset adjacency relationship is as follows: for any array node, array nodes whose row difference and column difference are not greater than 1 are determined as adjacent nodes.
[0022] In one optional implementation, the dual-code gating fusion of the spatial code tactile features and the temporal code tactile features includes: concatenating the spatial code tactile features and the temporal code tactile features to obtain a dual-code concatenated feature; inputting the dual-code concatenated feature into a gating network to obtain spatial code weights and temporal code weights; and performing a weighted summation of the spatial code tactile features and the temporal code tactile features according to the spatial code weights and temporal code weights to obtain the tactile master feature. Thus, the contribution ratio of spatial information and temporal information can be adaptively adjusted according to the current contact state.
[0023] In one alternative implementation, the visual context includes target location, target edge, target category, or expected contact area information; the linguistic context includes task instructions, action intentions, or contact intensity constraint information; the audio context includes contact sound, friction sound, knocking sound, or collision sound information; and the proprioceptive context includes robot end-effector velocity, joint posture, motion phase, or force control state information.
[0024] In one optional implementation, performing context routing scoring and weight normalization on the plurality of context candidate features based on the modal availability marker includes: calculating the routing score for each context candidate feature; setting the routing score corresponding to unavailable candidate context data to a preset suppression value according to the modal availability marker; normalizing the suppressed routing scores to obtain the context weights corresponding to each candidate context data; performing a weighted summation of the context candidate features in the available state according to the context weights to obtain the context aggregation feature; when all candidate context data are unavailable, setting the context aggregation feature to a preset zero vector or a preset default vector, and making the dendritic modulator output an all-1 gain vector and an all-0 bias vector, so that the context-enhanced tactile feature is equal to the tactile master feature, thereby causing the magnetic tactile perception method to revert to the output of the unmodulated tactile master feature. This allows for adaptation to scenarios with incomplete context, such as visual occlusion, missing language, audio noise, or proprioceptive asynchrony.
[0025] In one optional embodiment, the dendritic modulator includes a gain generation branch and a bias generation branch. The gain generation branch generates the gain vector based on the context aggregation features, and the bias generation branch generates the bias vector based on the context aggregation features. Modulating the tactile main feature using the gain vector and bias vector at the feature dimension level includes: multiplying each gain element in the gain vector by the feature value of the corresponding feature dimension in the tactile main feature to obtain the gain-modulated tactile feature; and superimposing each bias element in the bias vector onto the feature value of the corresponding feature dimension in the gain-modulated tactile feature to obtain the context-enhanced tactile feature. Therefore, the context information does not directly replace the tactile information, nor is it simply concatenated with the tactile features, but rather is used to amplify, suppress, or correct different feature dimensions in the tactile main feature.
[0026] This invention also provides a training method for a magnetic tactile network based on dual-code dendritic modulation, used to train a magnetic tactile perception network that executes the aforementioned magnetic tactile perception method based on dual-code dendritic modulation. The training method includes the following steps:
[0027] T1. Construct a training sample set, wherein each training sample in the training sample set includes a triaxial magnetotactile temporal array, contextual source data, modal availability markers, continuous quantity labels, and contact pattern labels.
[0028] T2. Input the triaxial magnetic tactile timing array into the tactile dual-code main path to obtain the tactile main features.
[0029] T3. Input the aforementioned context source data and modal availability tags into the context encoding and routing path to obtain context aggregation features.
[0030] T4. Generate a gain vector and a bias vector based on the context aggregation features, and use the gain vector and bias vector to modulate the tactile main features to obtain context-enhanced tactile features.
[0031] T5. Based on the context, enhance the tactile features to obtain continuous quantity prediction results and contact pattern prediction results.
[0032] T6. Train the magnetic tactile perception network according to the total training loss, which includes at least two of the following: continuous regression loss, contact pattern classification loss, gain stabilization loss, bias magnitude loss, context missing consistency loss, and context routing regularization loss.
[0033] In one alternative implementation, the continuous quantity regression loss is determined based on the difference between the continuous quantity prediction result and the continuous quantity label; the contact pattern classification loss is determined based on the difference between the contact pattern prediction result and the contact pattern label; the gain stabilization loss is used to constrain the deviation of the gain vector from unity gain; the bias magnitude loss is used to constrain the magnitude of the bias vector; and the context routing regularization loss is used to constrain the distribution state of the context weights.
[0034] In one optional implementation, the context-missing consistency loss is obtained as follows: Full context forward inference is performed on the same training samples to obtain a full context output; one or more of the following are randomly discarded: visual context, linguistic context, audio context, or proprioceptive context; and missing context forward inference is performed based on the updated modal availability label to obtain a missing context output; the context-missing consistency loss is determined based on the difference in continuous quantity predictions and the difference in contact pattern probability distributions between the full context output and the missing context output. By introducing the context-missing consistency loss, the magnetotactic sensing network can maintain relatively stable output results even when the context modality is incomplete.
[0035] In one optional implementation, the training method includes: a first stage, freezing the tactile dual-code main pathway used to obtain tactile main features, and training the context encoder, context routing module, dendritic modulator, and task output head; a second stage, unfreezing the back-layer network of the tactile dual-code main pathway and performing joint training with a learning rate lower than that of the first stage; and a third stage, randomly discarding at least one context source data during training and performing robust enhancement training based on context loss consistency. Through staged training, the stable representational ability of the tactile main pathway can be maintained first, and then the tactile main pathway can be gradually adapted to context modulation, thereby reducing the risk of excessive interference of context branches with tactile main features in the early stages of training.
[0036] The present invention also provides a magnetic tactile sensing system based on dual-code dendritic modulation, including a magnetic tactile array sensor, a context acquisition module, a magnetic tactile feature construction module, a tactile dual-code main path module, a context encoding module, a context routing module, a dendritic modulation module, and a task output module.
[0037] The magnetic tactile array sensor is installed at the robot's end effector to acquire a three-axis magnetic tactile temporal array. This array includes multiple time frames, multiple array nodes, and a three-axis magnetic field component corresponding to each node. The context acquisition module acquires contextual source data and modal availability markers related to the current contact task. The contextual source data includes at least two types of candidate contextual data from visual context, linguistic context, audio context, and proprioceptive context. The magnetic tactile feature construction module constructs extended magnetic tactile features based on the three-axis magnetic tactile temporal array. These extended features include original three-axis magnetic field features, magnetic field modulus features, temporal difference features, high-frequency residual features, and exponential smoothing features. The tactile dual-code main path module extracts spatial code tactile features through a spatial code branch and temporal code tactile features through a temporal code branch. It then performs dual-code gating fusion on the spatial code and temporal code tactile features to obtain the main tactile features. The context encoding module inputs candidate contextual data in an available state into the corresponding context encoder to obtain multiple candidate contextual features. The context routing module is used to perform routing scoring and weight normalization on the multiple context candidate features based on the modal availability marker to obtain context weights, and to obtain context aggregate features based on the context weights. The dendritic modulation module is used to generate gain vectors and bias vectors based on the context aggregate features, and to modulate the tactile main features using the gain vectors and bias vectors to obtain context-enhanced tactile features. The task output module is used to output one or more of the following based on the context-enhanced tactile features: contact position, normal deformation, tangential slip, and contact mode.
[0038] In one optional embodiment, the system further includes at least two of a visual acquisition module, a language input module, an audio acquisition module, and a proprioceptive acquisition module. The visual acquisition module is used to acquire information about the target location, target edge, target category, or expected contact area; the language input module is used to acquire task instructions, action intentions, or contact strength constraint information; the audio acquisition module is used to acquire information about contact sounds, friction sounds, knocking sounds, or collision sounds; and the proprioceptive acquisition module is used to acquire information about the robot's end effector velocity, joint posture, motion stage, or force control state.
[0039] In one optional embodiment, the system further includes a robot control module, which is used to adjust the gripping force, movement speed, contact posture, slip compensation, or obstacle avoidance action of the robot end effector based on one or more of the contact position, normal deformation, tangential slip, and contact mode. Thus, the magnetotactile sensing results of the present invention can not only be used for contact state recognition but also further participate in the closed-loop control of the robot end effector.
[0040] In one optional implementation, the task output module is further configured to output one or more of context weights, gain vectors, bias vectors, and context-enhanced tactile features as interpretable intermediate results of the robot's contact state perception process. By outputting the aforementioned interpretable intermediate results, the degree of influence of different context source data on the main tactile features in the current contact task can be analyzed.
[0041] The present invention also provides a robot, including a robot body, a robot end effector, and the aforementioned magnetic tactile perception system based on dual-code dendritic modulation, wherein the magnetic tactile array sensor is disposed on the robot end effector. The robot can acquire contact state based on the magnetic tactile array sensor and, combined with contextual information such as vision, language, audio, and proprioception, perform context-enhanced tactile perception of the current contact task, thereby adjusting the gripping force, movement speed, contact posture, slip compensation, or obstacle avoidance actions of the robot end effector.
[0042] The present invention also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the above-described magnetic tactile perception method based on dual-code dendritic modulation, or implements the above-described magnetic tactile network training method based on dual-code dendritic modulation.
[0043] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described magnetic tactile perception method based on dual-code dendritic modulation, or implements the above-described magnetic tactile network training method based on dual-code dendritic modulation.
[0044] Compared with the prior art, the present invention has at least the following beneficial effects:
[0045] 1) This invention is based on a three-axis magnetic tactile temporal array acquired by a magnetic tactile array sensor at the robot end effector. It constructs magnetic tactile extended features that can reflect the original magnetic field distribution, magnetic field strength, dynamic changes, high-frequency disturbances, and slow-changing trends. Then, it extracts the spatial and temporal representations of magnetic tactile sensation through spatial and temporal code branches, respectively, and obtains the main tactile features through dual-code gating fusion. At the same time, it acquires at least two types of candidate context data from visual context, linguistic context, audio context, and proprioceptive context, and performs routing scoring and weighted aggregation on the candidate context features according to modal availability labels to obtain context aggregated features. Finally, it generates gain and bias vectors from the context aggregated features, and enhances, suppresses, or corrects the main tactile features in a feature dimension-level modulation manner to obtain context-enhanced tactile features, and outputs robot contact state perception results such as contact position, normal deformation, tangential slip, and contact mode.
[0046] 2) This invention constructs original triaxial magnetic field features, magnetic field modulus features, time difference features, high-frequency residual features, and exponential smoothing features based on a triaxial magnetic tactile temporal array. This allows the static distribution, dynamic changes, high-frequency disturbances, and slow-changing trends in the magnetic tactile signal to be explicitly expressed, improving the clarity of the physical meaning of the magnetic tactile features and the ability to distinguish contact states. Simultaneously, this invention extracts spatial and temporal tactile features through spatial and temporal code branches respectively, and obtains the main tactile features through dual-code gating fusion. This enables the main magnetic tactile features to simultaneously express information such as the contact center, contact area diffusion, array spatial distribution, sliding direction, speed changes, vibration disturbances, and the contact establishment process, thereby improving the recognition ability of states such as pressing, sliding, vibration, and composite contact. Finally, this invention uses the main tactile features as the basis for contact state judgment, and uses visual context, linguistic context, audio context, and proprioceptive context as conditional modulation information, rather than simply splicing tactile features with other modalities. This avoids the excessive replacement of tactile physical basis by high-dimensional contextual features, improving the reliability of the perception results.
[0047] 3) Furthermore, this invention adaptively determines the contribution of different contextual source data to the current contact task through modal availability labeling, contextual routing scoring, and weight normalization mechanisms. It also performs suppression or backoff processing when some or all contextual information is missing, improving the system's robustness in real-world scenarios such as visual occlusion, missing language, audio noise, and proprioceptive asynchrony. This invention utilizes contextual aggregation features to generate gain and bias vectors, and uses these vectors to perform feature-dimensional modulation of the tactile main features, enabling the same tactile signal to receive different semantic interpretations in different task contexts, thus improving the robot's task adaptability in tactile perception. This invention improves the training stability, modulation controllability, and contextual missing robustness of the magnetic tactile perception network through a training method that includes continuous regression loss, contact pattern classification loss, gain stabilization loss, bias amplitude loss, contextual missing consistency loss, and contextual routing regularization loss. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the system structure of the magnetic tactile sensing system based on dual-code dendritic modulation according to the present invention;
[0049] Figure 2 This is a flowchart illustrating the magnetic tactile sensing method based on dual-code dendritic modulation of the present invention.
[0050] Figure 3 This is a schematic diagram of the construction process of the magnetic tactile extension feature in this invention;
[0051] Figure 4 This is a schematic diagram of the tactile dual-code main pathway in this invention;
[0052] Figure 5 This is a schematic diagram of the context routing and dendritic modulation process in this invention;
[0053] Figure 6 This is a schematic diagram illustrating an application scenario of the present invention;
[0054] Figure 7 This is a schematic diagram of the training process of the present invention. Detailed Implementation
[0055] The present invention will be further described in detail below with reference to embodiments. It should be understood that the following embodiments are only used to explain the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Equivalent substitutions, simple modifications, or conventional adjustments made by those skilled in the art based on the content of this specification without departing from the concept of the present invention should all fall within the scope of protection of the present invention.
[0056] Example 1
[0057] As shown in the figure, this embodiment provides a magnetic tactile sensing method based on dual-code dendritic modulation, which is applied to scenarios where a robot's end effector performs contact tasks. The contact tasks can include grasping, pressing, sliding detection, edge tracking, insertion assembly, human-robot interaction, or other tasks requiring the robot's end effector to determine the contact state.
[0058] The robot's end effector is equipped with a magnetic tactile array sensor. The magnetic tactile array sensor comprises multiple array nodes, each used to acquire three-axis magnetic field components. The robot can also be configured with at least two of the following: a vision acquisition module, a speech input module, an audio acquisition module, and a proprioception acquisition module. The vision acquisition module can be an RGB camera, a depth camera, or an RGB-D camera; the speech input module can be a text input interface, a speech recognition module, or a task planning module; the audio acquisition module can be a single-channel microphone or a microphone array; and the proprioception acquisition module can be a robot controller, a joint encoder, an end effector velocity estimation module, or a force control status reading module.
[0059] The method in this embodiment includes the following steps.
[0060] S1. Acquire the perception data generated by the robot end effector when performing a contact task. The perception data includes a three-axis magnetic tactile temporal array collected by a magnetic tactile array sensor installed on the robot end effector, as well as contextual source data and modal availability markers related to the current contact task. The three-axis magnetic tactile temporal array includes multiple time frames, multiple array nodes, and a three-axis magnetic field component corresponding to each array node. Before or after constructing the three-axis magnetic tactile temporal array, the three-axis magnetic field components are subjected to contactless baseline subtraction, zero-point drift compensation, and normalization to obtain a baseline-corrected three-axis magnetic tactile temporal array. The contextual source data includes at least two types of candidate contextual data from visual context, linguistic context, audio context, and proprioceptive context. The modal availability markers are used to characterize whether each type of candidate contextual data is available in the current inference sample.
[0061] When the robot's end effector performs a contact task, the magnetotactile array sensor collects triaxial magnetic field data at a preset sampling frequency. The sampling frequency can be set from 50Hz to 500Hz, preferably 100Hz or 200Hz. During each inference, a time window is selected from T consecutive frames of triaxial magnetic field data, where T can be set from 16 to 128, preferably 32. The magnetotactile array can be 3×3, 4×4, 5×5, or other array configurations; this embodiment uses a 3×3 array as an example. Each array node outputs magnetic field components in the Bx, By, and Bz directions in each time frame.
[0062] The triaxial magneto-tactile temporal array can be represented as a tensor of size B×T×H×W×C, where B represents the batch size, T represents the time window length, H and W represent the number of rows and columns of the magneto-tactile array, respectively, and C represents the number of magnetic field channels, with C=3. For example, in a single online inference, B=1, T=32, H=3, W=3, and C=3, then the triaxial magneto-tactile temporal array is a tensor of size 1×32×3×3×3. In offline training, B can be set to 16, 32, or 64.
[0063] The above-mentioned triaxial magnetic tactile timing array can be further written in the following tensor form:
[0064]
[0065] While acquiring a three-axis magnetotactile temporal array, contextual source data and modal availability tags related to the current contact task are obtained. The contextual source data includes at least two types of candidate contextual data from visual context, linguistic context, audio context, and proprioceptive context.
[0066] Visual context can include target location, target edge, target category, or expected contact area information. For example, the vision acquisition module identifies the location of the cup to be grasped, the edge of the cup wall, the cup category, and the expected contact area of the gripper. Language context can include task instructions, action intentions, or contact strength constraints. For example, language input could be "gently grip the cup," "slide slowly along the workpiece edge," "determine if the object is slipping," or "insert into the hole and avoid collision." Audio context can include contact sounds, friction sounds, knocking sounds, or collision sound information. Proprioceptive context can include robot end-effector velocity, joint posture, motion stage, gripper opening / closing amount, or force control status information.
[0067] Modal availability markers are used to characterize whether various candidate context data are available in the current inference sample. Specifically, when the visual image is acquired normally and the target detection confidence is greater than a preset threshold, the modal availability marker corresponding to the visual context is set to available; when the vision is obstructed, the lighting is abnormal, or the target detection confidence is lower than a preset threshold, the modal availability marker corresponding to the visual context is set to unavailable. When the language command is empty or the speech recognition confidence is lower than a preset threshold, the modal availability marker corresponding to the language context is set to unavailable. When the audio signal-to-noise ratio is lower than a preset threshold or the environmental noise is too strong, the modal availability marker corresponding to the audio context is set to unavailable. When the proprioceptive data delay exceeds a preset time threshold, the robot controller does not return status information, or the status timestamps are out of sync, the modal availability marker corresponding to the proprioceptive context is set to unavailable.
[0068] Modal availability tags can be represented by a vector consisting of 0s and 1s. For example, if visual context, linguistic context, audio context, and proprioceptive context are all available, the modal availability tag can be [1, 1, 1, 1]; if visual context is unavailable, linguistic and proprioceptive contexts are available, and audio context is unavailable, the modal availability tag can be [0, 1, 0, 1].
[0069] The modal availability tag can be represented as the following vector:
[0070]
[0071] In one specific implementation, to reduce the impact of initial installation bias, changes in the ambient magnetic field, and zero-point drift of the magnetic tactile array sensor on subsequent sensing results, at least N0 frames of triaxial magnetic field data are collected as non-contact baseline data when the robot's end effector is not in contact with the target object, and the mean non-contact baseline value of each array node in the three directions of Bx, By, and Bz is calculated respectively:
[0072]
[0073] in, Represents array node The mean value of the non-contact baseline in the c-th magnetic field direction, Indicates the nth frame and the nth frame in the contactless state. Array node, magnetic field component in the c-th direction.
[0074] During online inference or offline training, the mean of the non-contact baseline for the corresponding array node and the corresponding magnetic field direction is subtracted from the currently acquired triaxial magnetic field components to obtain the baseline-subtracted triaxial magnetic field components:
[0075]
[0076] in, This represents the magnetic field component after baseline subtraction.
[0077] Furthermore, the triaxial magnetic field components after baseline subtraction can be normalized based on the mean and standard deviation of each magnetic field channel in the training sample set:
[0078]
[0079] in, This represents the normalized magnetic field components. Let represent the mean of the c-th magnetic field direction in the training sample set. This represents the standard deviation of the c-th magnetic field direction in the training sample set. It is a very small positive number, used to avoid the denominator being zero.
[0080] While acquiring a three-axis magnetotactile temporal array, contextual source data and modal availability tags related to the current contact task are obtained. Contextual source data includes at least two candidate contextual data types from visual context, linguistic context, audio context, and proprioceptive context. Modal availability tags can be represented as:
[0081]
[0082] in, , , , These represent whether the visual context, language context, audio context, and proprioceptive context are available, respectively; 1 is taken when the corresponding context is available, and 0 is taken when the corresponding context is unavailable.
[0083] In one specific embodiment, labels for contact position, normal deformation, and tangential slip are obtained through a calibration device. The calibration device includes a fixture for fixing the magnetotactile array sensor, a displacement platform capable of applying a known indentation along the normal direction, a two-dimensional translation platform capable of applying a known displacement along the tangential direction of the sensor surface, and a displacement sensor or motion controller for recording the indentation and tangential displacement. During calibration, different normal indentations and different tangential displacements are applied to multiple preset positions on the surface of the magnetotactile array sensor, and a three-axis magnetotactile time-series array is simultaneously acquired. The preset positions where contact is applied are used as contact position labels, the normal indentation is used as normal deformation labels, and the tangential displacement or the change in tangential displacement within adjacent time windows is used as tangential slip labels. For data acquired in actual robot tasks, corresponding continuous quantity labels can be generated based on the external displacement platform, robot end-effector pose, gripper displacement, force sensor, manual annotation, or high-confidence rule control results.
[0084] S2. Based on the triaxial magnetic tactile timing array, construct magnetic tactile extended features for each time frame and each array node. The magnetic tactile extended features include original triaxial magnetic field features, magnetic field magnitude features, time difference features, high-frequency residual features, and exponential smoothing features.
[0085] For each time frame and each array node in the triaxial magnetic tactile temporal array, the normalized triaxial magnetic field components are first retained as the original triaxial magnetic field features:
[0086]
[0087] in, Indicates the array node in frame t. The original triaxial magnetic field characteristics at the location.
[0088] Furthermore, the magnetic field modulus characteristics are calculated. These characteristics are obtained by taking the square root of the sum of the squares of the three-axis magnetic field components at the same time frame and the same array node:
[0089]
[0090] in, Indicates the array node in frame t. The magnetic field modulus characteristics at that location, It is a very small positive number. The magnetic field modulus characteristic is used to reflect the overall strength of the magnetic field change at the array node, thereby characterizing the local contact strength, the degree of compression, or the degree of magnetic structure deformation.
[0091] Furthermore, the temporal difference characteristics are calculated. These characteristics are obtained based on the differences in the three-axis magnetic field components of the same array node in adjacent time frames:
[0092]
[0093] Right now:
[0094]
[0095] in, Indicates the array node in frame t. The three-axis temporal difference features are used. For the first frame in the time window, the temporal difference features can be set to a zero vector, or the temporal difference features of the second frame can be copied.
[0096] Furthermore, the exponential smoothing feature is calculated. The exponential smoothing feature is obtained by performing an exponential moving average along the time dimension based on the magnetic field modulus feature:
[0097]
[0098] in, Indicates the array node in frame t. Exponential smoothing characteristics at that point. This represents the smoothing coefficient, with a value ranging from 0.1 to 0.5, preferably 0.25. For the first frame in the time window, we can set:
[0099]
[0100] Furthermore, the high-frequency residual characteristics are calculated. These characteristics are derived from the difference between the magnetic field modulus characteristics and the exponential smoothing characteristics:
[0101]
[0102] in, Indicates the array node in frame t. The high-frequency residual characteristics at the location. These high-frequency residual characteristics can highlight information such as frictional micro-vibrations, instantaneous slippage, collisional abrupt changes, and rapid changes in contact boundaries.
[0103] Furthermore, the characteristics of the magnetic field modulus variation are calculated:
[0104]
[0105] in, Indicates the array node in frame t. Characteristics of magnetic field modulus variation at a given location.
[0106] Through step S2, the original triaxial magnetotactile temporal array is expanded into a magnetotactile extended feature that includes the original triaxial magnetic field characteristics, magnetic field magnitude characteristics, time difference characteristics, high-frequency residual characteristics, and exponential smoothing characteristics. This magnetotactile extended feature simultaneously expresses static spatial distribution, dynamic changes, high-frequency perturbations, and slow variation trends, providing input for subsequent spatial code branches and temporal code branches.
[0107] S3. Input the magnetic tactile extended features into the tactile dual-code main path, extract spatial code tactile features from the magnetic tactile extended features through the spatial code branch, extract temporal code tactile features from the magnetic tactile extended features through the temporal code branch, and perform dual-code gating fusion on the spatial code tactile features and temporal code tactile features to obtain the tactile main features.
[0108] The tactile dual-code main path includes a spatial code branch, a temporal code branch, and a dual-code gating fusion module. The spatial code branch is mainly used to extract spatial information such as the contact center, contact area diffusion, and array spatial distribution. The temporal code branch is mainly used to extract temporal information such as sliding direction, speed change, vibration disturbance, and contact establishment process. The dual-code gating fusion module is used to adaptively fuse spatial code tactile features and temporal code tactile features according to the current contact state.
[0109] S3.1, extract spatial code tactile features from the magnetic tactile extended features through spatial code branching; wherein, the spatial code branching takes node features including the original triaxial magnetic field features, magnetic field modulus features and exponential smoothness features as input, and performs spatial topology modeling based on the local adjacency relationship between array nodes in the magnetic tactile array to obtain spatial code tactile features reflecting the contact center position, contact area diffusion and array spatial distribution.
[0110] In the spatial code branch, for each time frame, each array node in the magnetic tactile array is considered a graph node. The node features of each graph node can include the original triaxial magnetic field features, magnetic field magnitude features, and exponential smoothing features, represented as:
[0111]
[0112] After processing by the node embedding module, local graph attention module, node attention pooling module, and temporal attention pooling module, the spatial code tactile features are obtained:
[0113]
[0114] in, Represents the tactile characteristics of spatial codes. This represents the spatial feature extraction function corresponding to the spatial code branch. The local graph attention module only interacts with array nodes that satisfy a preset adjacency relationship. The preset adjacency relationship is as follows: for any array node, array nodes whose absolute difference between their row position and column position is not greater than 1 and whose absolute difference between their column positions is not greater than 1 are determined as adjacent nodes.
[0115] In the timecode branch, for each time frame, the time difference features, high-frequency residual features, and magnetic field magnitude variation features of all array nodes are expanded and concatenated to form the time-series input features of that time frame:
[0116]
[0117] in, This represents the temporal input feature corresponding to frame t. After processing by a single-frame projection module, a temporal convolution module, a recurrent dynamic coding module, and a temporal attention pooling module, the temporal code tactile feature is obtained:
[0118]
[0119] in, Indicates time-code tactile characteristics. This represents the time-series feature extraction function corresponding to the timecode branch.
[0120] After obtaining the spatial code tactile features and the temporal code tactile features, the two are concatenated to obtain the dual-code concatenated features:
[0121]
[0122] The concatenated features of the two codes are input into the gating network to obtain the spatial code weights and temporal code weights:
[0123]
[0124] in, Indicates spatial code weights, Represents the timecode weight and satisfies:
[0125]
[0126] Subsequently, the spatial code tactile features and temporal code tactile features are weighted and summed according to spatial code weight and temporal code weight to obtain the main tactile features:
[0127]
[0128] in, This represents the primary tactile feature. Through the aforementioned dual-code gating fusion, the spatial code can have a relatively high weight in static pressing or contact positioning tasks; while the temporal code can have a relatively high weight in sliding detection, anti-slip control, or surface friction detection tasks. This allows the primary tactile feature to adaptively adjust the contribution ratio of spatial and temporal information according to the current contact state.
[0129] S3.3, perform dual-code gating fusion on the spatial code tactile features and the temporal code tactile features to obtain the main tactile feature; wherein, the dual-code gating fusion includes: concatenating the spatial code tactile features and the temporal code tactile features to obtain a dual-code concatenated feature; inputting the dual-code concatenated feature into a gating network to obtain spatial code weights and temporal code weights; and performing a weighted summation of the spatial code tactile features and the temporal code tactile features according to the spatial code weights and temporal code weights to obtain the main tactile feature.
[0130] Specifically, spatial code tactile features and temporal code tactile features are concatenated to obtain dual-code concatenated features. These features are then input into a gating network to obtain spatial code weights and temporal code weights. The gating network may include fully connected layers, nonlinear activation layers, and a normalized output layer. The normalized output layer may employ a softmax function to ensure that the sum of the spatial code weights and temporal code weights is 1. Subsequently, the spatial code tactile features and temporal code tactile features are weighted and summed according to their respective weights to obtain the main tactile features.
[0131] Dual-code gating fusion can be further represented as:
[0132]
[0133]
[0134]
[0135] For example, when a robot performs static pressing or contact positioning tasks, the spatial code weight can be relatively high to highlight the contact area distribution and the contact center position; when a robot performs sliding detection, anti-slip control, or surface friction detection tasks, the temporal code weight can be relatively high to highlight the sliding direction, speed changes, and high-frequency disturbances. Through dual-code gating fusion, the tactile main features can adaptively adjust the contribution of spatial and temporal information according to the current contact state.
[0136] S4. Input the context source data into the corresponding context encoder. Generate corresponding context candidate features for candidate context data in an available state. Generate preset placeholder features or default features for candidate context data in an unavailable state. Map candidate context data of different categories to the same context feature dimension through the corresponding context encoder or linear mapping layer to obtain multiple context candidate features.
[0137] Context encoders include at least two of the following: visual context encoders, language context encoders, audio context encoders, and proprioceptive context encoders. Different context encoders may have different input dimensions, but the same output dimension, so that multiple context candidate features can be routed and aggregated in the same context feature space.
[0138] Specifically, the visual context can be acquired by the visual acquisition module and extracted by an object detection network, an image coding network, or a visual Transformer. The visual context can include information about the target location, target edges, target category, or expected contact area. For example, when grasping a cup-shaped object, the visual context can include the cup's edge, the cup wall area, and the expected contact area of the gripper.
[0139] The linguistic context can be obtained from text input, speech recognition, or task planning modules, and extracted through word embedding models, recurrent encoders, or language encoders. The linguistic context can include task instructions, action intentions, or contact strength constraints. For example, "gently clamp the cup" corresponds to a low contact strength constraint, "stable handling of an object" corresponds to a high requirement for grasping stability, and "detecting burrs along the edge" corresponds to the task objectives of edge tracking and high-frequency disturbance detection.
[0140] The audio context can be acquired by the audio acquisition module and extracted through short-time Fourier transform, Mel spectrum extraction, one-dimensional convolutional network, or audio encoder. The audio context can include information on contact sound, friction sound, knocking sound, or collision sound.
[0141] Proprioceptive context can be read from the robot controller and includes information such as robot end-effector velocity, joint posture, motion phase, gripper opening / closing, or force control status. For example, during the handling phase, end-effector velocity and gripper opening / closing can help determine if there is a risk of slippage.
[0142] In one specific implementation, the original features of visual context can be 128-dimensional, the original features of language context can be 128-dimensional, the original features of audio context can be 64-dimensional, and the original features of proprioceptive context can be 16-dimensional. Through the corresponding context encoder, the above-mentioned context features of different dimensions are all mapped to 128-dimensional context candidate features. Thus, different context candidate features can be routed, scored, weighted, and aggregated in the same context feature space, avoiding a certain modality from having an unreasonable advantage in the fusion process simply because of its higher original dimension.
[0143] For the m-th type of context source data, its input can be represented as: Modal availability can be represented as .when When the context source data is input into the corresponding context encoder, context candidate features are obtained:
[0144]
[0145] in, This represents the m-th type of context encoder. Let m represent the contextual candidate feature of the m-th class. When When this occurs, it indicates that the context source data for that type is unavailable, and the corresponding context candidate feature can be set to a preset placeholder feature, a zero vector, or a default feature:
[0146]
[0147] By using a corresponding context encoder or linear mapping layer, candidate context data of different categories are mapped to the same context feature dimension, resulting in multiple context candidate features. Thus, different context candidate features can be routed, scored, weighted, and aggregated within the same context feature space, preventing a particular modality from having an unreasonable advantage in the fusion process simply because its original dimension is higher.
[0148] S5. Perform context routing scoring and weight normalization on the multiple context candidate features according to the modal availability label to obtain context weights, and perform weighted aggregation on the multiple context candidate features according to the context weights to obtain context aggregate features.
[0149] Specifically, a routing score is calculated for each context candidate feature. This routing score can be obtained by a routing scoring network, which may include fully connected layers, non-linear activation layers, and linear output layers. For each context candidate feature, the routing scoring network outputs a scalar score, which represents the importance of that contextual information to the current magnetotactic interpretation.
[0150] For each type of contextual candidate feature, the route scoring network calculates the corresponding route score:
[0151]
[0152] in, Represents the routing score of the context candidate feature of class m. This represents the routing score network. Routing scores are suppressed based on modal availability tags.
[0153]
[0154] in, This represents the route score after suppression processing, where C is a sufficiently large positive number, such as 10000, so that the weight of unavailable context data in subsequent normalization is close to zero.
[0155] Subsequently, the route scores after suppression are normalized to obtain the context weights corresponding to each candidate context data:
[0156]
[0157] in, This represents the context weight corresponding to the m-th class of context data. The context aggregate features are obtained by weighted summation of the available context candidate features based on their context weights.
[0158]
[0159] in, This indicates context-aggregated features.
[0160] When all candidate context data is unavailable, i.e.:
[0161]
[0162] Set the context aggregation features to a preset zero vector or a preset default vector, and make the dendritic modulator output an all-1 gain vector and an all-0 bias vector:
[0163]
[0164] At this point, the context-enhanced tactile features become equal to the primary tactile features, causing the magnetotactile perception method to revert to the unmodulated primary tactile pathway output. Through this reversion mechanism, even with visual occlusion, missing language, excessive audio noise, or proprioceptive asynchrony, the system can still output basic contact state perception results based on the primary magnetotactile pathway.
[0165] S6. Input the context aggregation feature into the dendritic modulator, which generates a gain vector and a bias vector corresponding to the tactile main feature dimension, and uses the gain vector and bias vector to perform feature dimension-level modulation on the tactile main feature to obtain context-enhanced tactile features.
[0166] The dendritic modulator includes a gain generation branch and a bias generation branch. The gain generation branch generates a gain vector based on contextual aggregated features, while the bias generation branch generates a bias vector based on contextual aggregated features. The dimensions of both the gain and bias vectors correspond to the dimensions of the haptic master features. For example, when the haptic master features are 128-dimensional, both the gain and bias vectors are also 128-dimensional.
[0167] The gain generation branch can include a first fully connected layer, a nonlinear activation layer, and a second fully connected layer. The bias generation branch can also include a first fully connected layer, a nonlinear activation layer, and a second fully connected layer. The nonlinear activation layer can use ReLU, GELU, tanh, or other nonlinear activation functions. To avoid excessive alteration of the haptic main features by contextual information, the gain vector can be generated using residual gain, causing the gain vector to vary around unity gain. To avoid excessive alteration of the haptic main features by contextual information, the gain vector is generated using residual gain.
[0168]
[0169] in, Represents the gain vector. This represents the gain generation function. This represents the gain modulation amplitude coefficient. It can be set to 0.1 to 1.0, preferably 0.5. This formula causes the gain vector to vary around unity gain, thereby preventing contextual information from completely covering the main tactile features.
[0170] The bias vector is generated using an amplitude-limited method:
[0171]
[0172] in, This represents the bias vector. This represents the bias generating function. This represents the bias amplitude coefficient. It can be set to 0.05 to 0.2, preferably 0.1. This formula is used to limit the magnitude of the bias vector to avoid excessive translation of the tactile main features by the context information.
[0173] Subsequently, the tactile main features are modulated at the feature dimension level using the gain vector and bias vector to obtain context-enhanced tactile features:
[0174]
[0175] in, This indicates context-enhanced haptic features. Indicates the main tactile features, Represents the gain vector. This represents the bias vector. This indicates element-wise multiplication along the feature dimension.
[0176] In this way, contextual information does not directly replace tactile information, nor is it simply concatenated with tactile features. Instead, it is used to amplify, suppress, or shift and correct different feature dimensions in the main tactile features. For example, when the verbal instruction is "touch to confirm," the dendritic modulator can enhance tactile feature dimensions related to slight contact and normal deformation threshold, while suppressing tactile feature dimensions related to forceful grasping. When the task is "detecting slippage," the dendritic modulator can enhance tactile feature dimensions related to tangential slip, temporal difference, and high-frequency residuals. When the visual context indicates that the robot's end effector is approaching the target edge, the dendritic modulator can enhance tactile feature dimensions related to the contact center location and local spatial distribution.
[0177] S7, input the context-enhanced tactile features into the task output head and output the robot contact state perception result. The task output head includes one or more of the following: contact position output head, normal deformation output head, tangential slip output head, and contact mode output head. The contact position is the position of the contact center in the magnetic tactile array coordinate system or the physical coordinate system of the sensor surface. The normal deformation is the amount of indentation or normalized indentation of the flexible contact layer of the magnetic tactile sensor along the normal direction. The tangential slip is the amount of displacement, cumulative slip, or slip trend of the contact center along the tangential direction of the sensor surface. The contact mode includes one or more of the following: no contact, static pressing, sliding contact, collision contact, and composite contact.
[0178] The task output head can include a task fusion layer, a continuous regression head, and a contact pattern classification head. First, the context-enhanced haptic features and the context aggregation features are fused to obtain the task fusion features. The fusion method can be concatenation followed by input to a fully connected layer, or weighted fusion or attention fusion. In one specific implementation, the context-enhanced haptic features and the context aggregation features are concatenated and then input to a task fusion layer consisting of a fully connected layer, a non-linear activation layer, and a normalization layer to obtain the task fusion features.
[0179] Task fusion features can be represented as:
[0180]
[0181] in, Indicates the characteristics of task output. This indicates the task fusion layer. This indicates a feature splicing operation.
[0182] The continuous regression head outputs continuous parameters such as contact position, normal deformation, and tangential slip based on the task output characteristics.
[0183]
[0184] In one specific implementation, the prediction result of continuous quantities can be expressed as:
[0185]
[0186] in, and This indicates the position of the contact center in the magnetic tactile array coordinate system or the physical coordinate system of the sensor surface. This indicates normalized normal deformation. and This represents the slip components of the contact center along the sensor surface in two tangential directions. The tangential slip amount can be determined according to... and The calculation yielded:
[0187]
[0188] in, Indicates the tangential slip. It is a very small positive number.
[0189] The contact pattern classification head outputs a contact pattern probability distribution based on the task output features:
[0190]
[0191] in, Represents the probability distribution of contact patterns. This indicates a contact pattern classification head. Contact patterns can include one or more of the following: no contact, static pressing, sliding contact, collision contact, and combined contact. Ultimately, the robot's contact state perception results include one or more of the following: contact position, normal deformation, tangential slip, and contact pattern.
[0192] In one specific implementation, the robot controller generates control adjustment quantities based on the contact position, normal deformation, tangential slip, and contact mode output by the task output head. When the contact position deviates from the expected contact area, the robot controller adjusts the lateral position of the end effector to bring the contact center back to the expected contact area; when the normal deformation is below the target deformation range, the robot controller increases the end effector normal feed or gripper closure; when the normal deformation is above the target deformation range, the robot controller decreases the end effector normal feed or reduces the gripper clamping force; when the tangential slip exceeds a preset slip threshold, the robot controller increases the gripping force, reduces the end effector movement speed, or adjusts the contact posture; when the contact mode is a collision or abnormal compound contact, the robot controller performs deceleration, avoidance, or stopping actions. Thus, the magnetotactile perception results and the robot end effector control actions form a closed-loop correlation.
[0193] S8, output one or more of the context weights, gain vectors, bias vectors and context-enhanced tactile features as interpretable intermediate results to characterize the modulation effect of different context source data on the current robot contact state perception results.
[0194] While completing the output of the contact state perception results, it can also output one or more of the following as interpretable intermediate results: context weights, gain vectors, bias vectors, and context-enhanced tactile features, which can be used to characterize the modulation effect of different context source data on the current robot contact state perception results.
[0195] For example, by observing the context weights, we can determine whether the current inference sample relies more on visual, linguistic, audio, or proprioceptive context; by observing the gain vector, we can determine which tactile feature dimensions the context enhances; and by observing the bias vector, we can determine how the context semantically shifts and corrects the main tactile features. These interpretable intermediate results can be used for model debugging, task analysis, anomaly diagnosis, and robot control strategy optimization.
[0196] Example 2
[0197] This embodiment provides a training method for a magnetic tactile network based on dual-code dendritic modulation. T1. Construct a training sample set. Each training sample in the training sample set includes a triaxial magnetic tactile temporal array, contextual source data, modal availability markers, continuous quantity labels, and contact pattern labels. Continuous quantity labels include one or more of contact position labels, normal deformation labels, and tangential slip quantity labels; contact pattern labels include at least two of non-contact, static pressing, sliding contact, collision contact, and composite contact.
[0198] T2. Input the triaxial magnetic tactile timing array into the tactile dual-code main path to obtain the tactile main features.
[0199] T3. Input the context source data and modal availability tags into the context encoding and routing path to obtain the context aggregation features.
[0200] T4. Generate gain vectors and bias vectors based on contextual aggregation features, and use the gain vectors and bias vectors to modulate the tactile main features to obtain context-enhanced tactile features.
[0201] T5. Obtain continuous quantity prediction results and contact pattern prediction results by enhancing tactile features based on context.
[0202] T6. Train the magnetic tactile perception network based on the total training loss. The total training loss includes at least two of the following: continuous regression loss, contact pattern classification loss, gain stabilization loss, bias magnitude loss, context missing consistency loss, and context routing regularization loss. The total training loss can be expressed as:
[0203]
[0204] in, Indicates total training loss. This represents the loss during continuous regression. This represents the contact pattern classification loss. Indicates the gain stabilization loss. Indicates the bias amplitude loss. This represents the context-missing consistency loss. Indicates the context routing regular expression loss; , , , , and These represent the weighting coefficients of the corresponding loss terms.
[0205] The regression loss for continuous quantities is determined based on the difference between the continuous quantity prediction results and the continuous quantity labels, and can be achieved using the Smooth L1 loss.
[0206]
[0207] in, This indicates the number of samples in the training batch. This represents the continuous prediction result for the nth training sample. This represents the continuous label of the nth training sample.
[0208] The contact pattern classification loss is determined based on the difference between the predicted probability of the contact pattern and the contact pattern label, and cross-entropy loss can be used:
[0209]
[0210] in, Indicates the number of contact pattern categories. This represents the true label of the c-th contact pattern in the n-th training sample. This represents the predicted probability of the c-th contact pattern in the n-th training sample.
[0211] Gain stabilization loss is used to constrain the deviation of the gain vector from unity gain, and can be expressed as:
[0212]
[0213] in, This represents the gain vector corresponding to the nth training sample.
[0214] The bias magnitude loss, used to constrain the magnitude of the bias vector, can be expressed as:
[0215]
[0216] in, This represents the bias vector corresponding to the nth training sample.
[0217] The missing context consistency loss is determined by the output difference between full context forward inference and missing context forward inference. For the same training sample, continuous output is obtained under full context conditions. and contact mode probability distribution Randomly discard one or more of the visual context, linguistic context, audio context, or proprioceptive context, and perform forward inference based on the updated modal availability tag to obtain a continuous output. and contact mode probability distribution The context-missing consistency loss can be expressed as:
[0218]
[0219] in, Indicates the KL divergence. This represents the weight of the probability distribution consistency loss.
[0220] Context routing regularization loss is used to constrain the distribution of context weights. In one implementation, it can be determined based on the difference between the distributions of context weights and target weights.
[0221]
[0222] in, This represents the context weight vector corresponding to the nth training sample. This represents the target context weight distribution corresponding to the nth training sample. The target context weight distribution can be set to a uniform distribution across available modalities based on modality availability labels, or it can be set to a distribution with different modality preferences based on task priors. This loss can reduce the risk of routing weights being concentrated on a single context modality in the long run, improving the stability and interpretability of context routing.
[0223] In one specific training process, the training method comprises three stages. The first stage freezes the tactile dual-code master pathway used to obtain tactile main features, and trains the context encoder, context routing module, dendritic modulator, and task output head. The second stage unfreezes the back-layer networks of the tactile dual-code master pathway and performs joint training using a learning rate lower than that of the first stage. The third stage randomly discards at least one context source data during training and performs robust enhancement training based on context loss consistency. Through staged training, the stable representational ability of the tactile master pathway can be maintained first, and then the tactile master pathway can be gradually adapted to context modulation, thereby reducing the risk of excessive interference of context branches with tactile main features in the early stages of training.
[0224] Example 3
[0225] This embodiment provides a magnetic tactile sensing system based on dual-code dendritic modulation, used to implement the magnetic tactile sensing method described in Embodiment 1.
[0226] The system includes a magnetic tactile array sensor, a context acquisition module, a magnetic tactile feature construction module, a tactile dual-code main path module, a context encoding module, a context routing module, a dendritic modulation module, a task output module, and a robot control module.
[0227] A magnetic tactile array sensor is installed at the robot's end effector to acquire a three-axis magnetic tactile temporal array. A context acquisition module acquires contextual source data and modal availability markers relevant to the current contact task. A magnetic tactile feature construction module constructs extended magnetic tactile features based on the three-axis magnetic tactile temporal array. A dual-code main path module extracts spatial code tactile features through a spatial code branch and temporal code tactile features through a temporal code branch, then performs dual-code gating fusion on the spatial and temporal code features to obtain the main tactile features. A context encoding module inputs candidate context data in an available state into the corresponding context encoder to obtain multiple candidate context features. A context routing module performs routing scoring and weight normalization on multiple candidate context features based on modal availability markers to obtain context weights, and then obtains context aggregate features based on these weights. A dendritic modulation module generates gain and bias vectors based on the aggregated context features, and uses these vectors to modulate the main tactile features, resulting in context-enhanced tactile features. The task output module is used to enhance haptic features based on context and output one or more of the following: contact position, normal deformation, tangential slip, and contact mode.
[0228] The robot control module adjusts the gripping force, movement speed, contact posture, slip compensation, or obstacle avoidance actions of the robot's end effector based on one or more of the following: contact position, normal deformation, tangential slip, and contact mode. For example, when the contact mode is stable pressing and the normal deformation reaches the target range, the robot control module maintains the current contact posture; when a continuous increase in tangential slip is detected and classified as sliding or compound contact, the robot control module increases the gripping force or adjusts the gripper posture; when the contact mode is collision or abnormal contact, the robot control module reduces the end effector movement speed, performs obstacle avoidance actions, or stops the current task; when the contact position deviates from the expected contact area, the robot control module adjusts the end effector trajectory to bring the contact point back to the target area.
[0229] Example 4
[0230] This embodiment provides an application scenario where a robot grasps a cup-shaped object and performs anti-slip control.
[0231] In this scenario, the robot's end effector is a two-finger gripper, with a 3×3 magnetic tactile array sensor mounted on the fingertip of each gripper. An RGB-D camera is positioned above the robot to identify the position, outline, and intended contact area of the cup-shaped object. The operator inputs the verbal command, "Gently grip the cup and move it to the tray." The robot controller provides the gripper opening and closing amount, end-effector velocity, and motion phase as proprioceptive context. A microphone collects the friction or collision sounds between the gripper and the cup as audio context.
[0232] When the robot begins grasping, the magnetic tactile array sensor collects 32 consecutive frames of triaxial magnetic field data. Based on this data, the system constructs original triaxial magnetic field features, magnetic field modulus features, temporal difference features, high-frequency residual features, and exponential smoothing features. The spatial code branch identifies whether the current contact center is located within the expected contact area of the cup wall based on the magnetic field distribution of each node in the magnetic tactile array; the temporal code branch determines whether there is a slippage trend based on the temporal difference and high-frequency residual. The dual-code gating fusion module generates the main tactile features based on the current state.
[0233] Simultaneously, the visual context provides the cup wall position and the expected gripping area, the linguistic context provides the contact strength constraint of "gently gripping," the proprioceptive context provides the gripper closing speed and end effector movement stage, and the audio context provides friction sound characteristics. The context routing module calculates context weights based on modal availability tags and candidate context features. When the gripper first contacts the cup, the visual and linguistic contexts have higher weights, used to guide the contact position and limit the gripping force; as the robot begins to move the cup, the proprioceptive and audio contexts have higher weights, used in conjunction with tactile signals to determine whether slippage has occurred.
[0234] The dendritic modulator generates gain and bias vectors based on contextual aggregation features to modulate the main tactile features. When the verbal command instructs for a light touch, feature dimensions associated with excessive normal deformation are suppressed or corrected; when the timecode branch detects enhanced high-frequency residuals and the audio context detects friction sounds, feature dimensions associated with tangential slip are enhanced. The task output module outputs the contact position, normal deformation, tangential slip amount, and contact pattern. When the contact pattern changes from pressing to compound contact and the tangential slip amount exceeds a preset threshold, the robot control module increases the gripper force or adjusts the gripper posture until the tangential slip amount decreases to a safe range. If the visual context is unavailable due to occlusion, the modal availability flag sets the visual context to unavailable, the system automatically suppresses the visual context weights, and continues to perform anti-slip control relying on magnetotactile, proprioceptive, and audio context.
[0235] Example 5
[0236] This embodiment provides an application scenario for a robot to perform tactile tracking along the edge of a workpiece.
[0237] In this scenario, a flexible tactile probe with a magnetic tactile array sensor is mounted on the robot's end effector. The vision acquisition module identifies the approximate location of the workpiece edge, and the verbal command is "slowly slide along the workpiece edge and detect burrs." The proprioceptive context provides the end effector's speed and posture along the edge direction, and the audio context is used to acquire the sound generated by the friction between the probe and the workpiece surface.
[0238] During edge tracking, the spatial code branch determines whether the probe is off-center from the edge based on the contact area offset on the magneto-tactile array, while the temporal code branch determines whether burrs, protrusions, or abnormal friction occur based on temporal differences and high-frequency residuals. Visual context provides the initial edge position at the start of the task, linguistic context emphasizes the task objectives of "slow sliding" and "detecting burrs," proprioceptive context provides the end-effector sliding speed, and audio context provides abnormal friction sounds. The context routing module dynamically adjusts the weights of each context based on the current state. When the probe enters a visually occluded area, the visual context weight decreases; when both high-frequency residuals and audio friction sounds increase simultaneously, the audio context weight increases.
[0239] A dendritic modulator enhances tactile feature dimensions related to edge contact offset, tangential slip, and high-frequency perturbations based on contextual aggregation features. The task output module outputs the contact position, tangential slip amount, and contact pattern. The robot control module adjusts the end effector's lateral position based on the contact position offset, keeping the probe sliding along the workpiece edge; when a suspected burr or abnormal contact is detected, the robot reduces its speed or records the abnormal location. This scenario illustrates that the present invention can be used for robot edge tracking, surface inspection, and defect perception tasks.
Claims
1. A magnetic tactile sensing method based on dual-code dendritic modulation, characterized in that, Includes the following steps: S1. Acquire the perception data generated by the robot end effector when performing a contact task. The perception data includes a three-axis magnetic tactile timing array collected by a magnetic tactile array sensor set on the robot end effector, as well as context source data and modal availability markers related to the current contact task. S2. Construct magnetic tactile extended features based on the triaxial magnetic tactile temporal array, wherein the magnetic tactile extended features include original triaxial magnetic field features, magnetic field modulus features, time difference features, high-frequency residual features, and exponential smoothing features; S3. Input the magnetic tactile extended features into the tactile dual-code main path, extract spatial code tactile features through the spatial code branch, extract temporal code tactile features through the temporal code branch, and perform dual-code gating fusion on the spatial code tactile features and temporal code tactile features to obtain the tactile main features; S4. Input the context source data into the corresponding context encoder, generate corresponding context candidate features for candidate context data in an available state, generate preset placeholder features or default features for candidate context data in an unavailable state, and map candidate context data of different categories to the same context feature dimension through the corresponding context encoder or linear mapping layer to obtain multiple context candidate features. S5. Perform context routing scoring and weight normalization on the multiple context candidate features according to the modal availability tag to obtain context weights, and perform weighted aggregation on the multiple context candidate features according to the context weights to obtain context aggregate features; S6. Input the context aggregation feature into the dendritic modulator, generate a gain vector and a bias vector corresponding to the tactile main feature dimension, and use the gain vector and bias vector to perform feature dimension-level modulation on the tactile main feature to obtain context-enhanced tactile features; S7. Input the context-enhanced tactile features into the task output head and output the robot contact state perception results. The robot contact state perception results include one or more of the following: contact position, normal deformation, tangential slip, and contact mode.
2. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S1, the triaxial magnetic tactile temporal array is represented as a tensor of B×T×H×W×C, where B represents the batch size, T represents the time window length, H and W represent the number of rows and columns of the magnetic tactile array, respectively, and C represents the number of magnetic field channels, with C=3.
3. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S2, the magnetic field magnitude feature is obtained by taking the square root of the sum of squares of the three-axis magnetic field components of the same array node in the same time frame; the time difference feature is obtained by taking the difference of the three-axis magnetic field components of the same array node in adjacent time frames; the exponential smoothing feature is obtained by taking the exponential moving average of the magnetic field magnitude feature along the time dimension; and the high-frequency residual feature is obtained by taking the difference between the magnetic field magnitude feature and the exponential smoothing feature.
4. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S3, the spatial code branch takes node features including the original triaxial magnetic field features, magnetic field modulus features and exponential smoothness features as input, and performs spatial topology modeling based on the local adjacency relationship between array nodes in the magnetic tactile array to obtain spatial code tactile features that reflect the contact center position, contact area diffusion and array spatial distribution. The timecode branch takes temporal features, including time difference features, high-frequency residual features, and magnetic field modulus variation features, as input and performs dynamic modeling along the time dimension to obtain timecode tactile features that reflect sliding direction, velocity changes, vibration disturbances, and contact establishment process.
5. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 4, characterized in that, The spatial code branch includes a node embedding module, a local graph attention module, a node attention pooling module, and a temporal attention pooling module connected in sequence; wherein, the local graph attention module only performs information interaction between array nodes that satisfy a preset adjacency relationship; the temporal code branch includes a single-frame projection module, a temporal convolution module, a cyclic dynamic coding module, and a temporal attention pooling module connected in sequence.
6. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S3, the dual-code gating fusion of the spatial code tactile features and the temporal code tactile features includes: concatenating the spatial code tactile features and the temporal code tactile features to obtain a dual-code concatenated feature; inputting the dual-code concatenated feature into a gating network to obtain spatial code weights and temporal code weights; and performing a weighted summation of the spatial code tactile features and the temporal code tactile features according to the spatial code weights and temporal code weights to obtain the tactile master feature.
7. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S4, the visual context includes target location, target edge, target category, or expected contact area information; the language context includes task instructions, action intentions, or contact intensity constraint information; the audio context includes contact sound, friction sound, knocking sound, or collision sound information; and the proprioceptive context includes robot end-effector velocity, joint posture, motion stage, or force control state information.
8. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S5, Performing context routing scoring and weight normalization on the multiple context candidate features based on the modal availability tag includes: calculating the routing score for each context candidate feature respectively; Based on the modal availability flag, the routing score corresponding to the unavailable candidate context data is set to a preset suppression value; the route score after suppression is normalized to obtain the context weight corresponding to each candidate context data. The context aggregate features are obtained by weighted summation of the context candidate features in the available state according to the context weights. When all candidate context data are unavailable, the context aggregation feature is set to a preset zero vector or a preset default vector, and the dendritic modulator outputs an all-1 gain vector and an all-0 bias vector, so that the context-enhanced tactile feature is equal to the tactile main feature, thereby causing the magnetic tactile perception method to fall back to the output of the unmodulated tactile main feature.
9. The magnetic tactile sensing method based on dual-code dendritic modulation according to claim 1, characterized in that, In step S6, the dendritic modulator includes a gain generation branch and a bias generation branch. The gain generation branch generates the gain vector based on the context aggregation feature, and the bias generation branch generates the bias vector based on the context aggregation feature. Modulating the tactile main feature using the gain vector and bias vector at the feature dimension level includes: multiplying each gain element in the gain vector with the feature value of the corresponding feature dimension in the tactile main feature to obtain the gain-modulated tactile feature; and superimposing each bias element in the bias vector onto the feature value of the corresponding feature dimension in the gain-modulated tactile feature to obtain the context-enhanced tactile feature.