Multimodal human-machine interaction chip based on adaptive threshold pulse neural network

By using a multimodal human-computer interaction chip with an adaptive threshold spiking neural network, integrating visual, pressure, and electromyographic signal modalities and dynamically adjusting the membrane potential threshold, the high power consumption and insufficient single-modal perception in SNN hardware implementation are solved, achieving efficient and real-time human-computer interaction.

CN120780159BActive Publication Date: 2025-11-11TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511164223.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-11
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Existing spiking neural networks (SNNs) suffer from challenges in hardware implementation, including high training algorithm difficulty and excessively high pulse emission rate due to the dynamic characteristics of membrane potential. These issues affect system power consumption and real-time performance. Furthermore, their single-mode signal sensing capability is limited, making it difficult to accurately capture the user's multi-dimensional intent.

Method used

An adaptive threshold spiking neural network is adopted, which combines visual, pressure and surface electromyography signal modalities. By dynamically adjusting the membrane potential threshold of spiking neurons, an adaptive threshold SNN algorithm is integrated to optimize hardware design, reduce power consumption and improve computational accuracy.

Benefits of technology

It achieves efficient fusion of multimodal signals, reduces chip power consumption, improves computational accuracy and real-time performance, enhances the accuracy and reliability of human-computer interaction, and is suitable for complex environments and highly dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780159B_ABST
    Figure CN120780159B_ABST
Patent Text Reader

Abstract

This invention provides a multimodal human-machine interaction chip based on an adaptive threshold spiking neural network, comprising: a sensor module for real-time acquisition of input data from three modalities: visual signals, pressure signals, and surface electromyography (sEMG) signals; an acquisition module for analog-to-digital conversion and preprocessing of pressure signals and sEMG signals to generate feature tensors; and a recognition module comprising three adaptive threshold spiking neural networks (CNNs) that process the feature tensors of the three modalities respectively, perform cross-modal feature fusion, and output the robot's motion intention. The adaptive threshold spiking neural network dynamically adjusts the membrane potential threshold of the spiking neurons, reducing power consumption while maintaining computational accuracy. This invention, by combining the input data from the three modalities with the adaptive threshold spiking neural network algorithm for feature fusion, can comprehensively capture the operator's motion intention and the environmental interaction state, significantly improving the accuracy and adaptability of robot motion generation and enhancing the naturalness and reliability of human-machine collaboration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of neural network chips, and more specifically, to a multimodal human-computer interaction chip based on an adaptive threshold spiking neural network. Background Technology

[0002] In recent years, spiking neural networks (SNNs) have demonstrated significant advantages in fields such as robot control, pattern recognition, and real-time reasoning due to their bio-inspired spatiotemporal information processing capabilities and ultra-low power consumption. Compared with traditional artificial neural networks (ANNs), SNNs, through the asynchronous sparse communication mechanism of pulse sequences, can more efficiently simulate the dynamic characteristics of biological nervous systems, making them particularly suitable for resource-constrained embedded hardware scenarios. However, the hardware implementation of SNNs faces two major challenges: first, their training algorithms are not yet mature, making direct training of SNNs difficult and limiting model accuracy and generalization ability; second, in hardware deployment, the dynamic characteristics of neuron membrane potential can easily lead to excessively high emission rates, especially when the input signal strength is high. Excessively high pulse emission frequencies can significantly increase system power consumption, limiting practical applications.

[0003] To overcome training challenges, existing technologies propose strategies to transform well-trained ANNs (such as Convolutional Neural Networks, CNNs) into SNNs. This involves using the mature gradient descent algorithm of ANNs for model training and then implementing network transformation through a pulse coding mechanism. However, such methods still face bottlenecks in hardware deployment. For example, traditional SNNs typically employ a fixed threshold mechanism. When the input signal intensity fluctuates significantly, the membrane potential of neurons may frequently exceed the threshold, leading to uncontrolled pulse emission rate, increased power consumption, and impacted real-time performance. Furthermore, in the field of human-computer interaction, the perception capability of single-modal signals (such as visual or electromyographic signals only) is limited, making it difficult to accurately capture the user's multi-dimensional intentions, resulting in delayed interaction responses or decision-making biases. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this invention is to provide a multimodal human-computer interaction chip based on an adaptive threshold spiking neural network.

[0005] A multimodal human-computer interaction chip based on an adaptive threshold spiking neural network, provided by the present invention, includes:

[0006] Sensor module: Used to acquire input data in real time from three modalities: visual signals, pressure signals, and surface electromyography (sEMG) signals;

[0007] Acquisition module: used to perform analog-to-digital conversion and preprocessing on the pressure signal and surface electromyography (sEMG) signal to generate feature tensors;

[0008] Recognition module: It includes three adaptive threshold SNNs transformed by convolutional neural networks (CNNs), which process the feature tensors of the three modalities respectively, perform cross-modal feature fusion, and output the robot's motion intention;

[0009] The adaptive threshold SNN reduces power consumption while maintaining computational accuracy by dynamically adjusting the membrane potential threshold of spiking neurons.

[0010] Preferably, the sensor module includes:

[0011] A camera: used to acquire visual images of the robot's surrounding environment in real time;

[0012] Five pressure sensors: used to detect the interaction force between the robot and the target object or environment in real time;

[0013] Eight surface electromyography (EMG) sensors: mounted on the operator's upper arm for real-time acquisition of EMG signals.

[0014] Preferably, each adaptive threshold SNN of the recognition module is transformed through the following steps:

[0015] Replace the activation function of the CNN with the ReLU function and remove all bias terms;

[0016] Replace the max pooling layer with an average pooling layer;

[0017] Add a pulse sequence coding layer before the convolutional layer of the CNN to convert the input data into a pulse sequence;

[0018] The trained CNN weights are transferred to the SNN, and the artificial neurons are replaced with spiking neurons.

[0019] Preferably, the mechanism of the adaptive threshold SNN includes:

[0020] The initial threshold is determined based on the maximum activation of adjacent layers, and the initial threshold is adjusted by a scaling factor.

[0021] During the inference process, the threshold is dynamically adjusted based on the weighted sum and maximum value of the postsynaptic potential of the current input data. If the current threshold is exceeded, the threshold is updated to the maximum value to shorten the membrane potential trigger time and optimize the emissivity.

[0022] Preferably, the feature fusion generates a unified motion intent by weighting and integrating the outputs of each modality after pulse counters of the three SNNs. The motion intent includes information such as the robot's trajectory, torque, or pose.

[0023] Preferably, the spiking neuron adopts a leakage integral-triggered LIF model, and its membrane potential dynamic update formula is:

[0024]

[0025] Where t and n represent the time step and the number of layers, respectively; Representation of characteristic tensor and time input The membrane potential generated after the internal state coupling of the spiking neuron in the previous time step; It is used to determine the output pulse tensor The threshold; It is Step function, when hour, ,otherwise ; This indicates the reset potential set after activating the output pulse; Reflects the attenuation factor; spatial characteristics From the original input, use fully connected (FC) or convolutional (Conv) operations. Extracted from;

[0026] When the membrane potential exceeds the threshold, a pulse is triggered and the membrane potential is reset to the preset potential. If no pulse is triggered, the membrane potential is updated according to the attenuation factor.

[0027] Preferably, the acquisition module filters and normalizes the pressure signal; extracts time-frequency domain features from the sEMG signal, including root mean square value and power spectral density; and stores the processed signal as a three-dimensional feature tensor for real-time reading by the SNN.

[0028] Preferably, the chip supports real-time control of the robot's dexterous hand, and the motion intention directly drives the actuator to adjust the gripping force, motion trajectory, or obstacle avoidance path.

[0029] Preferably, the pulse sequence coding layer uses rate coding or time coding to convert continuous data into a pulse sequence, and the frequency of the encoded pulse is proportional to the intensity of the input signal.

[0030] Preferably, the chip is integrated into the robot and performs collaborative computing with the cloud via a low-latency communication interface to achieve distributed decision-making and control.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] 1. This invention integrates input data from three modalities—vision, pressure, and surface electromyography (sEMG)—and combines them with an adaptive threshold SNN algorithm for feature fusion. This enables comprehensive capture of the operator's motion intentions and environmental interaction states, significantly improving the accuracy and adaptability of robot motion generation and enhancing the naturalness and reliability of human-robot collaboration. It innovatively introduces a dynamic threshold adjustment mechanism, adjusting the threshold in real time based on neuronal membrane potential to suppress excessively high pulse emission rates. This effectively reduces chip power consumption while maintaining computational accuracy, solving the problem of energy surges caused by membrane potential overflow in traditional SNNs and extending the device's battery life.

[0033] 2. This invention transforms a mature Convolutional Neural Network (CNN) into a Spiking Neural Network (SNN), and utilizes the mature training framework of CNN (such as ReLU activation and average pooling) to complete model optimization, thus avoiding the algorithmic bottleneck of directly training SNN, shortening the development cycle, and improving the model's transferability and actual deployment efficiency. The chip adopts a modular design (sensor module, acquisition module, and recognition module), combining hardware-level parallel computing and pulse sequence coding technology to achieve real-time processing of multimodal signals from acquisition and preprocessing to feature fusion, ensuring rapid output of motion intentions (millisecond-level response) and meeting the real-time requirements of highly dynamic human-computer interaction scenarios.

[0034] 3. This invention reduces hardware complexity and cost by simplifying sensor configuration (such as a fixed number of pressure and electromyography sensors), optimizing analog-to-digital conversion and preprocessing processes, and improving the system's anti-interference capability. It is suitable for complex industrial environments or medical rehabilitation and other fields with stringent requirements for equipment stability. The chip's modular design and adaptive algorithm framework support flexible expansion to other modal signals (such as voice and temperature) and can be adapted to different types of robots or exoskeleton devices, possessing broad application potential, such as intelligent manufacturing, remote control, and rehabilitation medicine. Attached Figure Description

[0035] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0036] Figure 1 This is a schematic diagram of the multimodal human-computer interaction chip described in this invention;

[0037] Figure 2 This is a schematic diagram of the multi-channel acquisition module of the present invention;

[0038] Figure 3 This is a schematic diagram of the spiking neural network with adaptive threshold described in this invention. Detailed Implementation

[0039] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0040] Example 1:

[0041] Reference Figure 1 and Figure 2 According to the present invention, a multimodal human-computer interaction chip based on an adaptive threshold spiking neural network includes: a sensor module for real-time acquisition of input data from three modalities: visual signals, pressure signals, and surface electromyography (sEMG) signals; an acquisition module for performing analog-to-digital conversion and preprocessing on the pressure signals and sEMG signals to generate feature tensors; and a recognition module comprising three adaptive threshold spiking neural networks (CNNs) that process the feature tensors of the three modalities respectively, perform cross-modal feature fusion, and output the robot's motion intention. The adaptive threshold spiking neural networks (CNNs) reduce power consumption while maintaining computational accuracy by dynamically adjusting the membrane potential threshold of the spiking neurons.

[0042] The sensor module includes: one camera for real-time acquisition of visual images of the robot's surrounding environment; five pressure sensors for real-time detection of the robot's interaction with the target object or environment; and eight surface electromyography (EMG) sensors mounted on the operator's upper arm for real-time acquisition of EMG signals.

[0043] Each adaptive threshold SNN in the recognition module is transformed through the following steps: the activation function of the CNN is replaced with the ReLU function and all bias terms are removed; the average pooling layer is used instead of the max pooling layer; a spiking sequence encoding layer is added before the convolutional layer of the CNN to convert the input data into a spiking sequence; the weights of the trained CNN are transferred to the SNN and the artificial neurons are replaced with spiking neurons.

[0044] The mechanism of adaptive threshold SNN includes: determining the initial threshold based on the maximum activation of adjacent layers and adjusting the initial threshold through a scaling factor; during inference, dynamically adjusting the threshold according to the weighted sum of the postsynaptic potentials of the current input data and the maximum value, and updating it to the maximum value if the current threshold is exceeded, so as to shorten the membrane potential triggering time and optimize the emissivity.

[0045] Feature fusion generates a unified motion intent by weighting and integrating the outputs of each modality after pulse counters of the three SNNs. This motion intent includes the robot's trajectory, torque, or pose information. Specifically, the output pulse sequence of each SNN is first processed... Count the total number of pulses within a fixed time window. (m=1,2,3 correspond to visual, stress, and electromyographic modalities, respectively). A weighted sum of pulse counts from the three modalities is performed to generate a unified intent. :

[0046] ,in To ensure normalization. The weights are determined through offline training or online adaptive learning. The weighted average... The mapping is to the robot's specific motion parameters (such as trajectory and torque), which is achieved through empirical lookup tables or regression models.

[0047] The spiking neuron uses a leakage integral-triggered LIF model, and its membrane potential dynamic update formula is as follows:

[0048]

[0049] Where t and n represent the time step and the number of layers, respectively; Representation of characteristic tensor and time input The membrane potential generated after the internal state coupling of the spiking neuron in the previous time step; It is used to determine the output pulse tensor The threshold; It is Step function, when hour, ,otherwise ; This indicates the reset potential set after activating the output pulse; Reflects the attenuation factor; spatial characteristics From the original input, use fully connected (FC) or convolutional (Conv) operations. Extracted from the membrane potential; when the membrane potential exceeds the threshold, a pulse is triggered and the membrane potential is reset to the preset potential; if no pulse is triggered, the membrane potential is updated according to the attenuation factor.

[0050] The acquisition module filters and normalizes the pressure signal; it extracts time-frequency features from the sEMG signal, including the root mean square value and power spectral density; and it stores the processed signal as a three-dimensional feature tensor for real-time reading by the SNN. The chip supports real-time control of the robot's dexterous hand, and the motion intention directly drives the actuator to adjust the gripping force, motion trajectory, or obstacle avoidance path.

[0051] The pulse sequence coding layer uses rate coding or time coding to convert continuous data into a pulse sequence, with the frequency of the encoded pulses being proportional to the intensity of the input signal. The chip is integrated into the robot and collaborates with the cloud via a low-latency communication interface to achieve distributed decision-making and control.

[0052] Example 2:

[0053] This invention provides a multimodal human-computer interaction chip based on an adaptive threshold SNN algorithm. This chip can process inputs from three modalities: vision, pressure, and electromyography (EMG). It rapidly processes and analyzes the user's movement intentions and the robot's (hand's) physical state information using the SNN algorithm, fusing user intentions and robot state to generate robot motion, thus improving the quality and efficiency of human-computer interaction. Vision is derived from a camera on the robot, EMG data from the operator's surface EMG information, and pressure data from the interaction force between the robot and the target object (or environment, or person). During chip operation, it acquires current images through the camera, pressure data through a pressure sensor, and real-time EMG signals through an EMG signal sensor. The data from these three modalities are then transmitted to the SNN for recognition, feature fusion, and output of the movement intention, thus realizing the multimodal human-computer interaction chip. The chip includes a sensor module, an acquisition module, and a recognition module. The sensor module includes a camera for real-time image acquisition, five pressure sensors for real-time pressure data acquisition, and eight electromyography (EMG) sensors mounted on the upper arm for real-time EMG signal acquisition. The acquisition module performs analog-to-digital conversion and preprocessing on the acquired pressure and EMG signals to convert them into feature tensors that serve as input to a multimodal neural network (SNN). The recognition module mainly includes three SNNs with adaptive thresholds, converted from a CNN, used to process the feature tensors and perform feature fusion across the three modalities, outputting the intended movement. This invention can acquire visual images, pressure data, and EMG signals in real time, perform real-time calculations based on the acquired multimodal signals, and output the intended movement to support actuator decision-making and control, demonstrating high practical application value.

[0054] A multimodal human-computer interaction chip based on an adaptive threshold SNN algorithm includes a sensor module, an acquisition module, and a recognition module. The chip acquires visual, interactive force, and sEMG data in real time through sensors, analyzes the user's movement intention and the robot's (hand) physical state information, and fuses the user's intention and robot state to generate the robot's motion, thereby improving the quality and efficiency of human-computer interaction.

[0055] The sensor module includes a camera for real-time image acquisition, five pressure sensors for real-time pressure data acquisition, and eight electromyography (EMG) sensors mounted on the upper arm for real-time acquisition of surface EMG signals. The acquisition module performs analog-to-digital conversion and preprocessing on the acquired pressure and surface EMG signals, storing them in internal memory as feature tensors for the SNN input.

[0056] The recognition module mainly includes three SNNs with adaptive thresholds transformed from CNNs, used to process feature tensors and perform feature fusion of three modalities. The original CNN uses average pooling and ReLU function, and is converted into an SNN after pulse conversion. Feature fusion is performed after each SNN pulse counter to obtain a multimodal SNN.

[0057] The adaptive thresholding mechanism is designed by comparing the membrane potential of each neuron in each layer with a threshold. If the membrane potential exceeds the threshold by a certain amount, the threshold is appropriately increased and the emissivity is reduced, so as to reduce power consumption without sacrificing accuracy as much as possible.

[0058] This invention relates to the field of neural network chips, specifically a multimodal human-computer interaction chip based on an adaptive threshold SNN algorithm. The invention provides a multimodal human-computer interaction chip based on an SNN algorithm. This chip can process inputs from three modalities: vision, pressure, and electromyography (EMG). It rapidly processes and analyzes the user's movement intention and the robot's (hand's) physical state information using the SNN algorithm, fusing user intention and robot state to generate the robot's motion (trajectory, torque, pose, etc.), thus improving the quality and efficiency of human-computer interaction. Vision is derived from a camera on the robot, EMG data from the operator's surface EMG information, and pressure data from the interaction force between the robot and the target object (or environment, or person). During chip operation, the chip acquires the current image through the camera, acquires pressure data through a pressure sensor, and collects EMG signals in real time through an EMG signal sensor. Then, the data from these three modalities are transmitted to the SNN for recognition, feature fusion, and output of the movement intention, thus realizing the multimodal human-computer interaction chip. The chip includes a sensor module, an acquisition module, and a recognition module. The sensor module includes a camera for real-time image acquisition, five pressure sensors for real-time pressure data acquisition, and eight electromyography (EMG) sensors mounted on the upper arm for real-time EMG signal acquisition. The acquisition module performs analog-to-digital conversion and preprocessing on the acquired pressure and surface EMG signals to convert them into feature tensors that serve as input to a subnet mask (SNN). The recognition module mainly includes three SNNs converted from a CNN, used to process the feature tensors and perform feature fusion across the three modalities, outputting the intended movement. This invention can acquire visual images, pressure data, and EMG signals in real time, and can perform real-time calculations based on the acquired multimodal signals to output the intended movement, supporting dexterous hand decision-making and control, thus possessing high practical application value.

[0059] This invention provides a multimodal human-computer interaction chip based on the SNN algorithm. This chip can process inputs from three modalities: visual, pressure, and electromyographic signals. It rapidly processes and analyzes the user's movement intentions and the robot's (hand's) physical state information using the SNN algorithm, fusing user intentions and robot state to generate robot motion, thus improving the quality and efficiency of human-computer interaction. Since existing spiking neurons suffer from increased power consumption due to excessively high emissivity when the membrane potential is too high, this invention designs an adaptive threshold mechanism deployed on each neuron. By adjusting the threshold potential in real time, power consumption is reduced while minimizing impact on accuracy. Because the training algorithm for spiking neural networks is still relatively immature, this invention transforms traditional artificial neural networks into spiking neural networks and utilizes a more mature artificial neural network training algorithm to train a deep neural network based on artificial neural networks.

[0060] like Figure 1 A multimodal human-computer interaction chip based on an adaptive threshold SNN algorithm is proposed, comprising a sensor module, an acquisition module, and a recognition module. The chip acquires visual, interactive force, and sEMG data in real time through sensors, analyzes the user's movement intention and the robot's (hand) physical state information, and generates the robot's motion by fusing the user's intention and the robot's state, thereby improving the quality and efficiency of human-computer interaction.

[0061] like Figure 2 , 3 A multimodal human-computer interaction chip based on an adaptive threshold SNN algorithm is proposed. The acquisition module is used to perform analog-to-digital conversion and preprocessing on the acquired pressure and electromyography signals to convert them into feature tensors as inputs to the SNN. The recognition module mainly includes three SNNs with adaptive thresholds converted by CNN, which are used to process the feature tensors and perform feature fusion of the three modalities to output the motion intention.

[0062] The design steps for a neural network architecture are as follows:

[0063] To ensure that the output value of each neuron in the CNN is positive, an absolute value function (abs()) is added after the image preprocessing layer and before the convolutional layer to guarantee that the input values ​​of the convolutional layers are non-negative. The activation function of the neurons is replaced with the ReLU function. This speeds up the convergence of the original CNN, and the ReLU function is similar in properties to LIF neurons, minimizing the accuracy loss after network transformation.

[0064] Set all bias terms of the CNN to 0.

[0065] An average pooling layer is used instead of a max-pooling layer.

[0066] The network structure of SNN is the same as that of a pruned CNN, simply replacing the artificial neurons in the CNN with spiking neurons.

[0067] A network layer for pulse sequence generation is added before the convolutional layers of the CNN to encode the input image data and convert it into a pulse sequence.

[0068] The decision-making method is changed. Within a certain period of time, the number of pulses output by the fully connected layer is counted, and the category with the most pulses is taken as the final classification result.

[0069] All weights obtained from the pruned CNN training are transferred to the corresponding SNN.

[0070] The design steps for a spiking neuron are as follows:

[0071] This invention employs the LIF model as the spiking neuron model. The LIF model is one of the most commonly used spiking neuron models because it strikes a balance between the complex spatiotemporal dynamics of biological neurons and a simplified mathematical form. LIF neurons are suitable for large-scale SNN simulations and can be described by differential functions:

[0072]

[0073] in, It is a time constant. and This represents the membrane potential of the postsynaptic neuron and the input collected from the presynaptic neuron. Solving the formula yields a simple iterative representation of the LIF neuron. For ease of describing dynamic SNNs, the expression for a single-layer LIF-SN is given here:

[0074]

[0075] Where t and n represent the time step and the number of layers, respectively; Representation of characteristic tensor and time input The membrane potential generated after coupling (the internal state of the spiking neuron in the previous time step); It is used to determine the output pulse tensor The threshold; It is Step function, when hour, ,otherwise ; This indicates the reset potential set after activating the output pulse; This reflects the attenuation factor. Spatial characteristics. The original input can be processed through fully connected (FC) or convolution (Conv) operations. Extracted from. When using the Conv operation:

[0076]

[0077] in, , and These represent average pooling, batch normalization, and convolution operations, respectively. It is a weight matrix; It is a pulse tensor that contains only 0s and 1s; ( It is the number of channels. and (This refers to the channel size).

[0078] LIF layers will convert feature tensors and time input Integration into membrane potential Then, triggering and leakage mechanisms are used to generate pulse tensors for the next layer and new neuron states for the next time step. Specifically, when The number of entries is greater than the threshold At that time, pulse sequence The spatial output will be activated. The entries in will be reset to Time output Will be Decision, because It must be 0. After the Conv operation, all tensors have the same dimension.

[0079] The hardware implementation of the LIF model used in this invention employs digital circuit or mixed-signal circuit design. The specific steps are as follows: First, for membrane potential integration and leakage, a capacitor-resistor (RC) circuit is used to simulate the membrane potential integration and leakage process, where the capacitor stores charge (corresponding to the membrane potential). ), resistance controls leakage rate (by attenuation factor) (Decision). Secondly, in digital circuits, integration is implemented using an accumulator, and a multiplier is introduced. (as a coefficient) simulates leakage. Each time step... Then the membrane potential is updated as follows:

[0080] = + The membrane potential is monitored in real time using a comparator circuit. With threshold .when At that time, trigger pulse =1, and immediately reset the membrane potential to (Typically set to 0 or a fixed value). When not triggered, the membrane potential is updated according to the leakage rule. Conditional reset is achieved via a multiplexer (MUX): if a trigger pulse is detected, the output... Otherwise, output Furthermore, by integrating a dynamic threshold adjustment module (such as a programmable threshold register), an adaptive threshold mechanism is supported. The design steps of the adaptive threshold algorithm are as follows:

[0081] The first stage involves determining an initial threshold based on the maximum activation of adjacent layers. First, the maximum activation of each layer in the network is estimated, i.e., the activation greater than 99.9% of all activations. Then, a threshold is determined based on the maximum activation of adjacent layers. Since the threshold is obtained from a large subset of training samples, it is generally quite large. For the vast majority of input samples, even the maximum activation of units within a layer will be much lower than the maximum activation of adjacent layers. This results in insufficient emission within a layer to drive higher layers, leading to worse classification results. Therefore, [the following is missing from the original text: "will"] proportional factor Decrease it to obtain the initial threshold, that is:

[0082]

[0083] In the second stage, the threshold is dynamically adjusted based on the specific input data. At each time step, the maximum value of the weighted sum of postsynaptic potentials in each layer of the SNN is calculated. If the value is greater than the current threshold Then use replace This method adjusts the threshold only for a single test sample. During the dynamic adjustment in the inference phase, the maximum value of the weighted sum of postsynaptic potentials in the current layer is calculated. : in For synaptic weights, The spatiotemporal characteristics of the input pulse sequence. Comparison. Compared with the current threshold :like Then update the threshold: Otherwise, keep the threshold unchanged. Updated threshold The pulse trigger is used to determine the next time step. In addition, upper and lower limits of thresholds can be set to prevent extreme values ​​from causing system instability; and a smoothing factor (such as a moving average) can be introduced to avoid threshold abrupt changes.

[0084] It can dynamically set the optimal threshold corresponding to the current sample for each layer of neurons during the inference process, shortening the time for the membrane potential to cross the threshold while distinguishing input differences, thus making the emission rate more conducive to driving higher layers.

[0085] Through the aforementioned steps, the chip, during operation, uses a camera to acquire images in real time, five pressure sensors to acquire pressure data in real time, and eight electromyography (EMG) sensors mounted on the upper arm to acquire EMG signals in real time. The collected multimodal information, after preprocessing, is used in real-time by a subtraction neural network (SNN) with adaptive thresholds to accurately output the current movement intention, supporting the actuator in decision-making and control.

[0086] Those skilled in the art can understand this embodiment as a more specific description of Embodiment 1.

[0087] Those skilled in the art will understand that, besides implementing the system and its various devices, modules, and units provided by this invention in the form of purely computer-readable program code, the same functions can be achieved entirely through logical programming of the method steps, making the system and its various devices, modules, and units of this invention function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system and its various devices, modules, and units provided by this invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; alternatively, the devices, modules, and units for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0088] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A multimodal human-computer interaction chip based on an adaptive threshold spiking neural network, characterized in that, include: Sensor module: Used to acquire input data in real time from three modalities: visual signals, pressure signals, and surface electromyography (sEMG) signals; Acquisition module: used to perform analog-to-digital conversion and preprocessing on the pressure signal and surface electromyography (sEMG) signal to generate feature tensors; Recognition module: It includes three adaptive threshold SNNs transformed by convolutional neural networks (CNNs), which process the feature tensors of the three modalities respectively, perform cross-modal feature fusion, and output the robot's motion intention; The adaptive threshold SNN reduces power consumption while maintaining computational accuracy by dynamically adjusting the membrane potential threshold of spiking neurons. Each adaptive threshold SNN of the recognition module is transformed through the following steps: Replace the activation function of the CNN with the ReLU function and remove all bias terms; Replace the max pooling layer with an average pooling layer; Add a pulse sequence coding layer before the convolutional layer of the CNN to convert the input data into a pulse sequence; The trained CNN weights are transferred to the SNN, and the artificial neurons are replaced with spiking neurons; The mechanism of the adaptive threshold SNN includes: The initial threshold is determined based on the maximum activation of adjacent layers, and the initial threshold is adjusted by a scaling factor. During the inference process, the threshold is dynamically adjusted based on the weighted sum and maximum value of the postsynaptic potential of the current input data. If the current threshold is exceeded, the threshold is updated to the maximum value to shorten the membrane potential trigger time and optimize the emissivity. The spiking neuron uses a leakage integral-triggered LIF model, and its membrane potential dynamic update formula is as follows: Where t and n represent the time step and the number of layers, respectively; Representation of characteristic tensor and time input The membrane potential generated after the internal state coupling of the spiking neuron in the previous time step; It is used to determine the output pulse tensor The threshold; It is Step function, when hour, ,otherwise ; This indicates the reset potential set after activating the output pulse; Reflects the attenuation factor; spatial characteristics From the original input, use fully connected (FC) or convolutional (Conv) operations. Extracted from; When the membrane potential exceeds the threshold, a pulse is triggered and the membrane potential is reset to the preset potential. If no pulse is triggered, the membrane potential is updated according to the attenuation factor.

2. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The sensor module includes: A camera: used to acquire visual images of the robot's surrounding environment in real time; Five pressure sensors: used to detect the interaction force between the robot and the target object or environment in real time; Eight surface electromyography (EMG) sensors: mounted on the operator's upper arm for real-time acquisition of EMG signals.

3. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The feature fusion generates a unified motion intent by weighting and integrating the outputs of each modality after pulse counters of three SNNs. The motion intent includes the robot's trajectory, torque, or pose information.

4. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The acquisition module filters and normalizes the pressure signal; extracts time-frequency domain features from the sEMG signal, including root mean square value and power spectral density; and stores the processed signal as a three-dimensional feature tensor for real-time reading by the SNN.

5. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The chip supports real-time control of the robot's dexterous hand, and the motion intention directly drives the actuator to adjust the gripping force, motion trajectory, or obstacle avoidance path.

6. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The pulse sequence coding layer uses rate coding or time coding to convert continuous data into a pulse sequence. The frequency of the encoded pulse is proportional to the intensity of the input signal.

7. The multimodal human-computer interaction chip based on an adaptive threshold spiking neural network according to claim 1, characterized in that, The chip is integrated into the robot and collaborates with the cloud through a low-latency communication interface to achieve distributed decision-making and control.

Citation Information

Patent Citations

  • Robot dynamic obstacle avoidance method based on multi-modal pulse neural network

    CN116382267A

  • Method for converting artificial neural network into spiking neural network

    CN119862922A