Hardware acceleration platform and method for electroencephalogram epilepsy data classification
By designing a hardware acceleration platform for EEG epilepsy data classification, using the improved Spikformer model and accelerator, the problems of high computing complexity and high power consumption in the existing technology are solved, and efficient and low-power consumption of EEG epilepsy data classification is achieved.
Patent Information
- Application Number
- CN202510150694.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing EEG epilepsy data classification technology based on deep learning is high in complexity in long sequences, resulting in large calculation volume and high power consumption, which limits its application in resource-constrained environments.
A hardware acceleration platform for EEG epilepsy data classification is designed. Through the improved Spikformer model and accelerator, pulse depth separation convolution and pulse matrix dot product operations are used to reduce the computational complexity and improve efficiency.
By converting the EEG signal into Gram angle field images and classifying using improved Spikformer models and accelerators, computing power consumption is significantly reduced and application efficiency in resource-constrained environments are improved.
Smart Images

Figure CN119919775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electroencephalogram (EEG) technology, and in particular to a hardware acceleration platform for EEG epilepsy data classification. Background Art
[0002] Epilepsy is a common neurological disease, and electroencephalogram (EEG) can record changes in brain electrical signals. It is an important tool for epilepsy data classification and thus for diagnosing and monitoring epilepsy. In order to improve detection efficiency, deep learning is currently used to classify EEG epilepsy data.
[0003] Existing deep learning-based EEG epilepsy data classification usually uses the Transformer model. Transformer is a neural network architecture based on the self-attention mechanism, which allows the model to automatically focus on different parts of the sequence when processing sequence data and assign different weights to each part according to the importance of each part. It is suitable for processing time series data such as EEG. Although the Transformer architecture is powerful, the self-attention mechanism has a high complexity in the case of long sequences. There is a problem that the implemented self-attention mechanism is not easy to deploy on hardware because it contains calculations such as multiplication, square root, and division. The required design is too complex, the amount of calculation is large, and it occupies more resources, resulting in high power consumption when classifying EEG epilepsy data, which limits its application in resource-constrained environments (such as edge devices). Summary of the invention
[0004] The present invention provides a hardware acceleration platform and method for EEG epilepsy data classification, which solves the technical problem of high power consumption when performing EEG epilepsy data classification in the existing EEG epilepsy data classification technology.
[0005] A hardware acceleration platform for EEG epilepsy data classification provided by the first aspect of the present invention comprises: a PC terminal, an accelerator and an off-chip memory connected in sequence;
[0006] The PC is equipped with an image processor, which is used to obtain electroencephalogram signals and convert the electroencephalogram signals into Gram's angle field images;
[0007] The off-chip memory is used to store target weight parameters of the improved Spikformer model;
[0008] The accelerator is used to load the target weight parameters, and classify the electroencephalogram epilepsy data based on the Gram angle field image according to the improved Spikformer model, and output the classification result.
[0009] Furthermore, the processing of the improved Spikformer model includes:
[0010] Performing a pulse depth separation convolution operation on the Gram angular field image to generate a first pulse convolution feature;
[0011] Perform multiple pulse depth separation convolution operations on the first pulse convolution feature, and map them into a pulse query matrix, a pulse key matrix and a pulse value matrix respectively;
[0012] The pulse query matrix and the pulse key matrix are phase-ORed and then added column by column, and the pulses are nonlinearly mapped and then phase-ORed with the pulse value matrix to generate pulse attention features through pulse matrix dot product operation;
[0013] Performing a pulse full connection calculation on the pulse attention feature to determine a first pulse transformation feature;
[0014] Performing a residual connection between the first pulse convolution feature and the first pulse transformation feature, and outputting a bitwise addition feature;
[0015] Performing pulse full connection calculation on the bitwise addition feature to construct a second pulse transformation feature;
[0016] After the bitwise addition feature is connected with the second pulse transformation feature residual, a pulse full connection calculation is performed to output a classification result;
[0017] The pulse depth separation convolution operation includes sequentially performing a pulse depth convolution operation and a pulse point convolution operation.
[0018] Further, the accelerator includes a pulse detection model controller, a pulse matrix dot product component, a pulse convolution component and a pulse fully connected component, a pulse buffer and a model parameter buffer;
[0019] The pulse detection model controller is respectively connected to the pulse matrix dot product component, the pulse convolution component, the pulse fully connected component, the pulse buffer and the PC end;
[0020] The pulse buffer is respectively connected to the pulse matrix dot product component, the pulse convolution component, and the pulse fully connected component;
[0021] The model parameter buffer is connected to the pulse convolution component, the pulse fully connected component and the off-chip memory respectively;
[0022] The pulse detection model controller is used to control the pulse matrix dot product component, the pulse convolution component and the pulse fully connected component to calculate and output a classification result based on the Gram angular field image according to the improved Spikformer model;
[0023] The model parameter buffer includes a weight buffer and a neuron parameter buffer, the weight buffer is used to cache target weight parameters, and the neuron parameter buffer is used to cache neuron state parameters;
[0024] The pulse convolution component is used to load the target weight parameters and the neuron state parameters to perform pulse depth separation convolution operation to output the pulse convolution feature;
[0025] The pulse matrix dot product component is used to perform a pulse matrix dot product operation to output a pulse attention feature;
[0026] The pulse fully connected component is used to load the target weight parameter and the neuron state parameter to perform pulse fully connected calculation to output the pulse transformation feature;
[0027] The pulse buffer is used to cache pulse convolution features, pulse attention features, pulse transformation features and classification results, and perform bit-wise addition of features.
[0028] Further, the pulse matrix dot product component includes a pulse matrix dot product controller, a matrix dot product calculation core and a pulse matrix dot product buffer connected in sequence;
[0029] The pulse matrix dot product controller is connected to the pulse detection model controller;
[0030] The matrix dot product calculation core includes a pulse column summation unit and a phase and calculation unit;
[0031] The pulse column summing unit is used to perform column-by-column addition to output pulse summing characteristics;
[0032] The phase-by-phase AND calculation unit is used to perform phase-by-phase AND characteristics of output pulses;
[0033] The pulse matrix dot product buffer is used to buffer pulse summation characteristics and pulse phase and characteristics, and perform pulse nonlinear mapping.
[0034] Further, the pulse convolution component includes a pulse convolution controller and a processing unit array;
[0035] The pulse convolution controller is connected to the pulse detection model controller and the processing unit array respectively;
[0036] The processing unit array is connected to the pulse buffer and the model parameter buffer respectively;
[0037] The processing unit array includes a plurality of processing units in an array, and the processing units are used to perform event-driven pulse point convolution operations or pulse depth convolution operations.
[0038] Further, the impulse fully connected component includes an impulse fully connected controller and a linear unit;
[0039] The pulse fully connected controller is connected to the pulse detection model controller and the linear unit respectively;
[0040] The linear unit is connected to the pulse buffer and the model parameter buffer respectively;
[0041] The linear unit is used to perform pulse full connection calculation.
[0042] Furthermore, the process of obtaining the target weight parameter includes:
[0043] Using the training image set to train the improved Spikformer model to be trained, and determining the trained improved Spikformer model;
[0044] Perform weight quantization training on the trained improved Spikformer model to determine the target weight parameters.
[0045] A second aspect of the present invention provides a hardware acceleration method for EEG epilepsy data classification, comprising:
[0046] Acquiring an electroencephalogram signal, and converting the electroencephalogram signal into a Gram's angle field image;
[0047] The electroencephalogram epilepsy data is classified based on the Gram angle field image according to the improved Spikformer model, and the classification result is output.
[0048] A third aspect of the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the hardware acceleration method for classifying electroencephalogram epilepsy data as described in any one of the above items.
[0049] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the hardware acceleration method for classifying electroencephalogram epilepsy data as described in any one of the above items.
[0050] A fifth aspect of the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the hardware acceleration method for classifying electroencephalogram epilepsy data as described in any one of the above items.
[0051] It can be seen from the above technical solutions that the present invention has the following advantages:
[0052] The first aspect of the present invention provides a hardware acceleration platform for EEG epilepsy data classification, including: a PC terminal, an accelerator and an off-chip memory connected in sequence; an image processor is deployed on the PC terminal, and the image processor is used to obtain EEG signals and convert the EEG signals into Gram's angular field images; the off-chip memory is used to store the target weight parameters of the improved Spikformer model; the accelerator is used to load the target weight parameters, and classify the EEG epilepsy data based on the Gram's angular field image according to the improved Spikformer model, and output the classification result. Based on the above scheme, after the EEG signal is converted into the Gram's angular field image, the EEG epilepsy data is classified according to the improved Spikformer model through the accelerator, which can reduce the computational complexity and thus reduce the computational power consumption.
[0053] The second aspect of the present invention provides a hardware acceleration method for EEG epilepsy data classification, comprising: obtaining an EEG signal and converting the EEG signal into a Gram angular field image; performing EEG epilepsy data classification based on the Gram angular field image according to an improved Spikformer model, and outputting a classification result. Based on the above scheme, after converting the EEG signal into a Gram angular field image, the EEG epilepsy data is classified by using an improved Spikformer model, which can reduce the computational complexity and thus reduce the computational power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0055] Figure 1 A schematic diagram of the structure of a hardware acceleration platform for EEG epilepsy data classification provided by an embodiment of the present invention;
[0056] Figure 2 A schematic diagram of the structure of training and reasoning of the improved Spikformer model provided in an embodiment of the present invention;
[0057] Figure 3 A schematic diagram of the structure of an improved Spikformer model provided in an embodiment of the present invention;
[0058] Figure 4 A schematic diagram of data preprocessing of an electroencephalogram signal provided in an embodiment of the present invention;
[0059] Figure 5 A schematic diagram of the structure of a pulse self-attention module provided in an embodiment of the present invention;
[0060] Figure 6 A schematic diagram of feature processing mapping of the improved Spikformer model in an accelerator provided in an embodiment of the present invention;
[0061] Figure 7 A flowchart of the steps of a hardware acceleration method for classifying electroencephalogram epilepsy data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The embodiment of the present invention provides a hardware acceleration platform and method for EEG epilepsy data classification, which are used to solve the technical problem of high power consumption when performing EEG epilepsy data classification in the existing EEG epilepsy data classification technology.
[0063] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] Terminology explanation:
[0065] SNN: Spiking Neural Network.
[0066] Transformer: A deep learning network architecture.
[0067] LIF: leaky firing neuron.
[0068] BPTT: Back Propagation Through Time.
[0069] EEG: electroencephalogram.
[0070] Spikformerr: Spiking Transformer network architecture.
[0071] MLP (Multilayer Perceptron): A feed-forward artificial neural network.
[0072] See also Figure 1 , Figure 1 A schematic diagram of the structure of a hardware acceleration platform for EEG epilepsy data classification provided by an embodiment of the present invention.
[0073] The present invention provides a hardware acceleration platform for electroencephalogram epilepsy data classification, comprising: a PC terminal, an accelerator and an off-chip memory connected in sequence;
[0074] An image processor is deployed on the PC end, and the image processor is used to obtain the EEG signal and convert the EEG signal into a Gram's angle field image;
[0075] Off-chip memory for storing target weight parameters of the improved Spikformer model;
[0076] The accelerator is used to load the target weight parameters and classify the EEG epilepsy data based on the Gram angular field image according to the improved Spikformer model, and output the classification results.
[0077] It should be noted that the hardware acceleration platform for EEG epilepsy data classification of this embodiment performs data interaction between the PC, accelerator and off-chip memory connected in sequence: first, the acquired EEG signal is normalized and polar coordinate processed by the image processor deployed on the PC, and then converted into the corresponding Gram angular field image through Gram angular field (GASF) imaging processing. Then, the target weight parameter refers to the weight parameter used in actual application after the model is trained. When the accelerator receives the Gram angular field image sent by the PC, the accelerator loads the corresponding target weight parameter from the off-chip memory according to the hierarchical logic of the improved Spikformer model to complete the EEG epilepsy data classification, thereby outputting the classification result. Storing the target weight parameter in the off-chip memory can reduce the occupation of accelerator resources, and the hardware-friendly design of the improved Spikformer model makes it easier to deploy to FPGA hardware.
[0078] It can be understood that the improved Spikformer model in this embodiment is improved on the basis of Spikformer. Thanks to the computing characteristics of SNN, SNN has obvious advantages in power consumption compared with traditional neural networks of the same structure. Similarly, Spikformer, which combines SNN with transformer, inherits the characteristics of low power consumption, high energy efficiency, and the ability to process spatiotemporal information of pulse neural networks, as well as the powerful expressive ability of transformers. In the improved Spikformer model, in one implementation, simple operations and LIF neurons can be used to replace complex calculations such as softmax to solve the problem that the self-attention mechanism of the traditional Transformer architecture is not easy to deploy on hardware because it contains calculations such as multiplication, square root, and division, thereby achieving hardware friendliness.
[0079] In a specific implementation manner, the processing of the improved Spikformer model includes:
[0080] Performing a pulse depth separation convolution operation on the Gram angular field image to generate a first pulse convolution feature;
[0081] Perform multiple pulse depth separation convolution operations on the first pulse convolution feature, and map them into a pulse query matrix, a pulse key matrix, and a pulse value matrix respectively;
[0082] The pulse query matrix and the pulse key matrix are phase-ORed and then added column by column, and the pulses are nonlinearly mapped and then phase-ORed with the pulse value matrix to generate the pulse attention feature through the pulse matrix dot product operation;
[0083] Performing a pulse full connection calculation on the pulse attention feature to determine a first pulse transformation feature;
[0084] Perform a residual connection between the first pulse convolution feature and the first pulse transformation feature, and output a bitwise addition feature;
[0085] Perform pulse full connection calculation on the bitwise addition feature to construct the second pulse transformation feature;
[0086] After connecting the bitwise addition feature with the second pulse transformation feature residual, a pulse full connection calculation is performed to output the classification result;
[0087] The pulse depth separation convolution operation includes sequentially performing a pulse depth convolution operation and a pulse point convolution operation.
[0088] It should be noted that if Figure 2-5 As shown, the improved Spikformer model includes a pulse coding module, an encoder and a classification head connected in sequence; the pulse coding module includes a pulse depth separation convolution layer; the encoder includes a cascaded pulse self-attention (SSA) module with residual connection and an MLP module with residual connection, the MLP module includes at least one pulse fully connected layer, and the pulse fully connected layer includes a Linear unit, a BN unit and a LIF neuron; the pulse self-attention module includes multiple pulse depth separation convolution layers, LIF neurons and a pulse fully connected layer; the classification head includes a pulse fully connected layer;
[0089] In the specific implementation, first, the Gram angular field image of the input pulse coding module is subjected to a pulse depth separation convolution operation through a pulse depth separation convolution layer to generate a first pulse convolution feature. Then, in the encoder, the first pulse convolution feature is respectively subjected to pulse depth separation convolution operations based on multiple pulse depth separation convolution layers to be mapped into a pulse query matrix Q, a pulse key matrix K and a pulse value matrix V. The pulse query matrix and the pulse key matrix are phase-wise ANDed and then added column by column. The pulses are nonlinearly mapped according to the LIF neurons and then phase-wise ANDed with the pulse value matrix to generate a pulse attention feature through a pulse matrix dot product operation in the pulse self-attention module. The self-attention mechanism contained in this part allows the model to automatically pay attention to nonlinear data in the sequence when processing sequence data. The same part, and different weights are given according to the importance of each part, so that the model can better capture long-distance dependencies, which is helpful to improve the accuracy of the model in the EEG epilepsy data classification task. The pulse attention feature is calculated by the fully connected layer to determine the first pulse transformation feature, and the first pulse convolution feature is residually connected with the first pulse transformation feature to output the bitwise addition feature. The fully connected layer of the MLP module is used to perform pulse full connection calculation on the bitwise addition feature for further high-level feature extraction, and the second pulse transformation feature is constructed. After the bitwise addition feature is residually connected with the second pulse transformation feature, the pulse full connection calculation is performed through the fully connected layer of the classification head to realize the EEG epilepsy data classification, and the classification result is output;
[0090] Among them, the pulse deep separation convolution layer includes deep convolution, BN unit, LIF neuron, point convolution, BN unit and LIF neuron, and the processing process of the pulse deep separation convolution unit includes: performing pulse deep convolution operation on the input pulse convolution input feature through cascaded deep convolution, BN unit and LIF neuron to generate deep features, using cascaded point convolution, BN unit and LIF neuron to perform pulse point convolution operation on the deep features, and outputting pulse convolution output features; therefore, more specifically, the pulse deep separation convolution operation performed by the pulse deep separation convolution layer includes sequentially performing pulse deep convolution operation and pulse point convolution operation, the pulse deep convolution operation includes sequentially performing deep convolution operation, batch normalization and pulse nonlinear mapping, and the pulse point convolution operation includes sequentially performing point convolution operation, batch normalization and pulse nonlinear mapping;
[0091] It can be understood that the pulse coding module, encoder and classification head ensure that the data running therein are all pulse data by setting the pulse nonlinear mapping of the LIF neurons, thereby reducing the power consumption of the network operation; at the same time, since the input processing data has been converted from the EEG signal to the Gram angular field image, the pulse depth separation convolution operation that is more suitable for image feature extraction can be used to extract features of the image, thereby extracting the associated features and spatial hierarchical structure of the local information in the image. Accordingly, the pulse self-attention module of this embodiment no longer uses the linear layer to extract the corresponding QKV data, but uses the convolution operation that is more suitable for image data feature extraction. In order to reduce the number of parameters of the model, the self-attention operation composed of pulse depth separation convolution is adopted. Since the features obtained after the conversion of the LIF neurons are all pulse data, In this case, Q and K can be directly calculated by using the phase AND operation plus the addition operation. The pulse data obtained by the LIF neuron are only 0 and 1, so there is no need to limit the data range through the Scale and Softmax operations. The pulse data is naturally limited to 0 and 1. Then the obtained result is phase ANDed with V to obtain the pulse attention feature after the result weight distribution processing. The pulse self-attention module of this embodiment uses the pulse deep separation convolution to replace the QKV extraction operation composed of the pulse full connection layer, and replaces the traditional matrix multiplication, Scale operation and Softmax operation. Since the elements of the pulse data are only 1 or 0, the pulse data is very sparse, and the corresponding self-attention calculation consumption is also greatly reduced, which helps to reduce the number of network parameters and speed up the calculation.
[0092] It is understandable that many different neuron models can be used as basic units for building SNN networks, such as the Hodgkin-Huxley (HH) model, the Izhikevich (IZH) model, the Integred-and-Fire (IF) model, and the Leaky Integred-and-Fire (LIF) model. These neuron models can truly simulate the operation mode of biological nerves from the perspective of neuroscience; among them, the HH model has high accuracy but is too complex. Although the IZH model is a simplified version of the HH model, its biological accuracy is close to that of the HH model, but the computational complexity is still very high. The IF model is the simplest, with the smallest amount of computation and high hardware implementation efficiency. However, due to the oversimplification of the IF model, it may not be able to accurately simulate the complex behaviors of some biological neurons. Although the accuracy of the LIF model is slightly lower than that of the IZH model, and its computational amount is slightly larger than that of the IF model, the LIF model is simpler and more intuitive than the IZH. At the same time, its performance and biological interpretability are better than the IF model. It can also solve the problem of the lack of time dependence of the IF model and can provide the performance required by the SNN network, so it is widely used to implement various SNN networks.
[0093] LIF neurons integrate input current in a leaky manner , and at its membrane potential Crossing a fixed threshold from below When an action potential is triggered, a spike is emitted at this time, and the process is modeled as a nonlinear function, after which the membrane potential is reset to:
[0094] ;
[0095] In the formula, is a spike sequence The input signal, is a time constant The exponentially decaying membrane potential, is the membrane resistance, the peak The emission is expressed as a nonlinear function of threshold and potential.
[0096] In one specific embodiment, the accelerator includes a pulse detection model controller, a pulse matrix dot product component, a pulse convolution component and a pulse fully connected component, a pulse buffer and a model parameter buffer;
[0097] The pulse detection model controller is respectively connected with the pulse matrix dot product component, the pulse convolution component, the pulse fully connected component, the pulse buffer and the PC end;
[0098] The pulse buffer is connected to the pulse matrix dot product component, the pulse convolution component, and the pulse fully connected component respectively;
[0099] The model parameter buffer is respectively connected to the pulse convolution component, the pulse fully connected component and the off-chip memory;
[0100] The pulse detection model controller is used to control the pulse matrix dot product component, the pulse convolution component and the pulse fully connected component to calculate and output the classification results based on the Gram angular field image according to the improved Spikformer model;
[0101] The model parameter buffer includes a weight buffer and a neuron parameter buffer, the weight buffer is used to cache target weight parameters, and the neuron parameter buffer is used to cache neuron state parameters;
[0102] The pulse convolution component is used to load the target weight parameters and neuron state parameters to perform pulse depth separation convolution operation and output pulse convolution features;
[0103] The impulse matrix dot product component is used to perform impulse matrix dot product operations to output impulse attention features;
[0104] The pulse fully connected component is used to load the target weight parameters and neuron state parameters to perform pulse fully connected calculations to output pulse transformation features;
[0105] The pulse buffer is used to cache pulse convolution features, pulse attention features, pulse transformation features and classification results, and perform bitwise addition of features.
[0106] In a more specific embodiment, the pulse matrix dot product component includes a pulse matrix dot product controller, a matrix dot product calculation core, and a pulse matrix dot product buffer connected in sequence;
[0107] The pulse matrix dot product controller is connected to the pulse detection model controller;
[0108] The matrix dot product calculation core includes a pulse column summation unit and a phase and calculation unit;
[0109] The pulse column summing unit is used to perform column-by-column addition and output pulse summing characteristics;
[0110] The phase-by-phase AND calculation unit is used to perform phase-by-phase AND characteristics of the output pulse;
[0111] The pulse matrix dot product buffer is used to cache pulse summation characteristics and pulse phase and characteristics, and perform pulse nonlinear mapping.
[0112] In a more specific embodiment, the pulse convolution component includes a pulse convolution controller and an array of processing units;
[0113] The pulse convolution controller is connected to the pulse detection model controller and the processing unit array respectively;
[0114] The processing unit array is connected to the pulse buffer and the model parameter buffer respectively;
[0115] The processing unit array includes a plurality of processing units of the array, and the processing units are used to perform event-driven pulse point convolution operations or pulse depth convolution operations.
[0116] In a more specific embodiment, the impulse fully connected component includes an impulse fully connected controller and a linear unit;
[0117] The pulse fully connected controller is connected to the pulse detection model controller and the linear unit respectively;
[0118] The linear unit is connected to the pulse buffer and the model parameter buffer respectively;
[0119] Linear units are used to perform spike fully connected computations.
[0120] It should be noted that according to the framework of the improved Spikformer model, the entire model is mainly divided into three parts, such as Figure 6 As shown, it includes ① pulse encoding operation, ② SSA operation and ③ MLP operation. Correspondingly, the accelerator mainly includes a controller part, a cache part and a computing core part. The controller part includes a pulse detection model controller (Spikformer controller) and each computing core controller. The hierarchical logic of the improved Spikformer model is deployed in the pulse detection model controller, which is responsible for controlling the loading of storage data, the startup of the other computing core controllers and the loading and caching of cache data, so as to realize the correct operation of the data flow of the whole system. The computing core part includes a processing unit array (PE_Array) part for realizing pulse convolution operation, a linear unit (Linear) part for realizing pulse full connection calculation, and a matrix dot product (Matrix Dot-Product Core) part for realizing attention calculation. The cache part is used to cache the neuron state parameters and weights and intermediate output pulses in the calculation process.
[0121] Specifically, the accelerator is designed to include a pulse detection model controller (Spikformer controller), a pulse matrix dot product component, a pulse convolution component and a pulse fully connected component, a pulse buffer (Off-Chip Memory) and a model parameter buffer, the model parameter buffer includes a weight buffer and a neuron parameter buffer, the pulse convolution component includes a pulse convolution controller and a processing unit array, the processing unit array (PE_Array) includes a plurality of processing units (PE) in the array, each processing unit (PE) is used to simulate a pulse deep convolution structure of a cascaded deep convolution-BN unit-LIF neuron or a pulse point convolution structure of a cascaded point convolution-BN unit-LIF neuron, the pulse matrix dot product component includes a pulse matrix dot product controller, a matrix dot product calculation core and a pulse matrix dot product buffer connected in sequence, and the pulse fully connected component includes a pulse fully connected controller and a linear unit; in specific implementation:
[0122] First, the pulse encoding operation of the pulse encoding module is executed, namely ①-(a) pulse depth separation convolution operation: the pulse detection model controller calls the pulse convolution controller, the pulse convolution controller loads the corresponding target weight parameters from the off-chip memory through the weight buffer, and loads the neuron state parameters such as membrane potential from the neuron parameter buffer; based on the Gram angular field image, the data is sorted according to the deep convolution, and then input into the processing unit array for pulse depth convolution operation and the operation result is cached in the pulse buffer. At the same time, in the operation process of the LIF neuron, the new neuron state parameters will be output and cached in the neuron parameter buffer, and then , the operation result of the deep convolution is called from the pulse buffer and the data is arranged according to the point convolution and input into the processing unit array for pulse point convolution operation. The output pulse convolution feature is the first pulse convolution feature of the pulse encoding module and is cached in the pulse buffer; at the same time, multiple processing units (PE) of the processing unit array can operate in parallel and are designed to cooperate with event-driven to achieve only the response to non-zero pulses, that is, if the input pulse data is a non-zero pulse, the corresponding neuron weight and neuron state are loaded for calculation, and if the input is 0, the data will not be loaded for the corresponding operation, thereby speeding up the processing speed and energy consumption of the network;
[0123] Then, in the encoder part: 1) Execute the operation of the SSA module: the pulse detection model controller controls the pulse convolution controller to load the first pulse convolution feature according to the process of ①-(a) pulse depth separation convolution operation, and perform ②-(a) pulse depth separation convolution three times respectively, so as to obtain QKV cache to the pulse buffer; control the pulse matrix dot product controller to load the QKV pulse based on the matrix dot product calculation core to perform ②-(b) pulse matrix dot product operation, the pulse matrix dot product operation includes, through the phase and calculation unit using Q and K to perform the phase and operation to obtain the pulse phase and feature cache to the pulse matrix dot product buffer, using the pulse column summation unit to load the pulse phase and feature to perform the column summation operation to output the pulse summation feature cache to the pulse matrix dot product buffer, and perform pulse nonlinear mapping on the pulse summation feature in the pulse matrix dot product buffer, and continue to use the phase and calculation unit to use the pulse summation feature after nonlinear mapping to perform the phase and operation with V to obtain a new pulse summation feature. That is, the pulse attention features extracted by the pulse self-attention mechanism are cached in the pulse buffer; the pulse fully connected controller is controlled to load the pulse attention features, the corresponding target weight parameters and the neuron state parameters based on the linear unit to perform ②-(c) pulse full connection operation. The pulse full connection operation simulates the pulse full connection calculation, that is, the linear transformation of the Linear unit, the batch normalization of the BN unit and the pulse nonlinear mapping of the LIF neuron in sequence, and the output pulse transformation feature, that is, the corresponding first pulse transformation feature is cached in the pulse buffer, thereby completing the SSA operation; 2) In the pulse buffer, the first pulse convolution feature and the first pulse transformation feature are performed bitwise addition to obtain the bitwise addition feature, which is similar to the residual connection operation; 3) The pulse fully connected controller is controlled to load the bitwise addition feature to perform ③-(a) pulse full connection operation, so as to complete the pulse full connection calculation of the MLP module, and output the new pulse transformation feature, that is, the second pulse transformation feature is cached in the pulse buffer;
[0124] Finally, the pulse fully connected controller is controlled to load the second pulse transformation feature according to the pulse fully connected operation of ③-(a), complete the pulse fully connected calculation of the classification head, and output the new pulse transformation feature, that is, the classification result is cached in the pulse buffer.
[0125] In a specific implementation, the accelerator is also used to send the classification result to the PC for display. More specifically, the PC is connected to the pulse buffer.
[0126] In a specific implementation, the process of obtaining the target weight parameter includes:
[0127] Using the training image set to train the improved Spikformer model to be trained, and determining the trained improved Spikformer model;
[0128] Perform weight quantization training on the trained improved Spikformer model to determine the target weight parameters.
[0129] It should be noted that in order to enable the hardware acceleration platform based on the improved Spikformer model to show better performance, such as Figure 2 As shown in the figure, after the improved Spikformer model is built, in the training part of the Spikformer model, the training image set is used, and the Multi-Gaussian function is used as a substitute function of the BPTT algorithm to train the improved Spikformer model to be trained. After the trained improved Spikformer model is obtained, the pre-trained weight parameters are used for weight quantization. The weight quantization training is used to fine-tune the trained improved Spikformer model, which is beneficial to the improvement of hardware reasoning accuracy. After the quantization training is completed, the quantized weight parameters, i.e., the target weight parameters, can be loaded into the hardware architecture;
[0130] Weight quantization training aims to quantize the function The model's weight tensor W is quantized, including: first, setting a quantization factor n, converting the original continuous decimal value domain into a discrete integer value set by multiplying x by n and performing rounding operations; then, dividing the obtained integer value by n to reduce the value to a multiple of 1 / n; finally, using the clip() function to limit the range of the reduced value to ensure that all values are in the interval [a, b], where values less than a will be adjusted to a, and values greater than b will be adjusted to b; through this quantization process, the values originally represented by 32-bit floating point numbers can be effectively mapped to 8-bit integer values in the interval [a, b], thereby significantly reducing the computational complexity and improving processing efficiency.
[0131] In an embodiment of the present invention, the electroencephalogram signal is converted into a Gram angle field image for a simple data compression, and then the electroencephalogram epilepsy data is classified by the improved Spikformer model. During the feature processing, the feature maintains the pulse data form, and the inherent locality of the pulse deep separation convolution is used to effectively capture the local features and spatial hierarchy. The pulse-driven self-attention mechanism is used to replace the traditional self-attention mechanism, which can capture long-distance dependencies. The pulse self-attention mechanism uses binary pulse signals for calculation, and all matrix multiplications related to the pulse matrix can be converted into sparse addition, thereby converting the complex multiplication and matrix calculations in the self-attention mechanism into simple column-by-column addition operations and bit-by-bit and operations. In addition, the pulse data is sparse in nature, which can make the model more efficient when processing large-scale sparse data. At the same time, combined with the event-driven characteristics, it can achieve hardware friendliness as a whole, thereby realizing a hardware acceleration platform based on FPGA that supports the improved Spikformer model, accelerating the classification of electroencephalogram epilepsy data, and achieving high-performance, low-power electroencephalogram classification. In addition, due to the lower energy consumption, the improved Spikformer model can be deployed on edge devices, making it possible to apply it to wearable devices.
[0132] See also Figure 7 , Figure 7 A flowchart of the steps of a hardware acceleration method for classifying electroencephalogram epilepsy data provided by an embodiment of the present invention.
[0133] The present invention provides a hardware acceleration method for electroencephalogram epilepsy data classification, comprising:
[0134] Step 701: Acquire an electroencephalogram signal and convert the electroencephalogram signal into a Gram's angle field image.
[0135] It should be noted that when acquiring the electroencephalogram signal, the electroencephalogram signal is normalized and polar coordinate processed, and then converted into a corresponding Grammar angle field image through Grammar angle field imaging processing.
[0136] Step 702: Classify the EEG epilepsy data based on the Gram angular field image according to the improved Spikformer model, and output the classification result.
[0137] It should be noted that the improved Spikformer model includes a pulse encoding module, an encoder and a classification head, the pulse encoding module includes a pulse depth separation convolution layer, the encoder includes a pulse self-attention module and an MLP module, the MLP module includes at least one pulse fully connected layer, the pulse fully connected layer includes a Linear unit, a BN unit and a LIF neuron, the pulse self-attention module includes multiple pulse depth separation convolution layers, LIF neurons and pulse fully connected layers, and the classification head includes a pulse fully connected layer;
[0138] The processing process of the improved Spikformer model includes: performing a pulse depth separation convolution operation on the Gram angular field image to generate a first pulse convolution feature; performing multiple pulse depth separation convolution operations on the first pulse convolution feature to map it into a pulse query matrix, a pulse key matrix and a pulse value matrix respectively; using the pulse query matrix and the pulse key matrix to add them column by column after bitwise AND, and performing pulse nonlinear mapping and bitwise AND with the pulse value matrix to generate a pulse attention feature through a pulse matrix dot product operation; performing a pulse full connection calculation on the pulse attention feature to determine a first pulse transformation feature; performing a residual connection on the first pulse convolution feature and the first pulse transformation feature to output a bitwise addition feature; performing a pulse full connection calculation on the bitwise addition feature to construct a second pulse transformation feature; performing a pulse full connection calculation on the bitwise addition feature after residual connection on the bitwise addition feature and the second pulse transformation feature to output a classification result; the pulse depth separation convolution operation includes sequentially performing a pulse depth convolution operation and a pulse point convolution operation;
[0139] Among them, the pulse deep separation convolution layer includes deep convolution, BN unit, LIF neuron, point convolution, BN unit and LIF neuron; the processing process of the pulse deep separation convolution unit includes: performing pulse deep convolution operation on the input pulse convolution input feature through cascaded deep convolution, BN unit and LIF neuron to generate deep features, using cascaded point convolution, BN unit and LIF neuron to perform pulse point convolution operation on the deep features, and outputting pulse convolution output features.
[0140] In an embodiment of the present invention, the EEG signal is converted into a Gram angular field image for a simple data compression, and then the EEG epilepsy data is classified through the improved Spikformer model. During the feature processing, the features remain in the form of pulse data, and the inherent locality of the pulse deep separation convolution is used to effectively capture local features and spatial hierarchical structures. The complex multiplication and matrix calculations in the self-attention mechanism are converted into simple column-by-column addition operations and bit-by-bit AND operations, which can achieve high-performance, low-power EEG epilepsy data classification.
[0141] An embodiment of the present invention further provides a computer device, including a memory and a processor, wherein a computer program is stored in the memory; when the computer program is executed by the processor, the processor executes the steps of the hardware acceleration method for classifying electroencephalogram epilepsy data in any of the above embodiments.
[0142] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of the hardware acceleration method for classifying electroencephalogram epilepsy data as in any of the above embodiments are implemented.
[0143] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the hardware acceleration method for classifying electroencephalogram epilepsy data as in any of the above embodiments.
[0144] In the several embodiments provided in the present application, it should be understood that the disclosed hardware acceleration platform and method can be implemented in other ways. For example, the hardware acceleration platform embodiment described above is only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0145] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0148] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hardware acceleration platform for EEG epilepsy data classification, characterized in that: include: The PC side, accelerator and off-chip memory are connected in sequence; The PC is equipped with an image processor, which is used to obtain electroencephalogram signals and convert the electroencephalogram signals into Gram's angle field images; The off-chip memory is used to store target weight parameters of the improved Spikformer model; The accelerator is used to load the target weight parameters, and classify the electroencephalogram epilepsy data based on the Gram angle field image according to the improved Spikformer model, and output the classification result.
2. The hardware acceleration platform for EEG epilepsy data classification according to claim 1, characterized in that: The processing process of the improved Spikformer model includes: Performing a pulse depth separation convolution operation on the Gram angular field image to generate a first pulse convolution feature; Perform multiple pulse depth separation convolution operations on the first pulse convolution feature, and map them into a pulse query matrix, a pulse key matrix and a pulse value matrix respectively; The pulse query matrix and the pulse key matrix are phase-ORed and then added column by column, and the pulses are nonlinearly mapped and then phase-ORed with the pulse value matrix to generate pulse attention features through pulse matrix dot product operation; Performing a pulse full connection calculation on the pulse attention feature to determine a first pulse transformation feature; Performing a residual connection between the first pulse convolution feature and the first pulse transformation feature, and outputting a bitwise addition feature; Performing pulse full connection calculation on the bitwise addition feature to construct a second pulse transformation feature; After the bitwise addition feature is connected with the second pulse transformation feature residual, a pulse full connection calculation is performed to output a classification result; The pulse depth separation convolution operation includes sequentially performing a pulse depth convolution operation and a pulse point convolution operation.
3. The hardware acceleration platform for EEG epilepsy data classification according to claim 2, characterized in that: The accelerator includes a pulse detection model controller, a pulse matrix dot product component, a pulse convolution component and a pulse fully connected component, a pulse buffer and a model parameter buffer; The pulse detection model controller is respectively connected to the pulse matrix dot product component, the pulse convolution component, the pulse fully connected component, the pulse buffer and the PC end; The pulse buffer is respectively connected to the pulse matrix dot product component, the pulse convolution component, and the pulse fully connected component; The model parameter buffer is connected to the pulse convolution component, the pulse fully connected component and the off-chip memory respectively; The pulse detection model controller is used to control the pulse matrix dot product component, the pulse convolution component and the pulse fully connected component to calculate and output a classification result based on the Gram angular field image according to the improved Spikformer model; The model parameter buffer includes a weight buffer and a neuron parameter buffer, the weight buffer is used to cache target weight parameters, and the neuron parameter buffer is used to cache neuron state parameters; The pulse convolution component is used to load the target weight parameters and the neuron state parameters to perform pulse depth separation convolution operation to output the pulse convolution feature; The pulse matrix dot product component is used to perform a pulse matrix dot product operation to output a pulse attention feature; The pulse fully connected component is used to load the target weight parameter and the neuron state parameter to perform pulse fully connected calculation to output the pulse transformation feature; The pulse buffer is used to cache pulse convolution features, pulse attention features, pulse transformation features and classification results, and perform bit-wise addition of features.
4. The hardware acceleration platform for EEG epilepsy data classification according to claim 3, characterized in that: The pulse matrix dot product component includes a pulse matrix dot product controller, a matrix dot product calculation core and a pulse matrix dot product buffer connected in sequence; The pulse matrix dot product controller is connected to the pulse detection model controller; The matrix dot product calculation core includes a pulse column summing unit and a phase and calculation unit; The pulse column summing unit is used to perform column-by-column addition to output pulse summing characteristics; The phase-by-phase AND calculation unit is used to perform phase-by-phase AND characteristics of output pulses; The pulse matrix dot product buffer is used to buffer pulse summation characteristics and pulse phase and characteristics, and perform pulse nonlinear mapping.
5. The hardware acceleration platform for EEG epilepsy data classification according to claim 3, characterized in that: The pulse convolution component includes a pulse convolution controller and a processing unit array; The pulse convolution controller is connected to the pulse detection model controller and the processing unit array respectively; The processing unit array is connected to the pulse buffer and the model parameter buffer respectively; The processing unit array includes a plurality of processing units in an array, and the processing units are used to perform event-driven pulse point convolution operations or pulse depth convolution operations.
6. The hardware acceleration platform for EEG epilepsy data classification according to claim 3, characterized in that: The pulse fully connected component includes a pulse fully connected controller and a linear unit; The pulse fully connected controller is connected to the pulse detection model controller and the linear unit respectively; The linear unit is connected to the pulse buffer and the model parameter buffer respectively; The linear unit is used to perform pulse full connection calculation.
7. The hardware acceleration platform for EEG epilepsy data classification according to claim 1, characterized in that: The process of obtaining the target weight parameter includes: Using the training image set to train the improved Spikformer model to be trained, and determining the trained improved Spikformer model; Perform weight quantization training on the trained improved Spikformer model to determine the target weight parameters.
8. A hardware acceleration method for EEG epilepsy data classification, characterized in that: include: Acquiring an electroencephalogram signal, and converting the electroencephalogram signal into a Gram's angle field image; The electroencephalogram epilepsy data is classified based on the Gram angle field image according to the improved Spikformer model, and the classification result is output.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the hardware acceleration method for classifying electroencephalogram epilepsy data as claimed in claim 8.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 8 are implemented.
Citation Information
Patent Citations
Method for parallel acceleration of electroencephalogram signal processing process based on GPU
CN113420672A
Electroencephalogram signal classification method and apparatus, device, storage medium and program product
WO2022183966A1