Brain-computer interface-oriented efficient coding and decoding intelligent system

Through heterogeneous computing architecture and neural network accelerator optimization, the hardware calculation and power consumption problems of the brain-computer interface system in high-spatial-time resolution neural signal analysis are solved, and the efficient chipization of the brain-computer interface codec model is realized, improving the real-time performance and energy efficiency ratio of the system.

CN120335596APending Publication Date: 2025-07-18JILIN UNIV FIRST HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226358.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing brain-computer interface systems have challenges in hardware computing capabilities, signal processing efficiency and system power consumption in the analysis and decoding of high-spatial-time resolution neural signal, which is difficult to meet practical application needs.

Method used

It adopts a heterogeneous computing architecture, including embedded CPUs and neural network accelerators, combines reinforcement learning to optimize task mapping of deep neural networks, and adopts data multiplexing technology, clock gating and multi-level storage architecture to optimize hardware resource configuration and realize efficient on-chip neural network computing.

Benefits of technology

Under the condition that the accuracy is close to lossless, the storage and calculation amount of the codec model is reduced, and the chipization of the codec model with high spatio-temporal resolution brain-computer interface is achieved, which improves the real-time and energy efficiency ratio of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335596A_ABST
    Figure CN120335596A_ABST
Patent Text Reader

Abstract

The invention provides an efficient encoding and decoding intelligent system for a brain-computer interface. The efficient encoding and decoding intelligent system comprises an electroencephalogram acquisition module, an intelligent processing chip and a communication module, wherein the electroencephalogram acquisition module is used for acquiring electroencephalogram signals; the communication module is used for transmitting the electroencephalogram signal to the intelligent processing chip; the intelligent processing chip is used for calculating and analyzing the electroencephalogram signals in real time. According to the method, for the brain-computer interface signal with high temporal-spatial resolution, the storage and calculation amount of the coding and decoding model is effectively reduced under the condition of ensuring that the precision is nearly lossless, and the compression and calculation acceleration of the model are realized. According to the invention, a special hardware architecture is designed, the hardware resource configuration is optimized, the data carrying is reduced, and the chipping of the high-temporal-spatial-resolution brain-computer interface coding and decoding model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brain - computer interfaces, and particularly to an efficient encoding - decoding intelligent system for brain - computer interfaces. Background Art

[0002] In recent years, brain - computer interface (BCI) technology has made important progress in the fields of neurorehabilitation, medical diagnosis, brain science research, etc. However, existing brain - computer interface systems still face many challenges in data acquisition, signal processing, and real - time response capabilities. Especially in the analysis and decoding of neural signals with high spatio - temporal resolution, limited by factors such as hardware computing power, signal processing efficiency, and system power consumption, it is difficult to meet the actual application requirements.

[0003] Existing brain - computer interface chips mainly focus on basic signal acquisition, filtering, and low - level data processing, and have not yet formed a complete system architecture capable of supporting the decoding of neural signals with high spatio - temporal resolution. Currently, invasive brain - computer interface chips face key technical bottlenecks such as damage to the body by implanted electrodes and low - power high - precision signal processing. In contrast, although non - invasive brain - computer interface technology has advantages in terms of safety and applicability, there is still much room for improvement in chip architecture optimization, interdisciplinary integration, and system integrity. In particular, existing research mainly focuses on the optimization of single - point technologies such as low - power design, analog signal amplification, and wireless transmission, while the intelligent chip for brain - computer interfaces with high spatio - temporal resolution is still in the exploratory stage and has not yet formed a complete technical system. Summary of the Invention

[0004] In order to achieve the above - mentioned objects and other advantages of the present invention, the object of the present invention is to provide an efficient encoding - decoding intelligent system for brain - computer interfaces, including an electroencephalogram acquisition module, an intelligent processing chip, and a communication module; wherein,

[0005] The electroencephalogram acquisition module is used to acquire electroencephalogram signals;

[0006] The communication module is used to transmit the electroencephalogram signals to the intelligent processing chip;

[0007] The intelligent processing chip is used to perform real - time calculation and analysis on the electroencephalogram signals.

[0008] Further, the intelligent processing chip adopts a heterogeneous computing architecture, including an embedded CPU and a neural network accelerator; wherein, the embedded CPU is used for peripheral auxiliary computing, data transmission, and task scheduling, and the neural network accelerator is used to execute computationally intensive tasks in electroencephalogram signal processing.

[0009] Further, the neural network accelerator includes a computing unit, a storage unit, and a control unit.

[0010] Further, the neural network accelerator uses reinforcement learning to optimize the task mapping of the deep neural network, including the following steps:

[0011] Input information: neural network target model, platform constraints, optimization objectives;

[0012] RL agent training: The agent continuously generates resource allocation and task mapping schemes, and evaluates the quality of the schemes through environmental feedback.

[0013] Reward mechanism: The environment feedbacks a reward signal according to the computing performance and resource utilization rate to guide the policy network to optimize the decision-making.

[0014] Optimization result: Generate a resource allocation and computing mapping strategy that meets the application requirements to achieve efficient on-chip neural network computing.

[0015] Further, the on-chip data flow pattern of the neural network accelerator is determined according to the neural network structure, the utilization rate of the computing unit, and the number of memory accesses, including the systolic array pattern, the two-dimensional mapping pattern, and the single instruction multiple data pattern.

[0016] Further, the storage unit adopts data reuse technology and a multi-level storage architecture.

[0017] Further, the neural network accelerator adopts data gating and clock gating. The data gating is used to skip the calculation of zero-value data, and the clock gating is used to turn off the clock signal when the computing unit is idle.

[0018] Further, the neural network accelerator configures the parallelism of the computing unit and the size of the on-chip cache according to the on-chip resource situation of the FPGA, and conducts the overall architecture design of the on-chip system, including programmable logic (PL) design, accelerator interface optimization, and interconnection bus design.

[0019] Further, the neural network adopts a deep neural network, and uses knowledge distillation technology to transfer the knowledge of complex networks to lightweight models, pruning-based structured sparse optimization, and fixed-point quantization of network parameters and activation values.

[0020] Further, the fixed-point quantization of network parameters and activation values includes:

[0021] Data analysis: Collect and analyze the parameter and activation value distributions of each layer, and determine the optimal number of bits in combination with the task type and hardware characteristics.

[0022] Quantization mapping: Based on deterministic or random quantization strategies, map the parameters and activation values to a finite set of fixed-point numbers.

[0023] Accuracy recovery: Through network fine-tuning, compensate for the accuracy loss introduced during the quantization process.

[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0025] For the high spatio-temporal resolution brain-computer interface signals, the present invention can effectively reduce the storage and computational amount of the encoding and decoding model while ensuring nearly lossless accuracy, so as to realize the compression and computational acceleration of the model. The present invention designs a dedicated hardware architecture, optimizes the hardware resource configuration, reduces data transfer, and realizes the chipization of the encoding and decoding model of the high spatio-temporal resolution brain-computer interface.

[0026] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following takes the preferred embodiments of the present invention and combines with the accompanying drawings to elaborate in detail as follows. The specific implementation manners of the present invention are given in detail by the following embodiments and their accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0028] Figure 1 is a schematic diagram of an efficient encoding and decoding intelligent system for a brain-computer interface;

[0029] Figure 2 is a schematic diagram of data communication and synchronization of a multi-core (heterogeneous) system;

[0030] Figure 3 is a schematic diagram of a neural network accelerator;

[0031] Figure 4 is a schematic diagram of a method for allocating deep neural network computing tasks based on reinforcement learning;

[0032] Figure 5 is a flow chart of neural network compression;

[0033] Figure 6 is a flow chart of the overall verification experiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] In the accompanying drawings, for clarity, the shapes and dimensions may be enlarged, and the same reference numerals will be used throughout the drawings to indicate the same or similar components.

[0036] In the following description, terms such as center, thickness, height, length, front, back, rear, left, right, top, bottom, upper, lower, etc. are defined relative to the configurations shown in the respective drawings. In particular, "height" corresponds to the dimension from top to bottom, "width" corresponds to the dimension from left to right, and "depth" corresponds to the dimension from front to back. These are relative concepts and may accordingly change depending on their different positions and usage states. Therefore, these or other orientations should not be construed as restrictive terms.

[0037] Terms related to attachment, connection, etc. (e.g., "connect" and "attach") refer to the relationship in which these structures are directly or indirectly fixed or attached to each other through an intermediate structure, as well as a movable or rigid attachment or relationship, unless otherwise explicitly stated.

[0038] Embodiment 1

[0039] An efficient encoding and decoding intelligent system for a brain-computer interface (BCI) adopts model compression and computing acceleration technologies, an optimized scheduling strategy for computing tasks, and a high-performance chip architecture design method to improve the spatio-temporal resolution and real-time computing ability of neural signal processing. As Figure 1 shown, the system includes an electroencephalogram (EEG) acquisition module, an intelligent processing chip, and a communication module; wherein,

[0040] The EEG acquisition module is used to acquire EEG signals;

[0041] The communication module is used to transmit the EEG signals to the intelligent processing chip;

[0042] The intelligent processing chip is used to perform real-time calculation and analysis on the EEG signals.

[0043] Specifically, the EEG signals collected at the EEG acquisition end are transmitted wirelessly to the intelligent processing chip of the backend brain-computer interface for real-time calculation and analysis after signal conditioning. For this purpose, the system needs to integrate a wireless communication module to ensure the efficient transmission of EEG signals. According to the requirements of different rehabilitation devices, the output of the core computing system of EEG signals can be selected to transmit data wirelessly or wired.

[0044] The core computing system is based on an on-chip computing architecture to achieve efficient acceleration of neural network algorithms. For complex neural network operations, the data storage structure is optimized to ensure that the on-chip storage has sufficient bandwidth to avoid data transmission becoming a bottleneck in system performance. During the architecture design process, it is necessary to integrate each computing module and complete the control logic design to enable it to cooperate efficiently with an external host. In addition, drivers and compilers for the neural network accelerator need to be developed to support flexible configuration and invocation of algorithms.

[0045] Due to the characteristics of neural network computing tasks, not all operations can be efficiently executed on the accelerator, and some of the calculations are more suitable for CPU processing. Therefore, the system needs to reasonably partition the computing tasks and allocate different operations to the most suitable computing units to improve the overall computing efficiency. During this process, the interface between the accelerator and the CPU is optimized to reduce communication overhead and ensure that the system performance is not affected by data transmission bottlenecks when the two work together.

[0046] The present invention provides a computing task scheduling strategy for brain-computer interface encoding and decoding models for heterogeneous platforms, as Figure 2 shown. The present invention proposes a hardware acceleration scheme to improve real-time processing capabilities. The system adopts a heterogeneous computing architecture, where the CPU is mainly responsible for peripheral auxiliary computing, data transmission, and task scheduling, while computationally intensive tasks are completed by a dedicated accelerator.

[0047] Task allocation based on computing characteristics: Due to architectural limitations, embedded CPUs are not good at large-scale numerical calculations. Therefore, a dedicated hardware accelerator is designed to improve computing efficiency. The system architecture includes:

[0048] Dedicated accelerator: Responsible for computationally intensive tasks in electroencephalogram signal processing, such as deep neural network inference.

[0049] CPU tasks: Perform peripheral auxiliary calculations such as electroencephalogram signal acquisition and result post-processing.

[0050] Since there are strong data dependencies between tasks and the computing efficiency of different hardware varies greatly, it is necessary to reasonably partition the computing tasks and optimize data communication and synchronization to improve the overall processing efficiency.

[0051] Optimization of deep neural network computing tasks based on reinforcement learning: A neural network accelerator usually consists of processing elements (SIMD Core), storage units (Weight BRAM, Act BRAM), and control units (INSTR CTRL, DATA CTRL, COMPUTE CTRL), as Figure 3As shown in the figure. Under the condition of limited on-chip logic resources, it is crucial to efficiently utilize computing units. The present invention uses reinforcement learning (RL) to optimize the task mapping of deep neural networks to improve computing efficiency.

[0052] Computing task mapping strategy: Taking RNN / LSTM for processing serialized EEG signals as an example, neural network computing mainly involves large-scale matrix multiplication. For example, if the size of the computing array is 2×2, the calculation can be split into sub-block matrix operations:

[0053]

[0054] Reinforcement learning optimization strategy: Reinforcement learning (RL) continuously optimizes the policy network by interacting with the environment to maximize computing efficiency. As Figure 4 shown, the specific process is as follows:

[0055] Input information: Neural network target model, platform constraints (latency / power consumption), optimization objectives.

[0056] RL agent training: The agent continuously generates resource allocation and task mapping schemes (i.e., "actions") and evaluates the pros and cons of the schemes through environmental feedback.

[0057] Reward mechanism: The environment feeds back a "reward" signal according to computing performance and resource utilization to guide the policy network to optimize decisions.

[0058] Optimization result: Generate a resource allocation and computing mapping strategy that meets application requirements to achieve efficient on-chip neural network computing.

[0059] This method can optimize resource utilization in a heterogeneous computing environment, improve the inference efficiency of neural networks, and ensure that embedded devices meet the real-time requirements of the brain-computer interface system under limited computing power.

[0060] The present invention provides a method for designing an efficient data flow and computing architecture for a brain-computer interface chip. Based on the FPGA platform, the present invention designs an intelligent computing chip architecture suitable for EEG signal processing to achieve high-speed and low-power neural network forward acceleration computing and improve the computing efficiency and energy utilization rate of the system.

[0061] Data flow optimization and architecture selection: The data flow pattern of the neural network accelerator on the chip directly affects the architecture design and computing performance. Currently, common data flow patterns include:

[0062] Systolic Array, 2D-Mapping, Single Instruction Multiple Data (SIMD).

[0063] Different data flow patterns have their own advantages and disadvantages in terms of computing unit utilization, memory access frequency, power consumption, and computational complexity. For the specific neural network algorithm adopted in this system, it is necessary to deeply analyze the network structure, weigh the utilization of computing units against the number of memory accesses, and select the optimal data flow scheme to ensure the best balance between computing performance and energy consumption.

[0064] Memory Optimization and Memory Access Management: Memory design has an important impact on the power consumption and performance of the FPGA computing architecture. Memory access operations usually consume more energy and time than computing operations. Therefore, reducing the number of memory accesses is crucial. Optimization strategies include:

[0065] Data Reuse Technology: Reduce the memory access frequency, cache and reuse data on-chip to the greatest extent, and reduce the power consumption overhead caused by memory access.

[0066] Multi-level Memory Architecture: Store data hierarchically through on-chip caches (SRAM / BRAM), optimize the data flow, and improve the data acquisition efficiency.

[0067] Low-power Optimization Strategies: In addition to memory access optimization, there are a large number of zero values involved in the neural network calculation process. For example, many multipliers are zero in matrix multiplication, resulting in the ineffective occupation of computing resources. Therefore, the following low-power technologies are adopted to further optimize the calculation process:

[0068] Data Gating: Skip the calculation of zero-value data and reduce the occupation of redundant computing resources.

[0069] Clock Gating: Turn off the clock signal when the computing unit is idle to reduce the dynamic power consumption.

[0070] To reduce ineffective calculations, improve the energy efficiency ratio of computing units, and further reduce power consumption.

[0071] Computing Unit Parallelism and Architecture Design: According to the on-chip resources of the FPGA, reasonably configure the parallelism of computing units and the size of on-chip caches, which will directly determine the computing power of the accelerator. Subsequently, conduct the overall architecture design of the system-on-chip, including:

[0072] Programmable Logic (PL) Design: Ensure adaptation to different task requirements and improve flexibility.

[0073] Accelerator Interface Optimization: Achieve efficient communication with the CPU and memory.

[0074] Interconnection Bus Design: Optimize the data flow transmission and improve the coordination efficiency of computing and storage.

[0075] Through the above architecture optimization solutions, achieve efficient and low-power computing for electroencephalogram signal intelligent processing under the FPGA platform, and improve the real-time performance and energy efficiency ratio of the system.

[0076] The present invention provides a method for compressing deep neural networks for electroencephalogram (EEG) signal processing. Based on deep neural networks (DNNs), the present invention constructs an efficient and accurate intelligent processing system for brain-computer interfaces (BCIs). Since DNNs have a large amount of computation and slow running speed and are difficult to be directly applied to embedded devices, the present invention conducts in-depth research on the acceleration and compression technologies of neural networks, and realizes efficient real-time operation while maintaining the accuracy of the full-precision model. The optimization of neural networks mainly includes the following four aspects, as Figure 5 shown.

[0077] Neural network fitting based on knowledge distillation: Complex DNNs can improve the accuracy of EEG signal processing, but have a large computational overhead. Therefore, the present invention adopts the knowledge distillation (KD) technology to transfer the knowledge of complex networks to lightweight models, so that they can still maintain excellent performance in environments with limited computing resources.

[0078] Structured sparse optimization based on pruning. Neural network pruning can effectively reduce the amount of computation and storage requirements. The main advantages include:

[0079] Storage optimization: Compress the model through sparse coding to reduce storage occupancy.

[0080] Computation acceleration: The introduction of zero-valued weights reduces the amount of computation and improves the running efficiency.

[0081] However, unstructured pruning will increase the computational complexity of hardware accelerators and introduce problems such as load imbalance. Therefore, for the EEG signal sequence processing task, the present invention optimizes the matrix multiplication calculation mode and adopts structured sparse training (such as imposing sparse constraints on the weight matrix row / column by row / column) to improve the hardware adaptability.

[0082] Fixed-point quantization of weights and activations: Perform fixed-point quantization on network parameters and activation values. The specific process is as follows:

[0083] Data analysis: Collect and analyze the distribution of parameters and activation values of each layer, and determine the optimal number of bits in combination with the task type and hardware characteristics.

[0084] Quantization mapping: Map parameters and activation values to a finite set of fixed-point numbers based on deterministic or random quantization strategies.

[0085] Accuracy recovery: Through network fine-tuning, compensate for the accuracy loss introduced during the quantization process.

[0086] Through the above methods, the computational complexity of deep neural networks is reduced.

[0087] To verify the effectiveness of the integrated system, an electroencephalogram (EEG) experiment test of the brain-computer interface (BCI) integrated system is conducted in this embodiment. The experiment mainly examines the following factors: the effectiveness of EEG signal acquisition, the accuracy of the classification algorithm, and the applicability of the system to stroke patients.

[0088] Two motor imagery paradigms are adopted in the experiment to verify the effectiveness of the system architecture. The overall experimental process is as Figure 6 shown.

[0089] Experimental method: The subjects first undergo a 40-minute motor imagery BCI screening to select suitable subjects to verify the effectiveness of the algorithm. Subsequently, a 40-minute motor imagery BCI calibration is carried out, which is the same as the screening stage. The formal experiment is divided into Experiment 1 (30 minutes) and Experiment 2 (30 minutes), totaling 60 minutes, with appropriate rest periods set in between. The selected subjects will participate in repeated experiments twice a week for 6 months, and the experimental data is used to evaluate the reliability and repeatability of the system.

[0090] The motor imagery experimental paradigms include:

[0091] Paradigm 1: Agility test

[0092] This paradigm mainly examines the response speed of the system. The experimental process is as follows:

[0093] Cue phase (3 s): A small ball appears on the left or right side at the top of the screen.

[0094] Motor imagery phase (5 s): The small ball moves downward, and the subject needs to imagine the corresponding hand clenching action to catch the ball according to the position of the ball (left or right).

[0095] Feedback phase (2 s): The system displays the corresponding feedback (such as a red ball) on the screen according to the EEG signal classification result.

[0096] The experiment includes a training stage and a testing stage:

[0097] Training stage: 60 trials (30 times for each hand, randomly ordered), used to establish a classification model.

[0098] Testing stage: 60 trials (30 times for each hand, randomly ordered), used to evaluate the performance of the classifier.

[0099] The entire experiment consists of 120 trials, with a 2-s interval between trials.

[0100] Paradigm 2: Manipulation fineness test

[0101] This paradigm mainly examines the stability and accuracy of the system in fine hand movement imagery tasks. The experimental process is as follows:

[0102] Rest phase (2 s): The subject relaxes.

[0103] Preparation stage (3 s): The image of a clenched fist is displayed on the screen and a countdown is carried out.

[0104] Task stage (4 s): The subject imagines the right hand performing the alternating actions of opening and clenching (repeating every 1 s) according to the instructions.

[0105] Task interval (2 s): The subject can blink or adjust the state to reduce the interference of motion artifacts.

[0106] During the experiment, the subject needs to maintain a stable state, without blinking or moving, to ensure the quality of the EEG signals.

[0107] Experimental data analysis: After the experiment, EEG data of multiple people in multiple batches will be collected, and statistical analysis methods will be used to evaluate the computing performance, classification accuracy and system applicability of the intelligent chip under different task conditions. The experimental results will further verify the effectiveness of the chip architecture and provide a reference basis for system optimization.

[0108] The number of devices and the processing scale described here are used to simplify the description of the present invention. Applications, modifications and variations of the present invention will be apparent to those skilled in the art.

[0109] Although the embodiments of the present invention have been disclosed above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated examples here.

[0110] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.

[0111] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the focus of each embodiment is on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0112] The above description is only for the embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of one or more embodiments of this specification.

Claims

1. An efficient encoding and decoding intelligent system for a brain-computer interface, characterized in that: It includes an electroencephalogram (EEG) acquisition module, an intelligent processing chip, and a communication module; among them, the EEG acquisition module is used to acquire EEG signals; the communication module is used to transmit the EEG signals to the intelligent processing chip; the intelligent processing chip is used to perform real-time calculation and analysis on the EEG signals.

2. The efficient encoding and decoding intelligent system for a brain-computer interface according to claim 1, wherein: The intelligent processing chip adopts a heterogeneous computing architecture, including an embedded CPU and a neural network accelerator; among them, the embedded CPU is used for peripheral auxiliary computing, data transmission, and task scheduling, and the neural network accelerator is used to execute computationally intensive tasks in EEG signal processing.

3. The efficient encoding and decoding intelligent system for a brain-computer interface according to claim 2, characterized in that: The neural network accelerator includes a computing unit, a storage unit, and a control unit.

4. The efficient encoding and decoding intelligent system for a brain-computer interface according to claim 3, wherein: The neural network accelerator adopts reinforcement learning to optimize the task mapping of the deep neural network, including the following steps: Input information: neural network target model, platform constraints, optimization objectives; RL agent training: The agent continuously generates resource allocation and task mapping schemes, and evaluates the quality of the schemes through environmental feedback; Reward mechanism: The environment feedbacks a reward signal according to the computing performance and resource utilization rate to guide the policy network to optimize the decision-making; Optimization result: Generate a resource allocation and computing mapping strategy that meets the application requirements to achieve efficient on-chip neural network computing.

5. An efficient encoding and decoding intelligent system for a brain-computer interface according to claim 3, characterized in that: The data flow mode of the neural network accelerator on the chip is determined according to the neural network structure, the utilization rate of the computing unit, and the number of memory accesses, including the systolic array mode, the two-dimensional mapping mode, and the single instruction multiple data mode.

6. An efficient encoding and decoding intelligent system for a brain-computer interface according to claim 3, characterized in that: The storage unit adopts data reuse technology and a multi-level storage architecture.

7. The efficient encoding and decoding intelligent system for a brain-computer interface according to claim 3, wherein: The neural network accelerator adopts data gating and clock gating. The data gating is used to skip the calculation of zero-value data, and the clock gating is used to turn off the clock signal when the computing unit is idle.

8. The efficient encoding and decoding intelligent system for a brain-computer interface according to claim 3, characterized in that: The neural network accelerator configures the parallelism of the computing unit and the size of the on-chip cache according to the on-chip resource situation of the FPGA, and conducts the overall architecture design of the on-chip system, including programmable logic (PL) design, accelerator interface optimization, and interconnection bus design.

9. An efficient encoding and decoding intelligent system for a brain-computer interface according to claim 2, wherein: The neural network adopts a deep neural network, and uses knowledge distillation technology to transfer the knowledge of complex networks to lightweight models, pruning-based structured sparse optimization, and fixed-point quantization of network parameters and activation values.

10. An efficient encoding and decoding intelligent system for a brain-computer interface according to claim 9, characterized in that: The fixed-point quantization of network parameters and activation values includes: Data analysis: Collect and analyze the parameter and activation value distributions of each layer, and determine the optimal number of bits in combination with the task type and hardware characteristics; Quantization mapping: Based on deterministic or random quantization strategies, map the parameters and activation values to a finite set of fixed-point numbers; Accuracy recovery: Through network fine-tuning, compensate for the accuracy loss introduced during the quantization process.