Edge inference neural network coprocessor circuit for electroencephalogram signal processing

By designing an edge inference neural network coprocessor circuit for EEG signal processing, the problems of low computing efficiency and large hardware resource overhead in the prior art are solved, and efficient EEG signal processing is realized, which is suitable for resource-constrained devices.

CN120197660APending Publication Date: 2025-06-24BEIJING SONGGUO BRAIN MACHINE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510165630.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art has problems in the processing of EEG signal in low computing efficiency, high hardware resource overhead, and difficulty in adapting to complex tasks and high noise environments.

Method used

A coprocessor circuit of edge inference neural network for EEG signal processing is designed, including linear operation unit, nonlinear operation unit and convolutional operation unit, supporting 16-bit fixed-point number operation, and optimizing redundant reading operations in convolutional operation.

Benefits of technology

It improves computing efficiency and reduces hardware resource overhead. It is suitable for resource-constrained embedded devices, has cost and power consumption advantages, and supports a variety of machine learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197660A_ABST
    Figure CN120197660A_ABST
Patent Text Reader

Abstract

The invention discloses an electroencephalogram signal processing-oriented edge inference neural network coprocessor circuit, which supports 16-bit fixed-point number operation and comprises a linear operation unit, a nonlinear operation unit and a convolution operation unit, the linear operation unit performs vector and vector element-by-element summation, vector and vector element-by-element quadrature, vector and scalar element-by-element summation, vector and scalar element-by-element quadrature and vector inner product operation; the nonlinear operation unit is used for realizing e exponent operation, natural logarithm operation and maximum / minimum element searching operation; and the convolution operation unit performs single-channel one-dimensional convolution operation. According to the coprocessor designed for edge reasoning operation oriented to electroencephalogram signal processing, the hardware resource overhead can be greatly reduced on the premise that the performance is guaranteed, acceleration design is carried out for the most basic operation type in machine learning, rich bottom layer operation functions are provided on the software level, and various algorithm types can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to an edge inference neural network coprocessor circuit for electroencephalogram (EEG) signal processing. Background Art

[0002] The electroencephalogram (EEG) signal processing technology has undergone a rapid evolution from traditional algorithms to modern machine learning and deep learning methods in the past few decades. EEG signals are widely used in brain-computer interfaces, neurological disease diagnosis, emotion recognition and other fields due to their high temporal resolution and direct reflection of brain activities. However, due to its low signal-to-noise ratio, strong non-linearity and high-dimensional characteristics, signal processing has always been a complex technical challenge.

[0003] Early EEG processing mainly relied on traditional signal processing methods, focusing on analyzing the time-domain, frequency-domain and time-frequency domain characteristics of signals. For example, spectral features were extracted through Fourier transform, time-frequency information was captured using wavelet transform, or classifiers were constructed through statistical methods. These methods mainly relied on artificial feature design and had a strong dependence on domain knowledge, but were difficult to adapt to complex tasks and high-noise environments. Although these methods had high interpretability, their performance and generalization ability were limited by the feature extraction process.

[0004] With the improvement of computing power and the popularization of data-driven methods, machine learning has gradually become the mainstream technology for EEG signal processing. Through supervised learning models (such as support vector machines, random forests) and unsupervised learning techniques (such as principal component analysis, independent component analysis), machine learning methods can automatically learn features from data and improve the accuracy of classification and prediction. However, traditional machine learning still needs to rely on manual feature extraction and has limitations in processing high-dimensional data and time series correlations.

[0005] In recent years, the rise of deep learning has brought about a revolution in EEG signal processing. Deep learning automatically learns multi-level features from raw EEG signals through models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), etc., reducing the dependence on artificial feature engineering. Combined with advanced architectures such as Transformer, the model can more effectively capture the time series characteristics and global dependencies of EEG, significantly improving the performance of tasks such as classification and prediction. At the same time, the application of transfer learning and generative models further solves the problems of insufficient EEG data samples and individual differences, making EEG signal processing more robust and practical. The core of the above algorithms is vector operation, matrix operation, convolution operation, and non-linear activation function operation.

[0006] In the field of electroencephalogram (EEG) signal processing, some relatively mature existing applications include classification tasks such as sleep staging, epilepsy prediction, and emotion recognition. Deep learning algorithms have demonstrated excellent performance in the applications of these tasks, and the model scales of these algorithms are not large and the complexity is not high. Their operations in the inference stage are suitable for being implemented on edge devices. Therefore, designing a coprocessor for edge inference operations for EEG signal processing has important practical application value. Summary of the Invention

[0007] To solve the above problems existing in the prior art, the present invention provides an edge inference neural network coprocessor circuit for EEG signal processing.

[0008] The technical problems to be solved by the present invention are achieved through the following technical solutions:

[0009] The present invention provides an edge inference neural network coprocessor circuit for EEG signal processing, including an operation acceleration module. The operation acceleration module includes: a linear operation unit, a non-linear operation unit, and a convolution operation unit; the linear operation unit, the non-linear operation unit, and the convolution operation unit all support 16-bit fixed-point number operations;

[0010] The linear operation unit is used to perform vector operations. Among them, the vector operations include: element-by-element summation of vector and vector, element-by-element multiplication of vector and vector, element-by-element summation of vector and scalar, element-by-element multiplication of vector and scalar, and vector inner product operation;

[0011] The non-linear operation unit is used to implement non-linear operations. Among them, the non-linear operations include e exponential operation, natural logarithm operation, and finding the maximum / minimum element operation;

[0012] The convolution operation unit is used to perform single-channel one-dimensional convolution operations.

[0013] Compared with the prior art, the beneficial effects of the present invention are:

[0014] 1) In the coprocessor proposed by the present invention, it includes a linear operation, a non-linear operation, and a convolution operation unit. The linear operation unit solves basic vector and matrix operations. The non-linear operation unit provides operation support for activation functions in machine learning. The convolution operation unit supports convolution operations. Moreover, the convolution operation unit optimizes the redundant reading operations caused by repeatedly reading the convolution kernel and repeatedly reading some data segments. Within a single convolution operation, only the convolution kernel data needs to be read once, and the data sequence to be convolved needs to be read once to complete the convolution operation, greatly improving the operation efficiency. Thus, it solves the problem that in some design solutions of existing coprocessors, the acceleration type is single and insufficient to provide sufficient acceleration support for complex algorithms.

[0015] 2) The coprocessor designed in the present invention focuses on the flexibility of upper-layer software development. Since machine learning algorithms may have performance variations due to individual or environmental differences even in the same application scenario, in order to better adjust the algorithms according to the actual situation, the arithmetic units designed in the coprocessor of the present invention have strong generality and are all the most basic operations at the bottom layer that make up machine learning algorithms. A rich set of basic arithmetic functions are provided at the software level, and upper-layer developers can use these functions to build various complex algorithms, greatly improving the flexibility of the application, facilitating flexible adjustment of the algorithms according to different requirements in actual applications, and having a broader application scope. Thus, it solves the problem in some design schemes of existing coprocessors that the entire algorithm process is solidified in the hardware, making it inconvenient to adjust the algorithm during upper-layer development.

[0016] 3) The coprocessor of the present invention is designed to support 16-bit fixed-point number operations. Using 16-bit fixed-point numbers for operations increases the transmission efficiency of the memory interface and improves the utilization rate of the coprocessor's private memory compared to the commonly used 32-bit floating-point numbers in current machine learning. Therefore, it can support the use of a smaller private memory, with extremely low overhead in hardware resources and low overall power consumption. Moreover, since the logical operation complexity of fixed-point numbers is much lower than that of floating-point numbers, using 16-bit fixed-point numbers for operations can greatly reduce the overall computational complexity of the algorithm, making it very suitable for edge inference tasks for electroencephalogram signal processing. Thus, it solves the problem in some design schemes of existing coprocessors that the private storage used in the coprocessor is large and the arithmetic units are complex. Although they can process various data types and provide high-performance arithmetic acceleration, the hardware resource overhead is large, resulting in performance overkill and energy efficiency waste for edge inference applications for electroencephalogram signal processing.

[0017] The following will further elaborate on the present invention in detail in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings

[0018] Figure 1 is a schematic structural diagram of an edge inference neural network coprocessor circuit for electroencephalogram signal processing provided by an embodiment of the present invention;

[0019] Figure 2 is an architecture diagram of a linear arithmetic unit provided by an embodiment of the present invention;

[0020] Figure 3 is an architecture diagram of a non-linear arithmetic unit provided by an embodiment of the present invention;

[0021] Figure 4 is an architecture diagram of a convolutional arithmetic unit provided by an embodiment of the present invention;

[0022] Figure 5It is an exemplary input-output timing diagram of the shaping module provided by an embodiment of the present invention;

[0023] Figure 6 It is a schematic structural diagram of the CNN-BiGRU shallow hybrid classification network structure model provided by an embodiment of the present invention;

[0024] Figure 7 It is a schematic diagram showing the influence of the fractional bit width on the algorithm accuracy provided by an embodiment of the present invention;

[0025] Figure 8 It is a schematic flowchart of an exemplary process for performing edge inference operations when calling the custom instruction function of the coprocessor provided by an embodiment of the present invention;

[0026] Figure 9 It is the resource occupancy situation of the coprocessor when deployed on Xilinx A7 100T provided by an embodiment of the present invention. Detailed implementation manners

[0027] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0028] Figure 1 It is a schematic structural diagram of an edge inference neural network coprocessor circuit for electroencephalogram signal processing provided by an embodiment of the present invention. As Figure 1 shown, the coprocessor includes an operation acceleration module, and the operation acceleration module includes: a linear operation unit, a non-linear operation unit, and a convolution operation unit.

[0029] All kinds of network layers in machine learning algorithms can be represented by the combination of the most basic linear operations and non-linear operations. Linear operations include matrix operations, vector operations, and scalar operations. The coprocessor mainly completes vector operations. The main reasons are that matrix operations can be fully represented by the combination of vector operations. Therefore, accelerating vector operations can maximize the balance between hardware overhead and acceleration efficiency on the premise of meeting the operation requirements of machine learning. In the coprocessor of the present invention, the linear operation unit mainly completes vector-level operations, such as element-wise summation of vector and vector, element-wise product of vector and vector, element-wise summation of vector and scalar, element-wise product of vector and scalar, and vector inner product operation. Other types of operations are realized by combining the above basic types of operations. High-dimensional matrix operations rely on loop calling vector operation functions at the software level. Specifically, the linear operation unit is used to obtain the vector length and perform element-wise summation of vector and vector, obtain the vector length and perform element-wise product of vector and scalar, obtain the vector length and perform element-wise product of vector and vector, obtain the vector length and perform element-wise summation of vector and scalar, and obtain the vector length and calculate the inner product.

[0030] Exemplarily, Figure 2 is the architecture diagram of the linear operation unit in the coprocessor of the present invention. As Figure 2 shown, the linear operation unit includes: a control module, an operation unit, a cache, and an accumulator. The control module processes the decoding result from the instruction decoding module. The decoding result corresponds to different operation types. The control module is responsible for controlling the data flow of the input data and the output result of the operation logic. The input data has two data flows and can be stored in the cache unit or directly output to the input of the subsequent operation logic. The control module is also responsible for reading data from the cache and outputting it to the subsequent operation unit. The operation unit includes a multiplier and an adder. The operation result is output through a selector, and the selector is controlled by the control module. The input data of the operation unit comes from the cache on the one hand and the current input on the other hand. The cache is used to store partial operands. For example, when implementing the vector A + B operation, the vector A is first stored in the cache. When the data of vector B appears at the input, the data A in the cache is read and output to the subsequent operation unit together with B. The accumulator is used to accumulate the operation results and is mainly used to implement the vector dot product operation. For example, after the operation unit sequentially completes the element-by-element multiplication of vector A and vector B, the result enters the accumulator for accumulation, and finally the dot product result of vector A and vector B can be output.

[0031] The non-linear operation unit mainly accelerates the operations of various activation functions in machine learning, such as ReLU, sigmoid, tanh, etc. Such operations are different from linear operations and cannot be realized by simple multiplication and addition operations at the hardware level. The existing methods for implementing non-linear operations mainly include: look-up tables, piecewise linear approximation, Taylor expansion approximation, CORDIC iterative algorithm, etc. Considering that the non-linear activation functions in machine learning are roughly divided into 2 categories. One category is non-linear functions similar to ReLU, Leaky ReLU, PReLU, etc., which are non-linear functions composed of two piecewise linear functions. Such functions are relatively simple to process and can be calculated using a linear operation module in different intervals. Another category of activation functions such as Sigmoid, Tanh, ELU, Softmax, etc. Their basic non-linear operation parts include e exponential operation and reciprocal operation, and the operation of such non-linear activation functions can be realized by combining with linear operations. The non-linear operation in the coprocessor of the present invention uses a look-up table to implement the e exponential operation. Considering generality, the natural logarithm operation is also added to cooperate with the e exponential operation to implement the exponential and logarithm operations with any base. In addition, it also includes the function of taking the maximum and minimum values. The non-linear operation unit in the coprocessor of the present invention is used to implement the operations of activation functions (such as sigmoid, tanh, σ function, etc.) in machine learning, where the non-linear operations include e exponential operation, natural logarithm operation, and finding the maximum / minimum element operation.

[0032] Exemplarily, Figure 3 is the architecture diagram of the non - linear operation unit in the coprocessor of the present invention. As Figure 3 shown, the non - linear operation unit includes: a control module, a cache, an exponential lookup table, a natural logarithm lookup table, a comparator, and a selector. The control module has the same function as the control module in the linear operation unit. The cache has the same function as the control module in the linear operation unit. The exponential lookup table is a storage unit that stores the output results of the exponential function. That is, when a data is input, the lookup table will output the corresponding exponential operation function result. The natural logarithm lookup table has the same function as the exponential lookup table. When a data is input, it will output the corresponding natural logarithm function operation result. The comparator is used to compare the magnitudes of two input data. The comparator is controlled by the control module and can select to output the maximum value or the minimum value according to different operation types. The selector is used to select which logical function module the final output result comes from and is controlled by the control module.

[0033] The convolution operation unit of the present invention is mainly optimized for convolution operation. The key points of optimization are the caching of the convolution kernel and the processing of the movement of the convolution window. Convolution operation is essentially a linear operation, constantly repeating the dot - product operation between the convolution kernel and the sequence. The convolution kernel does not change within a long period. Although the convolution operation can also be completed using the linear operation module, frequently reading the convolution kernel data at the same address will consume a large number of memory access cycles. After a memory access request is initiated, it takes multiple cycles to receive valid data. Therefore, repeatedly accessing the same data segment will waste a lot of unnecessary cycles. In addition, when performing convolution operations, when the convolution kernel slides backward with a step size, the same data will participate in multiple operation processes. Although the way of incrementing the memory operation address can be used to read the partial data that needs to participate in the operation each time into the coprocessor, this requires the coprocessor to initiate multiple memory access requests, which will also generate a large amount of memory access cycle overhead. For the above two points, the coprocessor of the present invention uses a private RAM to store the convolution kernel for each convolution operation, and only needs to read the convolution kernel once in one convolution operation. In addition, before each convolution operation starts, the integer module in the convolution operation unit outputs the data segment corresponding to the current convolution kernel window according to the step size of the convolution operation and the size of the convolution kernel. In this way, each convolution operation only needs to perform one read operation on the convolution kernel data and the data to be convolved to start the convolution operation, avoiding the redundant read operations caused by the repeated reading of the convolution kernel and the repeated reading of some data to be convolved, and greatly improving the operation efficiency.

[0034] Exemplarily, Figure 4 is the architecture diagram of the convolution operation unit in the coprocessor of the present invention. As Figure 4As shown in the figure, the convolution operation unit includes: a control module, a shaping module, and an operation module. The function of the control module is the same as that of the above-mentioned control module, except that here the input data will be cached in the RAMs of the shaping module and the operation module according to different operation types. The RAM in the shaping module stores the input data of the current convolution operation, and the RAM in the operation module stores the convolution kernel data of the current convolution operation. The shaping module shapes and outputs the continuous input data stored in the RAM. Specifically, since the convolution kernel slides on the input data during the convolution operation, and the step size of each slide is usually smaller than the convolution kernel, after the convolution kernel slides, the data area it covers will partially overlap with that before the slide. The shaping module realizes this operation through a RAM and logical control, that is, the length of the data output by the shaping module each time is the same as the size of the convolution kernel, and the output data is the corresponding data segment after the convolution kernel slides. For example, Figure 5 is a timing diagram of the input and output of the shaping module. For a single-channel one-dimensional convolution operation, the input data is d0~d4, with a length of 5, the convolution kernel size is 3, and the step size is 1. Then, in this convolution operation, the convolution kernel needs to perform dot product operations on three parts of data, that is, the output of the shaping module is d0~d2, d1~d3, and d2~d4. The RAM in the operation module stores the convolution kernel data of the current convolution operation, and the operation unit realizes the dot product operation between the convolution kernel and the output data of the shaping module.

[0035] Continuing to refer to the above Figure 1 , the coprocessor of the present invention further includes an instruction interface and an instruction decoding module. The instruction interface is used to receive instructions sent by the main processor. The instructions include: an operation type code, a destination operand, and a source operand; wherein, the operation type code represents the operation to be performed by the linear operation unit, the non-linear operation unit, or the convolution operation unit, and the destination operand and the source operand respectively represent the address of the destination operand and the address of the source operand of the operation corresponding to the operation type code, or, the destination operand and the source operand respectively represent the operation input information and the operation output information required for the operation corresponding to the operation type code. The instruction decoding module is used to decode the instructions and cache the decoding results, and enable one or more of the linear operation unit, the non-linear operation unit, and the convolution operation unit according to the operation type code obtained by decoding.

[0036] Continuing to refer to the above Figure 1 , the coprocessor of the present invention further includes: a memory access interface and a memory access module. The memory access interface is used to exchange data with an external memory. The instruction decoding module is also used to enable the memory access module according to the operation type code, the destination operand, and the source operand. The memory access module is used to perform corresponding read / write operations on the external memory.

[0037] Continuing to refer to the above Figure 1, the coprocessor of the present invention further includes: a system configuration module. The instruction decoding module is further configured to enable the system configuration module according to the operation type code, the destination operand, and the source operand. The system configuration module includes a variety of registers, and the system configuration module is used to store the destination operand and the source operand, as well as clear the counter value, read the counter value, obtain the source operand address and the destination operand address, block the coprocessor, and reset the coprocessor.

[0038] In the present invention, the instruction interface is mainly responsible for handling the transfer of instructions from the main processor. The instruction contains an operation type code and destination / source operands, etc. The operation type code clearly indicates the type of operation to be performed. The operation type code is divided into two parts. In an N-bit operation type code, the high M bits represent the arithmetic units in the enabled arithmetic acceleration module, and the low N-M bits represent the specific operations to be performed in the enabled arithmetic unit. In the present invention, according to different operation type codes, the destination and source operands also represent different meanings. In the instructions in the initialization stage, the destination and source operands respectively indicate the destination and source operand addresses of the current arithmetic operation. In the instructions in the arithmetic stage, the destination and source operands indicate the arithmetic input / output information of the current arithmetic operation (for example, the arithmetic input information can be the convolution kernel size, stride, bias, and dimensions of the data to be convolved in the convolution operation, and the arithmetic output information can be the dimensions of the obtained convolution result, etc.). When the instruction interface receives these instructions, the instruction decoding module will decode them and enable the corresponding functional modules inside the coprocessor according to the operation type code. Specifically, the instruction decoding module determines the function of the current instruction according to the operation type code field on the instruction interface. For example, when the high M bits are encoded as 0, 1, 2, and 3, the system configuration, linear arithmetic, non-linear arithmetic, and convolution arithmetic modules are enabled respectively, and the low N-M bits encode the operation types inside each enabled module. Exemplarily, the instructions of the coprocessor are shown in Table 1. After decoding is completed, the instruction decoding module will enable the corresponding functional modules to perform corresponding operations. In the present invention, the instruction decoding module decodes the data on the instruction interface and controls the working state of the entire coprocessor according to the decoded information. The main tasks processed by the instruction decoding module include caching the destination and source operand addresses, caching the arithmetic input information and arithmetic output information, and enabling different arithmetic acceleration modules according to the operation type code, enabling the memory access module to read and write to external storage, etc.

[0039] Table 1

[0040]

[0041]

[0042] In the present invention, the design of the memory access interface aims to achieve direct and fast access between the coprocessor and the external memory. Through this interface, the coprocessor can bypass the main processor and directly interact with the memory, greatly improving the efficiency of data reading and writing and reducing the latency of data transmission. The implementation of the memory access interface varies in different instruction set architectures. In some processor architectures, a dedicated coprocessor memory access interface can be provided to share the cache with the main processor, thus directly realizing the function of the coprocessor's autonomous access to the external memory. In some other processor architectures, the coprocessor can interact with the external memory through the system bus and use DMA to implement the direct memory access operation of the coprocessor. When the coprocessor needs to obtain operands from the memory to execute instructions, the memory access interface generates corresponding memory access requests according to the memory address information (the above-mentioned source operand address) provided in the instructions and sends the requests to the external memory through a high-speed data bus. After receiving the requests, the external memory returns the required data to the memory access interface of the coprocessor, and the coprocessor can then quickly use this data for subsequent arithmetic processing. Similarly, after the coprocessor completes certain data processing tasks and needs to store the results back in the memory, the memory access interface can also efficiently write the data to the specified memory address (the above-mentioned destination operand address) to ensure data consistency and integrity. This direct memory access mechanism enables the coprocessor to fully utilize its performance advantages when processing large-scale data and complex algorithms, improving the operating efficiency of the entire system.

[0043] In the present invention, the memory access module realizes the autonomous memory access function of the coprocessor. Through the memory access interface, the coprocessor can actively initiate read and write requests to the external memory to realize the reading of electroencephalogram sensor data and network model parameters, and the write-back of operation results. Specifically, according to the current source and destination operand addresses in the register list cached in the system configuration module, the source operand data is read from the external memory and written into the internal cache, and the operation results output by the operation acceleration module are written back to the destination address of the external memory, thereby realizing the accurate transfer of data between the memory and the coprocessor. In addition, due to a large number of data boundary alignment problems in machine learning algorithms, the memory access module also supports unaligned access to the external memory.

[0044] In the present invention, the system configuration module includes a series of registers for storing the source and target operand addresses of the current operation, as well as some operation information used in the operation acceleration module, such as operation type, data length, fixed value, etc. Specifically, the system configuration module includes general-purpose registers, counting registers, and status registers. The general-purpose registers are used to store some operation information, such as the convolution kernel size, stride, bias, etc. in convolution operations, and are also used to store some intermediate results of operations. The counting register is a counter that continuously counts, facilitating debugging at the software level and calculating the number of cycles of program execution. The status register stores the working state of the coprocessor to indicate which working state the coprocessor is currently in, whether it occupies the memory bus, and whether it can process new coprocessor instructions, etc. The system configuration module plays an important role in the program flow control and exception handling in the coprocessor. Exemplarily, as shown in Table 1 above, the system configuration module is used to store the target operand and source operand, as well as clear the counter value, read the counter value, obtain the source operand address and target operand address, block the coprocessor, reset the coprocessor, etc.

[0045] In the present invention, the linear operation unit, non-linear operation unit, and convolution operation unit all support 16-bit fixed-point number operations. The coprocessor designed in the present invention performed sufficient quantization tests on the algorithm using fixed-point numbers with different bit widths instead of floating-point numbers during the preliminary algorithm testing stage. It was found that the operation error caused by using 16-bit fixed-point numbers (8-bit fractional part) to process electroencephalogram (EEG) signals would not affect the classification result. Therefore, the coprocessor was designed to only support 16-bit fixed-point number operations. Using 16-bit fixed-point numbers for operations increases the transmission efficiency of the memory interface and the utilization rate of the coprocessor's private memory. Specifically, in terms of the operation precision of the coprocessor, quantization processing was performed on the tests of the EEG signal processing algorithm through the software layer. For example, for the sleep staging algorithm, the CNN-BiGRU shallow hybrid classification network structure model was used. Specifically, the network model is as Figure 6 shown. The EEG data first undergoes 3 layers of convolution, and after each convolution, normalization, pooling, and ReLU activation function operations are performed. After three convolutions, it enters the BiGRU layer, and then after normalization and fully connected layers, the classification result is output. The total amount of operations of the entire model is 2.03 MOPs, and the number of network parameters to be trained is 15.51K. Using 32-bit floating-point numbers on Pytorch and performing 20-fold cross-validation on the dataset, the accuracy rate is 84.3%. On this basis, different types of data were used to fully test the algorithm, changing the floating-point numbers to fixed-point numbers and testing the fractional part bit width from 10 bits to 2 bits. Figure 7The effect of the decimal part bit width on the algorithm accuracy is shown. The results show that when there is no overflow in the integer part, the algorithm accuracy will begin to decline when the decimal part bit width is less than 6. When the decimal part bit width is less than 4, the algorithm accuracy is seriously affected. Finally, in order to leave a certain margin, the data type is quantized to a 16-bit fixed-point number, of which the decimal part is 8 bits. Under such quantization accuracy, the accuracy of the algorithm will not be affected.

[0046] In some embodiments, the coprocessor of the present invention can be applied to the following actual edge reasoning scenarios: the network parameters of the deep learning model are stored in the external memory, and the peripheral sensor initiates an interrupt request to the main processor every time it collects data of a certain length, and calls the coprocessor through the coprocessor instruction to perform reasoning operations. After receiving the instruction information, the coprocessor actively accesses the EEG data and network model parameters in the memory through the memory bus, performs reasoning operations and writes the results back to the memory. Such data collection, operation, and write-back processes are completed through DMA and the coprocessor memory access interface. The main processor is only responsible for processing DMA interrupts and coprocessor interface instructions, which greatly saves the time for the processor to move data from the memory and greatly reduces the processor occupancy. By actively completing data access and reasoning operations by the coprocessor, the collaborative work of the main processor and the coprocessor is realized. In other embodiments, there may be no peripheral sensors in the edge reasoning scenario, and the network parameters of the deep learning model and the pre-collected data are stored in the external memory.

[0047] The coprocessor instructions of the present invention construct the lowest-level operation functions based on Table 1. Based on these low-level operation functions, most operations in machine learning can be implemented at the software level. The present invention implements API encapsulation of a large number of basic operation functions based on these custom instructions at the C language level, which facilitates upper-level software development. When calling the coprocessor custom instruction function to perform edge reasoning operations, the flow chart is as follows: Figure 8 As shown. Figure 8 As shown, the process includes:

[0048] S1, configure the initial information of the coprocessor;

[0049] S2, configuring computing information of computing units;

[0050] S3, preload the training parameters of the inference model;

[0051] S4, read the peripheral sensor data and perform corresponding operations;

[0052] S5, feed the scalar result back to the processor / write the vector result back to the memory;

[0053] S6, determine whether the algorithm is finished, if not, return to the above S1 to continue execution; if yes, execute S7;

[0054] S7. Output the inference result.

[0055] The number of multiplier - adders and the size of the cache in the coprocessor designed by the present invention can be set according to actual requirements. The following gives a specific example of the coprocessor designed by the present invention: The data bit - width of the memory access bus of the coprocessor is 64 bits, and the data processed is 16 bits. Therefore, in the linear operation unit, 4 16 - bit multipliers, 4 16 - bit adders, and a 16 - bit accumulator are used, and the cache size is 128B; in the convolution operation unit, the operation unit is the dot - product operation of 4 16 - bit data, which consumes 4 16 - bit multipliers and 4 16 - bit adders, and the total size of two RAMs is 128B; in the non - linear operation unit, the size of the e - exponential lookup table is 256B, the size of the natural logarithm lookup table is 256B, and 4 comparators are also used. Therefore, the operation logic resources used by the coprocessor include 8 multipliers, 8 adders, and 4 comparators, and the total private storage in the module is 1KB. The present invention tests the coprocessor of this example based on the FPGA platform, and the tested algorithm is Figure 6 the model in, and runs the algorithm using the open - source Rocket processor and the coprocessor respectively. The test results at a main frequency of 50MHz are shown in Table 2. The processor needs 31.1M cycles to run the algorithm once, and the coprocessor needs 3.7M cycles, achieving an overall speed - up ratio of 8.37. The resource occupation situation of the coprocessor when deployed on Xilinx A7 100T is as Figure 9 shown, and the occupied LUT resources are 12231.

[0056] Table 2

[0057]

[0058] In summary, for the edge inference operation for electroencephalogram signal processing, the coprocessor designed by the present invention can greatly reduce the hardware resource overhead while ensuring performance, and is more suitable for deployment on resource - constrained embedded devices, having great advantages in terms of cost and power consumption. Moreover, the coprocessor of the present invention is designed to accelerate the most basic operation types in machine learning, and provides rich underlying operation functions at the software level, which can support various algorithm types.

[0059] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0060] In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. Certain measures are recited in mutually different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0061] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An edge inference neural network coprocessor circuit for EEG signal processing, characterized in that: The invention comprises a calculation acceleration module, wherein the calculation acceleration module comprises: a linear calculation unit, a nonlinear calculation unit and a convolution calculation unit; the linear calculation unit, the nonlinear calculation unit and the convolution calculation unit all support 16-bit fixed-point number calculation; The linear operation unit is used to perform vector operations, wherein the vector operations include: element-by-element summation of vectors, element-by-element product of vectors, element-by-element summation of vectors and scalars, element-by-element product of vectors and scalars, and vector inner product operations; The nonlinear operation unit is used to implement nonlinear operations, wherein the nonlinear operations include e-exponential operations, natural logarithm operations, and operations for finding the maximum / minimum element; The convolution operation unit is used to perform a single-channel one-dimensional convolution operation.

2. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 1, characterized in that: The convolution operation unit is used to use private RAM to store the convolution kernel of each convolution operation, read the convolution kernel once in a convolution operation, and before each convolution operation starts, the shaping module in the convolution operation unit outputs a data segment corresponding to the current convolution kernel window according to the step size of the convolution operation and the convolution kernel size.

3. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 1, characterized in that: Also includes: Instruction interface and instruction decoding module; The instruction interface is used to receive an instruction sent by a main processor, and the instruction includes: an operation type code, a target operand, and a source operand; wherein the operation type code represents the operation that the linear operation unit, the nonlinear operation unit, or the convolution operation unit needs to perform, and the target operand and the source operand respectively represent the target operand address and the source operand address of the operation corresponding to the operation type code, or the target operand and the source operand respectively represent the operation input information and the operation output information required for the operation corresponding to the operation type code; The instruction decoding module is used to decode the instruction and cache the decoding result, and enable one or more of the linear operation unit, the nonlinear operation unit and the convolution operation unit according to the operation type code obtained by decoding.

4. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 3, characterized in that: The operation type code has N bits, wherein the upper M bits of the N bits represent enabled operation units, and the lower NM bits of the N bits represent operation operations that the enabled operation units need to perform.

5. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 3, characterized in that: Also includes: Memory access interface and memory access module; The memory access interface is used to exchange data with the external memory; The instruction decoding module is further used to enable the memory access module according to the operation type code, the target operand and the source operand; The memory access module is used to perform corresponding read / write operations on the external memory.

6. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 3, characterized in that: Also includes: System configuration module; The instruction decoding module is further used to enable the system configuration module according to the operation type code, the target operand and the source operand; The system configuration module includes multiple registers, and the system configuration module is used to store the target operand and the source operand, as well as clear the counter count value, read the counter count value, obtain the source operand address and the target operand address, block the coprocessor, and reset the coprocessor.

7. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 6, characterized in that: The system configuration module includes: a general register, a counting register, and a status register; The general register is used to store the convolution kernel size, step size, bias and intermediate results of the convolution operation in the convolution operation; The counting register is a counter for continuously counting; The status register is used to store the working status of the coprocessor to indicate the current working status of the coprocessor.

8. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 1, characterized in that: The convolution operation unit is also used to obtain convolution information and data, and obtain the length of the vector to be convolved and the output length.

9. The edge inference neural network coprocessor circuit for EEG signal processing according to claim 1, characterized in that: The nonlinear operation unit implements the e-exponential operation by means of a lookup table, and adopts natural logarithm operation to cooperate with the e-exponential operation to implement the exponential and logarithmic operation of any base.