Hardware acceleration system for edge device and edge device

By adopting RISC-V cores and VLIW architecture accelerators in edge intelligent hardware systems, the problems of limited processing capabilities, high power consumption and high cost in the existing technology are solved, and efficient and low-power neural network computing is realized, which is suitable for the needs of high-computing scenarios.

CN119940431APending Publication Date: 2025-05-06WUHU RES INST OF XIAN UNIV OF ELECTRONIC SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411927812.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When handling complex AI tasks and large-scale data, existing edge intelligent hardware systems have limited processing capabilities, high power consumption, complex heat dissipation and high cost, making it difficult to meet the needs of high-computing scenarios.

Method used

It adopts the RISC-V core and the accelerator based on the VLIW architecture, and connects the RISC-V core, external memory, camera and accelerator through the system bus to achieve efficient computing and parallel processing of the neural network model.

Benefits of technology

It significantly improves the computing efficiency and performance of the neural network model, can effectively handle complex tensor data operations, reduces system power consumption, and has low hardware costs, making it suitable for deployment in environments with limited power and space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940431A_ABST
    Figure CN119940431A_ABST
Patent Text Reader

Abstract

The invention discloses an edge device-oriented hardware acceleration system and an edge device. The system comprises an RISC-V kernel, an external memory, a camera and an accelerator, the RISC-V kernel is used for sending a VLIW instruction to the accelerator; the VLIW instruction is obtained by performing decomposition and instruction conversion on operation of each network layer of the trained neural network; the accelerator is provided with a parallel processing structure based on a VLIW architecture and is used for configuring a calculation mode according to the calculation requirement of each VLIW instruction and performing parallel calculation of data operation of each VLIW instruction according to the configured calculation mode; the external memory is used for storing parameters of the trained neural network required when the accelerator performs parallel computing; the camera is used for capturing a to-be-processed image of the trained neural network. The system power consumption can be reduced, the parallel processing capability of the system is improved, and the hardware cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of chips and integrated circuits, and in particular relates to a hardware acceleration system and an edge device for edge devices. Background Art

[0002] Cloud computing, with its powerful computing and storage capabilities, supports a variety of applications through various software tools, covering many applications we use in daily life. However, as an extension of cloud computing, edge computing has gradually attracted attention because of its proximity to the device and its rapid response capabilities. Its instant feedback capabilities make it a powerful supplement to cloud computing.

[0003] The future development of the Internet of Things will show two main trends: massive connections and large-scale data. According to a report by IoTAnalytics, the number of global IoT connections will increase by 18% in 2022 to 14.3 billion, and is expected to exceed 29 billion by 2027. The data generated by these devices is huge. If all of it is uploaded to the cloud computing center for processing, it will place extremely high demands on computing power and bandwidth. Edge computing, because it is deployed near the device, can process and filter data in real time, reduce the cloud load, and adapt to the massive connection and data processing needs of the Internet of Things.

[0004] Edge intelligence is a concept that has gradually emerged with the development of the Internet of Things, artificial intelligence, and edge computing technologies. Edge intelligent devices can not only process data, but also make intelligent decisions, making devices more flexible and intelligent, reducing dependence on the cloud, and improving system adaptability and privacy protection capabilities. This evolution not only emphasizes processing speed, but also focuses on system intelligence and learning capabilities.

[0005] A common method of applying artificial intelligence in edge computing is to deploy neural networks on the end side. Since the operation of neural networks requires a lot of repeated calculations, designing hardware accelerators can significantly improve the operating efficiency of neural networks. This invention aims to optimize the edge intelligent system and add a neural network hardware accelerator to significantly improve the end-side artificial intelligence computing power.

[0006] There are currently two main types of hardware systems that can run on the edge and support the deployment of neural networks:

[0007] 1. Embedded systems: ARM Cortex series processors: These low-power processors are commonly used in embedded devices and support basic computing tasks and data processing. NVIDIA Jetson series: Designed specifically for edge AI computing, with powerful GPU performance, suitable for processing complex AI tasks.

[0008] 2. Dedicated accelerators: GPU (Graphics Processing Unit): such as NVIDIA Tesla and H / A series, widely used in edge AI applications that require high-performance computing. FPGA (Field Programmable Gate Array): such as Xilinx and Intel FPGA, used for flexible and efficient computing tasks.

[0009] However, facing the increasing demand for edge intelligent computing power, the existing technical solutions can no longer adapt well to the demand for artificial intelligence computing power. The main shortcomings are as follows:

[0010] 1. Limited processing power: Low-power processors such as the ARM Cortex series are low-cost and easy to deploy, but they have limited performance and weak processing power, and perform poorly in high-computing scenarios. They are unable to cope with complex AI tasks and large-scale data processing.

[0011] 2. Power consumption and heat dissipation: GPUs and high-performance processors, such as NVIDIA Jetson and Intel Xeon D, have strong computing capabilities but high power consumption and require complex heat dissipation solutions, making them unsuitable for deployment in power-constrained and space-constrained environments. Although FPGAs have excellent performance and can flexibly deploy algorithms and models, they also generate significant heat when running at high load for a long time, increasing the complexity of heat dissipation design.

[0012] 3. High cost: GPUs, high-performance processors, FPGAs, and some dedicated edge AI hardware are all expensive and are not suitable for large-scale popularization and deployment. Summary of the invention

[0013] In order to solve the above problems existing in the prior art, the present invention provides a hardware acceleration system and an edge device for edge devices.

[0014] The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0015] The present invention provides a hardware acceleration system for edge devices, comprising: a RISC-V core, an external memory, a camera and an accelerator; wherein the RISC-V core, the external memory, the accelerator and the camera are connected via a system bus;

[0016] The RISC-V core is used to send VLIW instructions to the accelerator; wherein the VLIW instructions are obtained by decomposing and converting the operations of each network layer of the trained neural network;

[0017] The accelerator has a parallel processing structure based on a VLIW architecture, configured to configure a computing mode according to the computing requirements of each VLIW instruction, and perform parallel computing of data operations of each VLIW instruction according to the configured computing mode;

[0018] The external memory is used to store the parameters of the trained neural network required for the accelerator to perform parallel computing;

[0019] The camera is used to capture the image to be processed of the trained neural network.

[0020] The present invention also provides an edge device, comprising the above-mentioned hardware acceleration system.

[0021] Compared with the prior art, the present invention has the following beneficial effects:

[0022] The present invention adopts RISC-V kernel as the control core of the hardware acceleration system, and transmits tensor instructions to the accelerator through the system bus, realizing efficient control and instruction transmission of the accelerator. The present invention adopts an accelerator with a parallel processing structure based on VLIW (very long instruction word) architecture to perform parallel processing of each VLIW instruction, which can process complex tensor data operations (such as transposition, dimension adjustment and data remapping), improves the system's ability to process complex tensor data, improves the system's parallel processing capability, significantly improves the efficiency and performance of neural network model calculations, and can be well adapted to high-computation scenarios; and, by designing an accelerator through the VLIW architecture, the circuit structure can be greatly simplified, thereby effectively reducing the system power consumption, so that the power consumption is low, and no complex heat dissipation solution is required. In addition, the hardware acceleration system provided by the present invention does not require GPU, high-performance processor, FPGA and some dedicated edge AI hardware, and the hardware cost is low.

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a structural diagram of a hardware acceleration system for edge devices provided by an embodiment of the present invention;

[0025] Figure 2 It is a schematic diagram of the deployment process of the neural network model provided by an embodiment of the present invention on a hardware acceleration system;

[0026] Figure 3 It is a schematic diagram of the data operation flow of the neural network model provided by an embodiment of the present invention on a hardware acceleration system. DETAILED DESCRIPTION

[0027] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0028] The present invention provides a hardware acceleration system for edge devices, comprising: a RISC-V core, an external memory, a camera and an accelerator; wherein the RISC-V core, the external memory, the accelerator and the camera are connected via a system bus;

[0029] The RISC-V core is used to send VLIW instructions to the accelerator; VLIW instructions are obtained by decomposing and converting the operations of each network layer of the trained neural network;

[0030] Exemplarily, the VLIW instructions stored in the RISC-V core are VLIW instructions obtained by decomposing the operations of each network layer of the trained neural network, converting the instructions, and optimizing the execution order of the instructions. The specific principles of decomposing the operations of each network layer of the trained neural network, converting the instructions, and optimizing the execution order of the instructions can be found in the following Figure 2 By optimizing the execution order of instructions, we can reduce computational redundancy and model size while ensuring that the specific processing logic of the trained neural network can be accurately implemented, thereby reducing computing time and resource consumption.

[0031] An accelerator having a parallel processing structure based on a VLIW architecture, the accelerator being used to configure a computing mode according to a computing requirement of each VLIW instruction, and to perform parallel computing of data operations of each VLIW instruction according to the configured computing mode;

[0032] The external memory is used to store the parameters of the trained neural network (e.g., weights, biases) required for the accelerator to perform parallel computing;

[0033] The camera is used to capture images to be processed by the trained neural network.

[0034] In some embodiments, the accelerator includes: a scheduling core, an internal cache, a DMA module, and a computing module with multiple computing cores; wherein the scheduling core is connected to the RISC-V core through a bus, and the scheduling core is also connected to the internal cache, the DMA module, and the computing module respectively. Specifically, the scheduling core receives VLIW instructions sent by the RISC-V core through the AXI-Lite interface; the DMA module reads and writes data with the external memory through the AXI interface to ensure high-speed data transmission; the scheduling core transmits signals to the DMA module and the computing module through the RX interface, and transmits data through the DX interface.

[0035] Specifically, the scheduling core is used to receive VLIW instructions sent by the RISC-V core, and to divide each VLIW instruction into scheduling instructions and calculation instructions, and to generate control instructions for the DMA module according to the scheduling instructions to control the DMA module to obtain data from the external memory or camera and return it, and at the same time indicate whether the returned data needs to be stored in the internal cache or sent to the calculation module; and the scheduling core is also used to send the calculation instructions and the data received from the DMA module that needs to be stored and sent to the calculation module to the calculation module. Here, the scheduling core coordinates the DMA module and the internal cache in this way to ensure that the calculation module can quickly access the required input data.

[0036] The multiple computing cores of the computing module constitute a computing cluster. The computing module is used to determine the need to call the corresponding computing core according to the computing instruction, and configure the computing mode of each computing core that needs to be called, and perform parallel computing of the data operation of the computing instruction according to the computing mode of each computing core that needs to be called. Here, each computing core is used to execute a computing instruction in each clock cycle, and the multiple computing cores included in the computing module are used to process multiple computing instructions in parallel. In some embodiments, the computing module has at least 8 computing cores, for example, the computing module has 8 to 32 computing cores, and these computing cores are arranged in a two-dimensional matrix. It should be noted that the number of computing cores included in the computing module can be set according to actual needs, and the present invention does not limit this.

[0037] Specifically, each computing core includes: M scalar ALU units, N vector ALU units, an internal memory corresponding to each scalar ALU unit, and an internal memory corresponding to each vector ALU unit. Exemplarily, each computing core may include 1 scalar ALU unit, 8 vector ALU units, an internal memory corresponding to each scalar ALU unit, and an internal memory corresponding to each vector ALU unit. Each scalar ALU unit is used to perform scalar calculations, and an internal memory corresponding to each scalar ALU unit is used to cache the calculation results of the scalar ALU unit. Each vector ALU unit is used to perform vector calculations, and an internal memory corresponding to each vector ALU unit is used to cache the calculation results of the vector ALU unit. The vector ALU unit and the scalar ALU unit are used together to perform interpolation calculations. Accordingly, the calculation mode of each computing core that needs to be called includes: the working mode of each scalar ALU unit and / or each vector ALU unit in the computing core.

[0038] In the present invention, each VLIW instruction includes an operation type, an operand position, a target ALU and a result storage position, wherein the scheduling instruction obtained by dividing each VLIW instruction includes: an operand position and a result storage position, and the calculation instruction obtained by dividing each VLIW instruction includes: an operation type and a target ALU. The operation type can be at least one of vector calculation, scalar calculation and interpolation calculation. The operand position is used to characterize the address where the data to be calculated is located. The result storage position is used to characterize the address where the calculation result obtained after the calculation is completed is stored. The target ALU is used to characterize the calculation core that needs to be called. It should be noted that the scheduling core can send the calculation instruction divided from a VLIW instruction to each calculation core, and each calculation core determines whether it is called according to the calculation core that needs to be called in the calculation instruction and the ALU that needs to be called in the calculation core that needs to be called, and after determining that it is called, it determines whether to configure the working mode of each scalar ALU unit, or to configure the working mode of each vector ALU unit, or whether the working modes of each scalar ALU unit and each vector ALU unit need to be configured according to the operation type in the calculation instruction. For example, when the operation type in the calculation instruction is vector calculation, the working mode of each vector ALU unit in the calculation core is configured; and when the operation type in the calculation instruction is scalar calculation, the working mode of each scalar ALU unit in the calculation core is configured. For example, each vector ALU unit that needs to be called can be configured to perform a parallel matrix operation mode, and each scalar ALU unit that needs to be called can be configured to perform a single data processing mode. Here, a vector ALU unit refers to an arithmetic logic unit that can process multiple data elements at the same time. These data elements are usually vectors (for example, a digital array), and matrix operations are essentially processing elements in a two-dimensional array. Therefore, through 8 vector ALU units, a calculation core can operate on matrices in parallel. Here, the scalar ALU unit executing a single data processing mode refers to the scalar ALU unit calculating common address calculations, loop counts, etc.

[0039] The internal cache is used to store the data that the DMA module obtains from the external memory or camera and needs to be cached, as well as the intermediate calculation results and data that need to be shared when the computing module performs parallel calculations. Here, by including the internal cache in the hardware acceleration system, the data access speed can be improved.

[0040] In some embodiments, the internal cache is also used to store image data processed by the trained neural network; and the hardware acceleration system for edge devices also includes: a display module. The scheduling core is also used to obtain the processed image data from the internal cache and send the processed image data to the display module for display via the system bus.

[0041] For example, Figure 1 is a schematic diagram of the structure of a hardware acceleration system for edge devices provided by an embodiment of the present invention, such as Figure 1 As shown in FIG. 1 , the computing module in the hardware acceleration system has a computing cluster composed of 16 computing cores, each of which includes a scalar ALU unit, an internal memory 1 corresponding to the scalar ALU unit, 8 vector ALU units, and 8 internal memories 2 corresponding to the 8 vector ALU units. Figure 1 As shown, the RISC-V core transmits tensor instructions to the accelerator through the AXI-Lite interface for accelerated processing, and the DMA in the accelerator obtains acceleration data from the external memory through the AXI interface for subsequent accelerated calculations.

[0042] In some embodiments, the trained neural network model is used as follows Figure 2 The process shown is deployed to the hardware acceleration system for edge devices mentioned above. Figure 2 As shown in Figure 1, the deployment process can be divided into three parts: model compilation and decomposition, instruction and data loading, and accelerator configuration. Figure 2 As shown, model compilation and decomposition may specifically include four steps: acquisition of neural network model, model stratification, operation decomposition and instruction scheduling. In the step of acquiring the neural network model, an arbitrary neural network model that has been trained is acquired, including the network architecture, weights and other parameters. In the step of model stratification, the acquired neural network model is decomposed by level, for example, the convolution layer, pooling layer, fully connected layer, etc. are separated for further processing. In the step of operation decomposition, the operation of each layer obtained by hierarchical decomposition is decomposed, for example, the convolution operation of the convolution layer is decomposed into matrix multiplication and addition operations. Through the step of operation decomposition, the high-level neural network operation can be converted into basic arithmetic and logical operations to adapt to the VLIW instruction format. In the step of instruction scheduling, the basic operation obtained by operation decomposition is converted into VLIW instructions, wherein the VLIW instructions include basic operation instructions, data transmission instructions and control instructions; thereafter, the execution order, repeatability and parallelism between the obtained VLIW instructions are determined to optimize the execution order and execution redundancy of the instructions, thereby improving the computational parallelism and reducing the idle time of computing resources. It should be noted that the model compilation and decomposition part can be performed by the existing compilation system. Figure 2As shown, instruction and data loading may specifically include three steps: instruction loading, model parameter loading, and raw data input. In the instruction loading step, all VLIW instructions after instruction scheduling are sent to the RISC-V core. In the model parameter loading step, the parameters of the neural network model (such as weights and biases) are stored in the external memory. In the raw data input step, the raw input data (such as image data) required by the neural network model is stored in the external memory. Figure 2 As shown, configuring the accelerator may specifically include the following three steps: configuring the scheduling core, setting the calculation mode, and reading the input data. In the step of configuring the scheduling core, the RISC-V core sends the VLIW instructions that need to be loaded this time to the scheduling core, and the scheduling core divides the VLIW instructions received this time into scheduling instructions and calculation instructions. According to the scheduling instructions, the parameters of the neural network model (such as weights, biases) are loaded from the external memory to the internal cache by controlling the DMA module, and the data to be calculated (image) is loaded from the external memory or the camera to the internal cache for subsequent neural network calculations. In the step of setting the calculation mode, the scheduling core sends the scheduling instructions obtained by this division to each calculation core in the calculation module, and the calculation core configures the working mode of each scalar ALU unit and / or each vector ALU unit of itself according to the received scheduling instructions. In the step of reading the input data, the data required to execute the calculation instructions this time is read from the internal cache, and it is ready to enter the calculation stage.

[0043] In some embodiments, the data operation process of the trained neural network model on the hardware acceleration system is as follows: Figure 3As shown. First, data loading and preparation are performed. Specifically, the RISC-V core sends VLIW instructions to the scheduling core through the AXI-Lite interface. The scheduling core receives the instructions and generates control signals according to the instructions and sends them to each execution module to coordinate DMA to read data from the external memory to the internal cache for fast access by the computing core. After that, parallel computing is performed. Specifically, the VLIW architecture allows multiple operations to be performed in one instruction cycle, and each VLIW instruction can control multiple ALU units. The computing tasks of neural networks usually include matrix multiplication, addition, activation functions, etc. These operations can be executed in parallel in hardware using multiple computing cores. The convolution layer is one of the most important computing layers in deep neural networks and is mainly used to extract image features. When calculating the convolution layer, the input image block, weight data, and convolution kernel are read from the internal cache. Since the convolution layer uses two-dimensional parallel computing, the computing core core is arranged in two dimensions to realize two-dimensional convolution calculation in one channel. Multiple computing cores constitute a computing cluster, thereby realizing parallel computing of multiple channels. A part of the ALU units in each computing core is responsible for the calculation of a part of the convolution kernel and the input feature map, and. The scalar ALU unit and the vector ALU unit perform convolution operations in parallel to perform matrix multiplication operations. After obtaining the output of the convolution layer, the output of the convolution layer is stored in the internal cache for use in the next layer of calculation. Since there will be overlapping areas when the convolution kernel slides on the input feature map during the convolution operation, considering the sharing characteristics of the convolution data, the data to be shared is stored in the internal cache, so that redundant calculations and data transmission can be reduced by sharing the convolution kernel data. The pooling layer is used to downsample the feature map to reduce the amount of data and computational complexity. Common pooling operations include maximum pooling and average pooling. When performing the convolution layer calculation, data is read from the output of the convolution layer to the row and column registers in the internal memory corresponding to each ALU unit that needs to be called. The pooling operation is implemented by calculating the data temporarily stored in the row register (Row-Reg) and column register (Col-Reg) in the internal memory corresponding to the ALU unit that needs to be called. When sliding the pooling window, the temporarily stored neurons are transmitted to the adjacent computing core through the local transmission port (TPout), thereby reducing the access to the on-chip cache (internal cache). For multiplication and accumulation paths including maximum pooling and average pooling, the ALU units used for calculation can be appropriately configured according to the specific operation. The vector ALU unit performs the pooling operation, and the data dimension is reduced by comparison or averaging, and the pooling result is stored in the internal cache. The activation function layer is used to introduce nonlinear factors so that the neural network can fit complex models. Commonly used activation functions include ReLU, Sigmoid, and Tanh. The activation function pipeline is divided into five sections. It adopts a hardware implementation based on piecewise linear interpolation. The matrix multiplication result is biased and the activation function is applied, which can efficiently complete the calculation of the activation function.These operations are also completed in parallel in the ALU unit to produce the final output. The calculation results are stored in the internal cache and wait to be transmitted to the next processing module or output device. The normalization layer is used to accelerate network training and improve model stability. Common normalization methods include batch normalization and local response normalization (LRN). For batch normalization calculations, it is achieved by converting into multiplication and addition operations, which utilizes parameter preprocessing and calculation optimization. The normalization operation stores the intermediate calculation results through the local cache and completes the corresponding normalization calculation in the calculation unit, reducing external data transmission and thus improving calculation efficiency. The fully connected layer is usually used in the last few layers of the network to connect all input and output neurons to complete the classification task. The calculation path of the fully connected layer is mainly a multiplication and accumulation operation path. The calculation core reads the output of the pooling layer and realizes the parallel calculation of multiple output neurons through time-sharing multiplexing of the multiplication and accumulation unit. The fully connected layer multiplies the input vector with the weight matrix, where the scalar ALU unit and the vector ALU unit in the calculation core perform this operation in parallel. The calculation results of the fully connected layer are stored in the internal cache and wait to be transmitted to the next layer for processing. The Softmax output layer is usually used for the output of multi-classification neural networks. It can be regarded as a normalized exponential function, which converts an N-dimensional vector containing arbitrary real numbers into another N-dimensional vector, so that the value of each element of the output vector is between 0 and 1, and the sum of all elements is 1. Since the Softmax calculation also involves exponential operations, the piecewise linear interpolation method is also used for approximate calculations. The final calculation result is stored in the internal cache for external output.

[0044] Exemplarily, the operation process of the hardware acceleration system can be as follows: image data and model parameters are stored in the internal cache through the scheduling core and the DMA module; the scheduling core generates a control signal according to the received VLIW instruction and sends it to each execution module respectively, instructing each module to execute in sequence; the ALU unit in each computing core performs parallel calculation according to the instruction, including operations such as convolution, activation, pooling, full connection and normalization; the calculation result is finally classified and output through the Softmax function.

[0045] It can be seen from the above that the hardware acceleration system provided by the present invention has the following characteristics:

[0046] Efficient data transmission design: The RISC-V core is used as the top-level control core of the system, and tensor instructions are transmitted to the accelerator through the AXILite bus to achieve efficient control and instruction transmission; the DMA module is used to achieve efficient data transmission between the accelerator and the external memory to ensure smooth data flow.

[0047] Dedicated module design: Introducing accelerators as dedicated modules to handle complex tensor data operations, such as transposition, dimension adjustment, and data remapping, improves the system's ability to handle complex tensor data; designing hardware threads (computing cores) to process each tensor instruction, realizing superimposed execution of instructions and data transmission, and improving the system's parallel processing capabilities.

[0048] Optimized design of computing tasks and instruction execution: After decomposing the operations of each network layer of the trained neural network, converting instructions, and optimizing the instruction execution order, the generated VLIW instructions are executed. This can reduce computing redundancy and model size while ensuring that the specific processing logic of the trained neural network can be accurately implemented, thereby reducing computing time and resource consumption; through hardware multi-threading mode, each computing core can execute one VLIW instruction per clock cycle, further improving the execution efficiency of the system.

[0049] Low-power design: The accelerator is designed using the VLIW architecture, which greatly simplifies the circuit structure and effectively reduces the system's power consumption. By optimizing the data transmission process, the system's energy consumption is minimized, ensuring the feasibility and sustainability of the design.

[0050] In summary, the hardware acceleration system provided by the present invention realizes efficient computing and real-time processing capabilities of the neural network model by adopting a VLIW architecture, combined with efficient data transmission and flexible computing core configuration; at the same time, through parallel instruction processing, efficient data transmission and storage management, flexible computing core configuration, parallel computing flow, real-time processing capabilities and low-power design, the efficiency and performance of the neural network model calculation are significantly improved, while power consumption and development costs are reduced. It has broad application prospects and significant technical advantages, and is an ideal choice for high-performance embedded systems and edge intelligent real-time applications.

[0051] The present invention also provides an edge device, which includes the above-mentioned hardware acceleration system, so that accelerated calculation of the neural network model can be achieved through the hardware acceleration system.

[0052] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0053] In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good effects.

[0054] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A hardware acceleration system for edge devices, characterized in that: include: RISC-V core, external memory, camera and accelerator; wherein the RISC-V core, the external memory, the accelerator and the camera are connected via a system bus; The RISC-V core is used to send VLIW instructions to the accelerator; wherein the VLIW instructions are obtained by decomposing and converting the operations of each network layer of the trained neural network; The accelerator has a parallel processing structure based on a VLIW architecture, configured to configure a computing mode according to the computing requirements of each VLIW instruction, and perform parallel computing of data operations of each VLIW instruction according to the configured computing mode; The external memory is used to store the parameters of the trained neural network required for the accelerator to perform parallel computing; The camera is used to capture the image to be processed of the trained neural network.

2. The hardware acceleration system for edge devices according to claim 1, characterized in that: The accelerator comprises: a scheduling core, an internal cache, a DMA module and a computing module having multiple computing cores; wherein the scheduling core is connected to the RISC-V core via a bus, and the scheduling core is also connected to the internal cache, the DMA module and the computing module respectively; The scheduling core is used to receive the VLIW instructions sent by the RISC-V core, and divide each VLIW instruction into a scheduling instruction and a calculation instruction, and generate a control instruction of the DMA module according to the scheduling instruction to control the DMA module to obtain data from the external memory or the camera and return it, and at the same time indicate whether the returned data needs to be stored in the internal cache or sent to the calculation module; The scheduling core is further used to send the calculation instruction and the data received from the DMA module and need to be stored and sent to the calculation module to the calculation module; The computing module is used to determine the computing cores that need to be called according to the computing instruction, configure the computing mode of each computing core that needs to be called, and perform parallel computing of the data operations of the computing instruction according to the computing mode of each computing core that needs to be called; The internal cache is used to store the data that the DMA module obtains from the external memory or the camera and needs to be cached, as well as the intermediate calculation results and data that need to be shared when the calculation module performs parallel calculations.

3. The hardware acceleration system for edge devices according to claim 2, characterized in that: The scheduling core receives the VLIW instructions sent by the RISC-V core through the AXI-Lite interface; the DMA module reads and writes data with the external memory through the AXI interface; the scheduling core transmits signals to the DMA module and the computing module through the RX interface and transmits data through the DX interface.

4. The hardware acceleration system for edge devices according to claim 2, characterized in that: Each computing core is used to execute one computing instruction in each clock cycle, and the multiple computing cores are used to process multiple computing instructions in parallel.

5. The hardware acceleration system for edge devices according to claim 2, characterized in that: The computing module has at least 8 computing cores.

6. The hardware acceleration system for edge devices according to claim 2 or 4, characterized in that: Each computing core includes: M scalar ALU units, N vector ALU units, an internal memory corresponding to each scalar ALU unit, and an internal memory corresponding to each vector ALU unit; Each scalar ALU unit is used for performing scalar calculations, and an internal memory corresponding to each scalar ALU unit is used for caching the calculation results of the scalar ALU unit; Each vector ALU unit is used for performing vector calculations, and an internal memory corresponding to each vector ALU unit is used for caching the calculation results of the vector ALU unit; Accordingly, the computing mode of each computing core that needs to be called includes: the working mode of each scalar ALU unit and / or each vector ALU unit in the computing core.

7. The hardware acceleration system for edge devices according to claim 6, characterized in that: M is equal to 1 and N is equal to 8.

8. The hardware acceleration system for edge devices according to claim 2, characterized in that: Each VLIW instruction includes an operation type, an operand position, a target ALU, and a result storage position, wherein the scheduling instruction obtained by dividing each VLIW instruction includes: the operand position and the result storage position, and the computing instruction obtained by dividing each VLIW instruction includes: the operation type and the target ALU; The operation types include: vector calculation, scalar calculation and interpolation calculation; The operand position is used to represent the address where the data to be calculated is located; The target ALU is used to characterize the computing core that needs to be called; The result storage location is used to represent the address where the calculation result obtained after the calculation is completed is stored.

9. The hardware acceleration system for edge devices according to claim 2, characterized in that: The internal cache is also used to store image data processed by the trained neural network; the hardware acceleration system for edge devices also includes: a display module; the scheduling core is also used to obtain the processed image data from the internal cache, and send the processed image data to the display module for display through the system bus.

10. An edge device, characterized in that: The hardware acceleration system comprises any one of claims 1 to 9.