Machine learning heterogeneous multi-model parallel hardware accelerator for mobile equipment
By designing a machine learning heterogeneous multi-model parallel hardware accelerator for mobile devices, the problems of insufficient compatibility when running in parallel with multiple models in the prior art, low memory bandwidth utilization, interference between computing tasks and lack of unified instruction control systems are solved, and efficient multi-task parallel processing and low latency computing effects are achieved.
Patent Information
- Application Number
- CN202510096252.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
When existing hardware accelerators for mobile devices are running in parallel, they face problems such as insufficient compatibility, low memory bandwidth utilization, interference between computing tasks and lack of unified instruction control systems.
Design a machine learning heterogeneous multi-model parallel hardware accelerator for mobile devices, including an input data cache unit, an input parameter cache unit, a multi-channel computing module, an accumulation unit, a nonlinear computing unit and a control unit. Through the design of multi-channel computing module, the parallel operation of multiple machine learning models is supported, which avoids the mutual influence between models and optimizes the data loading and computing process through alternating operation.
The efficient and low latency of multi-tasking parallel processing on mobile devices is achieved, avoiding mutual influence between models, improving acceleration efficiency and reducing power consumption, while improving the flexibility and scalability of the accelerator.
Smart Images

Figure CN119940435A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hardware design of integrated circuit artificial intelligence, and specifically relates to a hardware accelerator for machine learning tasks, which can achieve multi-model parallel acceleration for mobile devices and has both high efficiency and flexibility. Background Art
[0002] In recent years, with the rapid development of deep learning and machine learning technologies, various machine learning algorithms have been widely used in many fields such as speech recognition, image processing, natural language processing, etc. These applications have put forward higher requirements on computing performance and promoted the research and development of dedicated hardware accelerators.
[0003] At present, mainstream hardware accelerator designs usually focus on optimizing a single type of machine learning model. For example, Google's TPU (Tensor Processing Unit) is designed to accelerate matrix operations in neural networks, while NVIDIA's GPU (Graphics Processing Unit) has become one of the mainstream hardware platforms for deep learning training and reasoning with its powerful parallel computing capabilities. In addition, platforms such as Facebook's Glow and Xilinx's DPU optimize the execution efficiency of specific tasks through customized hardware architectures. However, with the increasing complexity of machine learning tasks in practical applications, the optimization of a single model can no longer meet the needs. In many scenarios, such as edge computing devices or mobile applications, the system needs to run multiple machine learning models at the same time. These models may each have different computing characteristics, such as k-Means clustering requires distance calculations, DNN requires matrix dot product operations, and logistic regression (LR) relies on gradient descent methods. How to efficiently support the parallel operation of multiple models on a single hardware platform has become an important research topic.
[0004] Existing mobile-oriented hardware accelerators face the following challenges when dealing with multi-model parallel execution: (1) Insufficient compatibility between different models: Various machine learning models have their own unique computing requirements at runtime. For example, DNN models rely on large-scale matrix multiplication operations, KNN requires a large number of distance calculations, and logistic regression (LR) focuses on gradient descent optimization. In existing hardware accelerators, the architecture design is often optimized for the computing characteristics of a certain type of model, while supporting other models is weak.
[0005] (2) Low memory bandwidth utilization: When multiple models are running in parallel, the loading of input data and parameters is frequent and complex, and memory bandwidth often becomes a bottleneck for system performance. Many existing hardware designs fail to fully optimize memory access strategies, resulting in bandwidth waste and data loading delays.
[0006] (3) Calculate the interference between tasks: In multi-tasking scenarios, computing tasks of different models may interfere with each other, such as contention for shared resources, reduction in cache hit rate, etc. These problems further reduce the performance of hardware accelerators.
[0007] (4) Lack of a unified command and control system: For hardware platforms that run multiple models in parallel, a flexible instruction control system needs to be designed to coordinate the running order, data flow, and resource allocation of multiple models. However, many current hardware accelerators fail to provide sufficient versatility and scalability in instruction set design.
[0008] From the perspective of the industry background, the research on multi-model hardware accelerators is of great significance. On the one hand, with the popularization of edge computing devices and the increase in computing power requirements of smartphones, multitasking has become a rigid demand. For example, in the field of autonomous driving, vehicles need to run multiple models such as target detection, path planning, and semantic segmentation at the same time; in intelligent medical diagnosis, equipment may need to process different types of image and data analysis tasks at the same time. On the other hand, in the field of academic research, the design of multi-model hardware accelerators provides a new direction for the exploration of efficient computing architectures, which has high theoretical value and practical significance.
[0009] In summary, the existing hardware accelerator design for mobile devices cannot fully meet the needs of multi-model parallel processing. Summary of the invention
[0010] The purpose of the present invention is to propose a hardware architecture that supports multi-task parallel processing and optimizes resource utilization, so as to achieve multi-task parallel operation while ensuring low latency under the condition of limited resources of mobile devices.
[0011] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: A machine learning heterogeneous multi-model parallel hardware accelerator for mobile devices, comprising: an input data cache unit, an input parameter cache unit, an output cache unit, a multi-channel computing module, an accumulation unit, a nonlinear operation unit and a control unit; The input data cache unit is used to store the target data to be accelerated by the machine learning model. The target data is loaded from the memory by DMA. The data is divided into blocks inside the unit. The amount of data stored in each basic storage unit SE is equal to the amount required by the basic processing unit PE. The input parameter cache unit is used to store the parameter data of the machine learning model. The parameter data is loaded from the memory by DMA. The data is divided into blocks inside the unit. The amount of data stored in each basic storage unit SE is equal to the amount required by the basic processing unit PE. The multi-channel computing module is a key design for the parallel operation of multiple models, including several acceleration channels, each of which is composed of several basic processing units PE. Different machine learning models run in parallel in different acceleration channels. The basic processing unit PE supports the accelerated operation of different machine learning models, so the channel acceleration also supports a variety of different machine learning models. And because different machine learning model data operations are in different acceleration channels, the mutual influence between different types of models is avoided. The multi-channel design can be expanded to support more parallel tasks and can adapt to more complex computing needs. The operation mode of the acceleration channel is that several PEs use alternating operation mode to calculate the data from the input data cache unit and the input parameter cache unit; The basic processing unit PE includes: a plurality of basic computing units CU, a multiplexer and an adder tree; the plurality of PEs are operated in an alternating operation acceleration method to form a computing mode of an acceleration channel; wherein: The basic computing unit CU includes: division operation, comparison operation, multiplication operation and addition and subtraction operation, and is used to calculate the input data of the PE, wherein one of the input data of the PE comes from the input data cache unit and the other comes from the input parameter cache unit; The multiplexer, when the instruction requires multiple operation combinations, performs calculation result reflux through the multiplexer and performs secondary operation; including: in the dot product operation of LR and DNN, multiplication operation is performed first and then addition operation is performed; in the distance calculation of KNN and k-Means, subtraction operation is performed first and then multiplication operation is performed; in the convolution operation of CNN, multiplication operation is performed first and then addition operation is performed; The adder tree accumulates and sums the calculation results of multiple CUs; The control unit is used to parse the instruction data into corresponding circuit control signals, complete the data reading from the input data cache unit and the input parameter cache unit, flow into the calculation of the multi-channel calculation module, and then flow into the calculation of the accumulation unit and the nonlinear operation unit, and finally output the calculated data to the output cache unit; The accumulating unit is used to store the partial sum of the data in the input data cache unit and the data in the input parameter cache unit after being calculated by the multi-channel operation module; The nonlinear operation unit is used to perform nonlinear operations of ReLU and Softmax on the data output from the accumulator unit, and calculate non-tensor operations in the machine learning model, including activation layer, pooling layer and softmax layer operations in CNN; The output buffer unit is used to store data calculated by the accumulation unit and the nonlinear operation unit.
[0012] The calculation is performed by several PEs in an alternating operation mode, specifically including: The PE is a basic processing unit, which is used to accelerate different types of machine learning models. All data required by the PE is provided by the SE. There are multiple PEs and SEs in an acceleration channel, where the number of PEs and SEs is the same. The corresponding relationship is that the SE provides data for the PE operation, and the SE and PE correspond one to one; The SE is a basic storage unit, which means that the input data cache unit and the input parameter cache unit divide the data into blocks and input them to the PE for operation. The amount of input data that the PE can process is the amount of data stored in the SE. The alternating operation means that the operation order of multiple PEs in an acceleration channel is sequential. After the previous PE completes the calculation, the next PE will start the calculation immediately. After the last PE completes the calculation, the first PE starts the calculation again. To implement this mechanism, it is necessary to ensure that the SE loads the data before the PE calculates. If there are n PEs in an acceleration channel, the calculation time of the PE is t1, and the loading time of the SE is k1, then k1 <= n × t1 must be satisfied, that is, continuous and uninterrupted calculation of the computing units in the acceleration channel can be achieved.
[0013] The subtlety of the design of the present invention lies in the use of space for time. By increasing the number of PEs in the acceleration channel, each PE can be guaranteed to complete the loading of its corresponding SE before the next round of calculation. From the perspective of the acceleration channel as a whole, the time-consuming simulation time can be saved, and the sequential and seamless calculation of each PE can be achieved. The multi-channel acceleration structure designed by the present invention avoids the mutual influence of multiple heterogeneous machine learning models during operation. At the same time, the channel design structure can improve the acceleration efficiency of the model and reduce power consumption. The instruction set designed by the present invention that matches the accelerator structure improves the flexibility and scalability of the accelerator. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a structural diagram of the accelerator operation state of the present invention; Figure 2 It is a block structure diagram of the input data buffer unit of the present invention; Figure 3 It is a block structure diagram of the input parameter buffer unit of the present invention; Figure 4 It is the design diagram of the basic processing unit PE of the present invention; Figure 5 is an example diagram of the alternating acceleration strategy of the present invention; Figure 6 It is the instruction structure diagram of the present invention. DETAILED DESCRIPTION
[0015] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0016] like Figure 1 As shown, a machine learning heterogeneous multi-model parallel hardware accelerator for mobile devices of the present invention includes: an input data cache unit, an input parameter cache unit, an output cache unit, a multi-channel computing module, an accumulation unit, a nonlinear operation unit and a control unit; like Figure 2 As shown, the data stored in the input data cache unit will be divided into multiple blocks to ensure that the smallest data block, that is, the input data dimensions required by the basic storage unit SE and PE are the same to support the operation of the acceleration channel; like Figure 3 As shown, the data stored in the input parameter cache unit will also be divided into multiple blocks to ensure that the smallest data block, that is, the input parameter dimensions required by the basic storage units SE and PE are the same to support the operation of the acceleration channel; When the machine learning model is a KNN or K-means model, the input data cache unit stores the target sample data, and the input parameter cache unit stores the reference sample data. When the machine learning model is a DNN or LR model, the input data cache unit stores the target data, and the input parameter cache unit stores the model weight parameter data; like Figure 4 As shown, the interior of the basic processing unit PE includes: multiple basic computing units CU, a multiplexer MUX and multiple adders; multiple PEs operate according to the alternating operation acceleration method to form a computing mode of the acceleration channel.
[0017] The basic computing unit CU includes: division operation, comparison operation, multiplication operation and addition and subtraction operation, and is used to calculate the input data of the PE, wherein one of the input data of the PE comes from the input data cache unit and the other comes from the input parameter cache unit; The multiplexer, when the instruction requires multiple operation combinations, performs calculation result reflux through the multiplexer and performs secondary operation; including: in the dot product operation of LR and DNN, multiplication operation is performed first and then addition operation is performed; in the distance calculation of KNN and k-Means, subtraction operation is performed first and then multiplication operation is performed; in the convolution operation of CNN, multiplication operation is performed first and then addition operation is performed; The adder tree accumulates and sums the calculation results of multiple CUs; like Figure 5The figure shows the process of alternating operation. First, (1) indicates that the SE data required by all PEs in the acceleration channel have been loaded. (2) indicates that PE1 starts to calculate the data corresponding to SE1. (3) indicates that PE2 starts to calculate immediately after PE1 calculates. At the same time, SE1 starts to load some data. (4) indicates that PE3 calculates immediately after PE2. At the same time, SE1 and SE2 load data. (5) indicates that SE1's data loading is completed. SE2 and SE3 continue to load data. (6) indicates that PE1 starts to calculate the result of SE1's reloaded data. At the same time, SE2 has completed loading and SE3 is continuing to load. Then PE2 seamlessly connects to the calculation after PE1, thus forming a loop. From a macro perspective, from PE1 to PE3, there is a continuous loop calculation. There is no time to load the SE corresponding to PE in the middle. Therefore, for the acceleration channel as a whole, the acceleration is achieved by saving the time of loading data. For some DNN models with huge parameter quantities, data loading occupies most of the overall time. Usually, the time consumed by simulation is dozens or even hundreds of times the calculation speed.
[0018] like Figure 6 The instruction structure table is shown, including: 4*CM, InputBuf, ParaBuf, ACC, NonLine, OutputBuf and PE components. CM is used to control the basic operators executed by the instruction, including dot product, distance calculation, counting, sorting, non-linear function and convolution, a total of 6 basic operator operations, InputBuf is used to control the loading of input data cache unit data, ParaBug is used to control the loading of input parameter cache unit data, ACC is used to control the execution of accumulation, when ACC completes the accumulation, it will enter the nonlinear operation unit NonLine for calculation, and finally enter the OutputBuf to save part and or output. The PE part is used to control the operation of all PEs in the acceleration channel, and the operations performed by PEs in the same channel are the same. Example
[0019] In this embodiment, the hardware design that can support up to 4 different models (DNN, LR, KNN, K-means, CNN) to run and accelerate at the same time is as follows: Four acceleration channels and a corresponding number of accumulation units are designed to ensure that the accumulation results of each acceleration channel are independent of each other. At the same time, the input data cache unit size is 8KB, the input parameter cache unit size is 8KB, and the number of computing units CU in a basic acceleration unit PE is 16, which can accept 16-dimensional input data for processing. Synopsys VCS is used to simulate the hardware design with a simulation accuracy of 1ns. Verilog is used as the design language, and the design library uses the TSMC 130nm process library. Under the 130nm process, the area is about 9.5mm², which is about 3mm in length and width. The overall power consumption is only 294mW, which is suitable for the small area and low power consumption characteristics of mobile devices. The relevant information about the proportion of each component of area and power consumption is shown in Table 1.
[0020] Table 1
[0021] At the same time, this embodiment realizes the simulation and running of four models of DNN, LR, KNN, and k-Means on the accelerator. The data set uses MNIST handwritten digit pictures, where the DNN model is composed of 4 fully connected layers, corresponding to the dimensions of 784, 1024, 512, and 10, respectively. There are 640 test instances, 784-dimensional feature input, and 10-dimensional output (a total of 10 classification results 0-9). There are 42,000 test instances of the LR logistic regression model, 784-dimensional feature input, and 10-dimensional output. There are 640 test cases for the KNN model, 160 reference cases, and K is set to 20. There are 640 test cases for the k-Means model, 784-dimensional feature input, and k is set to 10.
[0022] At the same time, in order to perform the inference step directly on the accelerator, DNN and LR were pre-trained in advance using Python, and the parameters of all input layers and hidden layers of DNN and LR were saved in .hex files. When the Synopsys VCS simulation model was run on the accelerator, the model parameters and image data were saved in the memory in advance. The simulation run needs to simulate the process of reading data from the memory to the input data cache unit and the input parameter cache unit, so it is achieved by setting the read data delay.
[0023] The execution process of an instruction first reads the CM part to determine the operator executed by this instruction, and then reads the data to be calculated into the cache unit, such as Figure 1In step 1, data will first be loaded from each model in the memory to the input data cache unit and the input parameter cache unit. In this embodiment, the network weight parameters of DNN and LR will be loaded into the input parameter cache unit, and the reference sample data of KNN and k-Means will be loaded into the input parameter cache unit. DNN, LR, KNN and k-Means will load handwritten digital image data into the input data cache unit. After the first round of reading the data, all PEs in the acceleration channel are calculated in sequence according to the control signal, and each PE corresponding to the data is loaded again. The cycle is repeated until the number of read iter iterations in the instruction is reached, and then the data is output to the ACC accumulation part. In this embodiment, taking DNN as an example, its input layer is 784 dimensions and the first layer is 1024 dimensions, and a PE can only process 16-dimensional data. Therefore, even after reaching the number of calculated iterations, the data output by the PE is still a partial sum, and 64 calculations are still required to completely calculate all the data in the first layer. After the ACC accumulation part is calculated, the data enters the NonLine nonlinear calculation unit, DNN and LR will perform softmax operations, KNN and k-Mmeans will perform sorting operations. After several cycles from reading data to PE calculation, to ACC accumulation and then to NonLine nonlinear calculation, the final result is output to OutputBuf, which is temporarily saved as a partial result or directly output.
Claims
1. A machine learning heterogeneous multi-model parallel hardware accelerator for mobile devices, characterized in that: include: An input data cache unit, an input parameter cache unit, an output cache unit, a multi-channel computing module, an accumulation unit, a nonlinear operation unit and a control unit; The input data cache unit is used to store the target data to be accelerated by the machine learning model. The target data is loaded from the memory by DMA. The data is divided into blocks inside the unit. The amount of data stored in each basic storage unit SE is equal to the amount required by the basic processing unit PE. The input parameter cache unit is used to store the parameter data of the machine learning model. The parameter data is loaded from the memory by DMA. The data is divided into blocks inside the unit. The amount of data stored in each basic storage unit SE is equal to the amount required by the basic processing unit PE. The multi-channel computing module includes several acceleration channels, each of which is composed of several basic processing units PE. Different machine learning models run in parallel in different acceleration channels. The basic processing unit PE supports the accelerated operation of different machine learning models. The operation mode of the acceleration channel is that several PEs use an alternating operation mode to calculate data from the input data cache unit and the input parameter cache unit. The basic processing unit PE includes: a plurality of basic computing units CU, a multiplexer and an adder tree; wherein: The basic computing unit CU includes: division operation, comparison operation, multiplication operation and addition and subtraction operation, and is used to calculate the input data of the PE, wherein one of the input data of the PE comes from the input data cache unit and the other comes from the input parameter cache unit; The multiplexer, when the instruction requires multiple operation combinations, performs calculation result reflux through the multiplexer and performs secondary operation; including: in the dot product operation of LR and DNN, multiplication operation is performed first and then addition operation is performed; in the distance calculation of KNN and k-Means, subtraction operation is performed first and then multiplication operation is performed; in the convolution operation of CNN, multiplication operation is performed first and then addition operation is performed; The adder tree accumulates and sums the calculation results of multiple CUs; The control unit is used to parse the instruction data into corresponding circuit control signals, complete the data reading from the input data cache unit and the input parameter cache unit, flow into the calculation of the multi-channel calculation module, and then flow into the calculation of the accumulation unit and the nonlinear operation unit, and finally output the calculated data to the output cache unit; The accumulating unit is used to store the partial sum of the data in the input data cache unit and the data in the input parameter cache unit after being calculated by the multi-channel operation module; The nonlinear operation unit is used to perform nonlinear operations of ReLU and Softmax on the data output from the accumulator unit, and calculate non-tensor operations in the machine learning model, including activation layer, pooling layer and softmax layer operations in CNN; The output buffer unit is used to store data calculated by the accumulation unit and the nonlinear operation unit.
2. The machine learning heterogeneous multi-model parallel hardware accelerator according to claim 1, characterized in that: The calculation is performed by several PEs in an alternating operation mode, specifically including: The PE is a basic processing unit, which is used to accelerate different types of machine learning models. All data required by the PE is provided by the SE. There are multiple PEs and SEs in an acceleration channel, where the number of PEs and SEs is the same. The corresponding relationship is that the SE provides data for the PE operation, and the SE and PE correspond one to one; The SE is a basic storage unit, which means that the input data cache unit and the input parameter cache unit divide the data into blocks and input them to the PE for operation. The amount of input data that the PE can process is the amount of data stored in the SE. The alternating operation means that the operation order of multiple PEs in an acceleration channel is sequential. After the previous PE completes the calculation, the next PE will start the calculation immediately. After the last PE completes the calculation, the first PE starts the calculation again. To implement this mechanism, it is necessary to ensure that the SE loads the data before the PE calculates. If there are n PEs in an acceleration channel, the calculation time of the PE is t1, and the loading time of the SE is k1, then k1 <= n × t1 must be satisfied, that is, continuous and uninterrupted calculation of the computing units in the acceleration channel can be achieved.