Cyclic partitioning and resource allocation method for reducing CPU empty equations of convolutional neural network of embedded device
By adopting a circular iterative chunking method and prefetching and memory access strategy on embedded devices, and exploring hardware allocation schemes combined with reinforcement learning, the problem of CPU space in convolutional neural networks is solved, and more efficient resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202510209156.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively reduce the CPU empty time of convolutional neural networks on embedded devices, especially in convolution and full connection operations, where there are a large number of data reuse problems.
A circular iterative chunking method is adopted, combined with prefetch and memory access strategy, the data locality of the convolutional neural network is optimized, and reinforcement learning is used to explore the best hardware allocation scheme to reduce CPU empty time.
It significantly reduces the overhead of read, write and memory access time, maximizes the reduction of CPU empty time, rationally utilizes the limited hardware resources of embedded devices, and improves the universality of the circular blocking method.
Smart Images

Figure CN119988033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of loop optimization and reinforcement learning, and in particular to a loop block division and resource allocation method for reducing CPU idle time of a convolutional neural network of an embedded device. Background Art
[0002] In today's booming computer technology, loop blocking is a typical key technology for enhancing data locality. It has made great achievements in reducing the communication between main memory and CPU, greatly improving the efficiency of system operation. The block size has a significant impact on the degree of performance improvement, so many works focus on the precise determination of the block size and propose a series of sophisticated optimization methods.
[0003] There is a lot of work in the area of loop chunking. Some researchers have proposed a clever use of sine and cosine functions to express a class of complex nonlinear optimization problems clearly and concisely; in addition, some researchers have proposed a constant-time algorithm to design the optimal chunk size specifically for doubly nested loops. These works mainly focus on designing the optimal loop chunking scheme, but there is a lot of data reuse in the basic operations of convolutional neural networks, which is a factor that must be taken into account. Write-mode-aware loop chunking (WMALT) makes a trade-off between write performance and write retention time, reducing the life cycle of write instances and maximizing the number of efficient and fast writes in the loop. The square chunking scheme (WET) determines the maximum chunk size based on the cache capacity, ensuring that writes from the cache to the main memory are only performed when the current chunk is executed.
[0004] At present, the application of convolutional neural networks is becoming more and more extensive, and the amount of data involved is growing explosively. If you want to use convolutional neural network models on embedded devices, you need to consider how to reasonably apply limited hardware resources. At the same time, although various studies have made progress in solving different problems, no research can minimize the CPU idle time for all basic operations of convolutional neural networks. Moreover, in these basic operations, especially in convolution operations and fully connected operations, although there is no data dependency, there is a lot of data reuse, so people should design a compilation-level, lower-level loop block scheme for convolutional neural networks running on embedded devices.
[0005] Reinforcement learning is a popular method in machine learning. It is good at solving decision-making problems and has shown extraordinary application value in many cutting-edge fields, such as the precise control process of robots, the performance of intelligent characters in games, and the safe driving of self-driving cars. At the same time, reinforcement learning is also good at solving multi-objective optimization problems. It can cleverly balance various goals so that the final solution fits the Pareto frontier and achieves the optimal allocation of resources. This technology is also flexible and can handle both discrete and continuous states.
[0006] In the operation mechanism of reinforcement learning, the agent observes the state of the environment, decides the next action based on the observation results, and then receives positive and negative rewards from the environment as feedback. Based on the feedback and past experience, the agent will continuously adjust its behavior strategy.
[0007] The combination of reinforcement learning and recurrent blocking also has precedents. In the actual operation process, we first use a new model to fully explore the suitable tile size. At the same time, with the help of the powerful tool of analyzing boundaries, we strictly limit the search space to avoid blind exploration, and then make fine adjustments to the square and non-square blocks based on past experience to ensure the scientificity and effectiveness of the plan in all aspects. Summary of the invention
[0008] The purpose of the present invention is to provide a loop iterative block method for improving data locality for the five basic operations of convolutional neural networks, and adopt a pre-fetch memory access strategy to realize computational hidden memory access, which is targeted and effective, so that the designed loop iterative block can significantly reduce the time overhead of reading and writing memory access, and minimize the CPU idle time. At the same time, the present invention uses reinforcement learning to explore the architecture, reasonably selects the hardware allocation plan according to the hardware configuration of the embedded device, and reasonably utilizes the limited hardware resources on the embedded device, thereby improving the universality of this loop block method. The present invention has carried out targeted design for the basic operations of convolutional neural networks, and has better solved the problems of inconsistent implementation of each operation cycle and different loop characteristics, especially on embedded devices with limited hardware conditions, it can reasonably utilize hardware resources and minimize the CPU idle time without affecting other processes, and has good application prospects.
[0009] The specific technical solution for achieving the purpose of the present invention is: A loop block and resource allocation method for reducing the CPU idle time of a convolutional neural network in an embedded device. Under given hardware conditions, the optimal loop block size is determined according to the loop characteristics of each basic operation of the convolutional neural network itself to obtain the minimum CPU idle time; at the same time, reinforcement learning is used to explore the embedded device resource allocation scheme to achieve the best balance between energy consumption, cache utilization and CPU idle time. The method includes the following specific steps: S1: Based on the cache capacity, number of computing cores and learning rewards of the current hardware device, the reinforcement learning controller selects the hardware allocation; S2: Get the type of the current convolutional neural network operation and the corresponding input data size; the convolutional neural network takes the image as input, and gets the size of the image input each time; S3: Analyze the characteristics of the loop and data reuse properties of the current operation type; S4: According to the characteristics of the loops of each operation, the data reuse characteristics and the hardware allocation, the block size constraint conditions are proposed, and the optimal block size that can minimize the CPU waiting time is obtained under the constraint conditions; the pre-fetch memory access strategy is adopted to make the memory access and calculation proceed in parallel to form a pipeline, and the required data is taken out in advance to calculate the hidden memory access, so as to further reduce the CPU waiting time; S5: Calculate the reward value based on the cache utilization, CPU idle time and energy consumption when the optimal block size is used, and feed it back to the reinforcement learning controller for learning; repeat steps S1-S4, and the reinforcement learning controller continuously selects and learns until the number of selections reaches the specified value, and the reinforcement learning controller selects the best resource allocation.
[0010] Furthermore, the step S3 of analyzing the characteristics of the loop corresponding to the current operation type and the data reuse characteristics specifically includes: S3-1: The convolution operation is implemented by a four-dimensional loop. The innermost two layers iteratively process the data in a window, and the outermost two-dimensional loop iteratively processes all windows of the input data. The innermost two layers of loops are expanded to obtain a two-dimensional loop, and each iteration processes a window. At this time, it is observed that there is a lot of data reuse in both the horizontal and vertical directions; S3-2: The pooling operation cuts the input image into windows of the same shape and size, performs the same operation on each window, and finally obtains the pooling result; the pooling operation is implemented as a four-dimensional loop, the innermost two layers traverse each data in the window, and the outermost two layers traverse each window in the input image; there is no data reuse in this loop, but there is room for improvement in the effective data load; S3-3: The activation operation performs a nonlinear transformation on each pixel of the input image, which is implemented by a two-dimensional loop. There is no data reuse in this loop, but there is room for improvement in the effective data load. S3-4: Batch normalization performs batch normalization on each pixel of the input image, which is implemented by a two-dimensional loop. There is no data reuse in this loop, but there is room for improvement in the effective data load. S3-5: The full connection operation is the weighted addition of the one-dimensional input vector and the weight matrix, which is implemented by a two-dimensional loop; the inner loop traverses the input vector, and the two-layer loop traverses the weight matrix. The multiplication of a vector element and a matrix element is an iteration; the inner loop is executed to obtain a result, and the input vector must be completely fetched into the cache during the execution process; the outer loop is executed to obtain the result matrix, and the weight matrix must be completely fetched into the cache during the execution process. The data requirement order of the weight matrix is inconsistent with the storage order. In order to use part of the data in the cache line, the same cache line will be repeatedly fetched into the cache, and the memory access overhead is extremely large; the data reuse degree of the input vector and the weight matrix is extremely high.
[0011] Further, the step S4 specifically includes: S4-1: According to the characteristics of the loop of each operation, data reuse and hardware allocation, the block size constraint is proposed; S4-2: Obtain a set of block size solutions according to the block size constraint. Each solution divides the iterations distributed in the two-dimensional space into rectangular areas, i.e., blocks. Data is read, processed, and results are written back in units of blocks. S4-3: A pre-fetch memory access strategy is adopted to enable calculation and memory access to be performed in parallel, so as to calculate hidden memory access and fetch the subsequent required data in advance: while calculating in one block, the memory access core first writes back the calculation result generated by the previous block, and then fetches the data required for the next block; in a multi-core system, one CPU core is dedicated to memory access, and other CPU cores jointly complete the calculation.
[0012] S4-4: Calculate the CPU idle waiting time for each solution in the block size solution set, compare and obtain the minimum CPU idle waiting time, and record the block size (i.e., the optimal block size) corresponding to the minimum CPU idle waiting time.
[0013] Further, the step S1 specifically includes: S1-1: Generate the search space of the reinforcement learning controller based on the cache capacity and number of computing cores of the current hardware device; S1-2: If it is the first selection, the reinforcement learning controller randomly selects a resource allocation; otherwise, the reinforcement learning controller selects a resource allocation in the search space based on the reward received from the previous selection.
[0014] Further, the step S5 specifically includes: S5-1: Calculate the reward value according to the cache utilization, CPU idle time and energy consumption at the optimal block size, and feed the reward value back to the reinforcement learning controller; S5-2: The reinforcement learning controller learns the reward value, and the learning result affects the subsequent selection of the reinforcement learning controller; when the number of selections of the reinforcement learning controller is less than the specified value, it returns to S1; when the number of selections is equal to the specified value, the reinforcement learning controller selects the best resource allocation.
[0015] Compared with the prior art, the present invention is more targeted, analyzing the loop characteristics of each operation of the convolutional neural network and designing a loop blocking scheme; it is also universal and can explore the optimal hardware allocation scheme on a given embedded device. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a flow chart of an embodiment of the present invention; Figure 2 This is a schematic diagram of the pipeline of the pre-fetch memory access strategy of the present invention; Figure 3 This is a flowchart of the reinforcement learning controller of the present invention. DETAILED DESCRIPTION
[0017] The present invention is described in detail below with reference to the accompanying drawings and embodiments. Example
[0018] This embodiment takes the convolution operation (convolution kernel size is 3×3, step size is 1) running on an embedded device with a cache of 64MB and 8 CPU cores as an example.
[0019] See also Figure 1 , the present invention is carried out according to the following steps: S101: First, obtain the cache capacity and number of CPU cores provided by the device. The reinforcement learning controller uses this as the search space and selects a hardware resource configuration combination in combination with the previous rewards. For example, if an embedded device provides 64MB of cache and 8 CPU cores, one data occupies 8B and one cache line can store 8 data, the controller can choose to put 16MB of cache and 3 cores (1 for memory access and 2 for computing) into subsequent use.
[0020] S102: Obtain the current operation type and input data. In this example, the operation type is a convolution operation, the convolution kernel size is 3×3, the step size is 1, and the size of the input image is 522×522.
[0021] S103: The current operation is convolution, and the cyclic characteristics of implementing the convolution operation are analyzed.
[0022] S104: According to the loop characteristics, data reuse characteristics and hardware allocation of the convolution operation, a block size constraint condition is proposed, and the optimal block size that can minimize the CPU idle time is obtained under the constraint condition; a pre-fetch memory access strategy is adopted to allow memory access and calculation to form a pipeline in parallel, and to take out the required data in advance to calculate hidden memory access, thereby further reducing the CPU idle time. In this embodiment, when 239MB cache, 1 computing core and 3 memory access cores are provided, the optimal block size is 8×150. The pipeline formed by the pre-fetch memory access strategy is shown in Figure 2 middle.
[0023] S105: The cache utilization rate at the optimal block size is 80%, the CPU idle time is 22381600 cycles, the energy consumption is 528.4nJ, and the calculated reward value is 0.82. The reward value is fed back to the reinforcement learning controller for learning.
[0024] Repeat steps S101-S105 until the number of selections reaches a specified value.
[0025] The reinforcement learning controller finally selected the best resource allocation as 131MB cache, 1 computing core and 1 memory access core. At this time, the optimal block size is 8×520, the cache utilization is 99%, the CPU waiting time is 22329200 cycles, and the energy consumption is 528.2nJ.
[0026] See also Figure 2 , the prefetch access strategy forms a pipeline as follows: When calculating at iteration T, the memory access core writes back the result generated by iteration T-1, and then fetches the data required for iteration T+1; similarly, when calculating at iteration T+1, the memory access core writes back the result generated by iteration T, and then fetches the data required for iteration T+2.
[0027] See also Figure 3 , the selection and learning process of the reinforcement learning controller is as follows: S101: The reinforcement learning controller makes a selection in the search space to determine the cache allocation amount and the CPU core usage.
[0028] S102: The selected hardware configuration is put into use to obtain an optimal block size, and three performance indicators, CPU idle waiting time, cache utilization and energy consumption, are obtained at the same time.
[0029] S103: Calculate the reward value based on the three performance indicators obtained, and return it to the reinforcement learning controller so that it can start a new round of selection based on it.
Claims
1. A loop block partitioning and resource allocation method for reducing CPU idle time of convolutional neural network in embedded devices, characterized in that: The method comprises the following specific steps: S1: Based on the cache capacity, number of computing cores, and learning rewards of the current hardware device, the reinforcement learning controller selects the hardware allocation; S2: Get the type of the current convolutional neural network operation and the corresponding input data size; the convolutional neural network takes the image as input, and gets the size of the image input each time; S3: Analyze the characteristics of the loop and data reuse properties of the current operation type; S4: According to the characteristics of the loops of each operation, the data reuse characteristics and the hardware allocation, the block size constraint conditions are proposed, and the optimal block size that can minimize the CPU waiting time is obtained under the constraint conditions; the pre-fetch memory access strategy is adopted to make the memory access and calculation proceed in parallel to form a pipeline, and the required data is taken out in advance to calculate the hidden memory access, so as to further reduce the CPU waiting time; S5: Calculate the reward value based on the cache utilization, CPU idle time and energy consumption when the optimal block size is used, and feed it back to the reinforcement learning controller for learning; repeat steps S1-S4, and the reinforcement learning controller continuously selects and learns until the number of selections reaches the specified value, and the reinforcement learning controller selects the best resource allocation.
2. The method for cyclic blocking and resource allocation according to claim 1, characterized in that: The step S3 of analyzing the characteristics of the loop and data reuse characteristics corresponding to the current operation type specifically includes: S3-1: The convolution operation is implemented by a four-dimensional loop. The innermost two layers iteratively process the data in a window, and the outermost two-dimensional loop iteratively processes all windows of the input data. The innermost two layers of loops are expanded to obtain a two-dimensional loop, and each iteration processes a window. At this time, it is observed that there is a lot of data reuse in both the horizontal and vertical directions; S3-2: The pooling operation cuts the input image into windows of the same shape and size, performs the same operation on each window, and finally obtains the pooling result; the pooling operation is implemented as a four-dimensional loop, the innermost two layers traverse each data in the window, and the outermost two layers traverse each window in the input image; there is no data reuse in this loop, but there is room for improvement in the effective data load; S3-3: The activation operation performs a nonlinear transformation on each pixel of the input image, which is implemented by a two-dimensional loop. There is no data reuse in this loop, but there is room for improvement in the effective data load. S3-4: Batch normalization performs batch normalization on each pixel of the input image, which is implemented by a two-dimensional loop. There is no data reuse in this loop, but there is room for improvement in the effective data load. S3-5: The full connection operation is the weighted addition of the one-dimensional input vector and the weight matrix, which is implemented by a two-dimensional loop; the inner loop traverses the input vector, and the two-layer loop traverses the weight matrix. The multiplication of a vector element and a matrix element is an iteration; the inner loop is executed to obtain a result, and the input vector must be completely fetched into the cache during the execution process; the outer loop is executed to obtain the result matrix, and the weight matrix must be completely fetched into the cache during the execution process. The data requirement order of the weight matrix is inconsistent with the storage order. In order to use part of the data in the cache line, the same cache line will be repeatedly fetched into the cache, and the memory access overhead is extremely large; the data reuse degree of the input vector and the weight matrix is extremely high.
3. The method for cyclic blocking and resource allocation according to claim 1, characterized in that: The step S4 specifically includes: S4-1: According to the characteristics of the cycles of each operation, data reuse and cache provision, propose block size constraints; S4-2: Obtain a set of block size solutions according to the block size constraint. Each solution divides the iterations distributed in the two-dimensional space into rectangular areas, i.e., blocks. Data is read, processed, and results are written back in units of blocks. S4-3: A pre-fetch memory access strategy is used to make calculation and memory access proceed in parallel, so as to calculate hidden memory access and fetch the subsequent required data in advance: while a block is being calculated, the memory access core first writes back the calculation result generated by the previous block, and then fetches the data required for the next block; in a multi-core system, one CPU core is dedicated to memory access, and other CPU cores jointly complete the calculation; S4-4: Calculate the CPU idle waiting time for each solution in the block size solution set, compare and obtain the minimum CPU idle waiting time, and record the block size corresponding to the minimum CPU idle waiting time, i.e., the optimal block size.
4. The method for cyclic blocking and resource allocation according to claim 1, characterized in that: The step S1 specifically includes: S1-1: Generate the search space of the reinforcement learning controller based on the cache capacity and number of computing cores of the current hardware device; S1-2: If it is the first selection, the reinforcement learning controller randomly selects a resource allocation; otherwise, the reinforcement learning controller selects a resource allocation in the search space based on the reward received from the previous selection.
5. The method for cyclic blocking and resource allocation according to claim 1, characterized in that: The step S5 specifically includes: S5-1: Calculate the reward value according to the cache utilization, CPU idle time and energy consumption at the optimal block size, and feed the reward value back to the reinforcement learning controller; S5-2: The reinforcement learning controller learns the reward value, and the learning result affects the subsequent selection of the reinforcement learning controller; when the number of selections of the reinforcement learning controller is less than the specified value, it returns to S1; when the number of selections is equal to the specified value, the reinforcement learning controller selects the best resource allocation.
Citation Information
Cited By
Accelerator multi-objective optimization method driven by adaptive reinforcement learning
CN120911543A
An Adaptive Reinforcement Learning-Driven Accelerator Multi-Objective Optimization Method
CN120911543B