A Sparse Network Load Balancing Scheduling Method for MPSoC
By adopting the MPSoC-oriented sparse network load balancing scheduling method on the embedded platform, the balanced loading of non-zero activation values and weight values is achieved, and combined with the hierarchical buffer mapping mechanism, the problem of load imbalance during sparse network inference acceleration is solved, and the computing efficiency is improved.
Patent Information
- Application Number
- CN202111164396.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The prior art has the problem of load imbalance when the inference acceleration of sparse networks on embedded platforms, resulting in low computing efficiency.
The sparse network load balancing scheduling method for MPSoC is adopted to dynamically construct feature map chunking and combine the convolution kernel to achieve balanced loading of non-zero activation values and weight values, and convolution calculation and mapping output are performed in combination with the hierarchical buffer mapping mechanism.
The computational load balancing of sparse networks in the inference stage is realized, and the inference acceleration performance of sparse networks is improved.
Smart Images

Figure CN113900803B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of embedded intelligent algorithm inference acceleration, and relates to a sparse network load balancing scheduling method for MPSoC. Background Art
[0002] Deep learning has been widely applied in fields such as image recognition, speech processing, natural language processing, language translation, and autonomous driving, and has become a key method for solving complex problems in recent years. However, with the improvement of model prediction accuracy, the number of layers of convolutional neural networks has become deeper and deeper, and the number of parameters and the amount of calculations have increased sharply, resulting in a large amount of computing resources, storage resources, and access bandwidth required in the inference stage, which limits its deployment and application on the embedded side.
[0003] Research shows that convolutional neural networks have a high degree of parameter redundancy. Without reducing the accuracy, by using model compression methods such as pruning and compression, the sparsity of most layers in convolutional neural networks can reach more than 70%, and some layers can even reach about 95%. This method can greatly reduce the effective number of network parameters and the amount of calculations, but it introduces a complex network structure and large asymmetry, making it difficult for traditional convolutional network inference acceleration methods to fully utilize the characteristics of sparse networks, and the inference acceleration effect is poor.
[0004] Currently, for the inference acceleration of sparse networks on embedded platforms, methods such as power gating, zero skipping, and Cartesian product - result hash mapping are mainly used. Power gating and zero skipping methods can be compatible with both dense networks and sparse networks, but the storage form of the network needs to adopt the dense network form and cannot effectively utilize the compressed storage characteristics of sparse networks. The Cartesian product - result hash mapping method can fully utilize the compressed characteristics of sparse networks, but due to the lack of consideration of the structural characteristics of sparse networks in the data scheduling part, the loading of non - zero weight values and activation values is unbalanced, resulting in the problem of unbalanced computing load in the inference stage and significantly reducing the inference acceleration performance of sparse networks. Summary of the Invention
[0005] Aiming at the problem of low computing efficiency caused by unbalanced load during the inference acceleration of sparse networks on embedded platforms in the prior art, the purpose of the present invention is to provide a sparse network load balancing scheduling method for MPSoC, which solves at least part of the above - mentioned technical problems.
[0006] An embodiment of the present invention provides a sparse network load balancing scheduling method for MPSoC, including:
[0007] S1: Obtain the feature map of the current input layer;
[0008] S2: Obtain the hardware configuration parameters of the computing platform, and achieve balanced loading of non-zero activation values by dynamically constructing feature map blocks;
[0009] S3: Obtain the weight parameters of the current network layer, and achieve balanced loading of non-zero weight values by grouping and merging convolution kernels;
[0010] S4: Based on the balanced loading of the non-zero activation values and the balanced loading of the non-zero weight values, adopt a hierarchical buffer mapping mechanism for convolution calculation and mapping output to achieve sparse network load balancing scheduling.
[0011] Further, the step S2 includes:
[0012] S21: The hardware configuration parameters include buffer hardware configuration parameters and computing array hardware configuration parameters; obtain the buffer hardware configuration parameters and feature map partition parameters, perform partition division of the current feature map, and obtain multiple feature map partitions;
[0013] S22: Statistically calculate non-zero activation value parameters and convolution parameters according to the feature map partitions; calculate the load rate of the current feature map partition according to the non-zero activation value parameters and the computing array hardware configuration parameters, and generate a load rate information table;
[0014] S23: Statistically calculate the number and proportion of activation values that cause mapping conflicts according to the non-zero activation value parameters and convolution parameters, and generate a conflict rate information table;
[0015] S24: Calculate the optimal feature map block scheme for the current feature map partition according to the load rate information table and the conflict rate information table of the current feature map partition to achieve balanced loading of non-zero activation values.
[0016] Further, the step S3 includes:
[0017] S31: Obtain the weight parameters, calculate the weight load rates under different filter groups, and determine the optimal filter group; after determining the optimal filter group, calculate the weight load rates under different channel groups and determine the optimal channel group;
[0018] S32: Merge and load the convolution kernels according to the weight parameters, the computing array hardware configuration parameters, and the determined optimal channel group to achieve balanced loading of weights.
[0019] Further, the adopting a hierarchical buffer mapping mechanism for convolution calculation and mapping output to achieve sparse network load balancing scheduling includes:
[0020] According to the computing array hardware configuration parameters, configure the computing array in the sparse network into a general convolution computing mode or a depthwise separable convolution computing mode; the general convolution computing mode is single-channel feature map input and multi-channel convolution kernel input; the depthwise separable convolution computing mode is multi-channel feature map input and multi-channel convolution kernel input;
[0021] Input the computing result output by the convolution computing mode into a two-level buffer for hierarchical buffer storage; map and output the computing result in the hierarchical buffer storage to achieve load balancing scheduling of the sparse network.
[0022] Further, it includes: the two-level buffer includes a first-level buffer and a second-level buffer;
[0023] The first-level buffer is a multi-layer storage structure for storing the convolution computing results of the feature map blocks in the current feature map partition;
[0024] The second-level buffer is a single-layer storage structure for storing the convolution computing results of the entire feature map partition.
[0025] Further, mapping and outputting the feature map convolution computing results in the hierarchical buffer storage to achieve load balancing scheduling of the sparse network includes:
[0026] Establish an output mapping table according to the storage location mapping relationship of the feature map during the convolution computing process; in the output mapping table, mark the data points with the same mapping conflict location in different layer storage spaces;
[0027] Input the convolution computing results of the current feature map block into the first-level buffer according to the output mapping table; input the block convolution computing results in the first-level buffer into the corresponding positions of the second-level buffer;
[0028] Output all the data in the second-level buffer to the external storage space to complete the mapping output of the current feature map partition.
[0029] A load balancing scheduling method for a sparse network oriented to MPSoC provided by an embodiment of the present invention, compared with the prior art, adopts an equilibrium loading strategy for non-zero activation values and weight values to achieve load balancing scheduling of the sparse network, achieving computational load balancing in the inference stage of the sparse network under the Cartesian product-result hash mapping computing paradigm based on hierarchical mapping, thereby improving the inference acceleration performance of the sparse network.
[0030] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0031] Figure 1 It is a block diagram of a sparse network load balancing scheduling method for MPSoC provided by an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of a sparse network equilibrium scheduling algorithm provided by an embodiment of the present invention;
[0033] Figure 3 It is a schematic diagram of an activation value dynamic chunking algorithm provided by an embodiment of the present invention;
[0034] Figure 4 It is a schematic diagram of a grouped equilibrium scheduling algorithm for weights provided by an embodiment of the present invention;
[0035] Figure 5 It is a schematic diagram of depthwise separable convolution calculation provided by an embodiment of the present invention;
[0036] Figure 6 It is a schematic diagram of ordinary convolution calculation provided by an embodiment of the present invention. Detailed Embodiment
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0038] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", "top / bottom end", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0039] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "provided with", "internally connected", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0040] A sparse network load balancing scheduling method for MPSoC provided by an embodiment of the present invention is as follows Figure 1 shown, including:
[0041] S1: Obtain the feature map of the input current layer;
[0042] S2: Obtain the hardware configuration parameters of the computing platform, and achieve balanced loading of non-zero activation values by dynamically constructing feature map blocks;
[0043] S3: Obtain the weight parameters of the current network layer, and achieve balanced loading of non-zero weight values by grouping and merging convolution kernels;
[0044] S4: Based on the balanced loading of the non-zero activation values and the balanced loading of the non-zero weight values, adopt a hierarchical buffer mapping mechanism for convolution calculation and mapping output to achieve sparse network load balancing scheduling.
[0045] A sparse network load balancing scheduling method for MPSoC provided by an embodiment of the present invention, compared with the prior art, adopts a balanced loading strategy for non-zero activation values and weight values to achieve sparse network load balancing scheduling, achieving computational load balancing in the inference stage of the sparse network under the Cartesian product-result hash mapping calculation paradigm based on hierarchical mapping, thereby improving the inference acceleration performance of the sparse network.
[0046] As Figure 2 shown, a sparse network load balancing scheduling method for MPSoC provided by an embodiment of the present invention includes three parts: a feature map balanced loading module, a weight balanced loading module, and a hierarchical mapping calculation module. Among them, the feature map balanced loading module achieves balanced loading of non-zero activation values by dynamically constructing feature map blocks, the weight balanced loading module achieves balanced loading of weight values by grouping and merging convolution kernels, and the hierarchical mapping calculation module achieves convolution calculation and mapping output by adopting a hierarchical buffer mapping mechanism.
[0047] As Figure 3 shown, the feature map balanced loading module consists of three parts: a feature map load rate evaluation unit, a feature map conflict rate evaluation unit, and a feature map block dynamic construction unit.
[0048] The feature map load rate evaluation unit first determines the size of the feature map partition processed each time according to the hardware configuration parameters of the output buffer, then counts the number of non-zero activation values in the current partition row by row, calculates the load rate of the current partition according to the hardware configuration parameters of the computing array, generates an information table of the feature map load rate of the current partition, and finally processes the partition feature map data in all channels in turn.
[0049] The feature map conflict rate evaluation unit first statistically calculates the number and proportion of activation values with mapping conflicts row by row based on the non-zero activation value position distribution, convolutional kernel size, and sliding step size in the current partition, generates the current feature map partition load rate information table, and finally processes the feature map partition data in all channels in sequence.
[0050] The feature map conflict rate evaluation unit first statistically calculates the number and proportion of activation values with mapping conflicts row by row based on the non-zero activation value position distribution, convolutional kernel size, and sliding step size in the current partition, generates the current feature map partition conflict rate information table, and then processes the feature map partition data in all channels in sequence.
[0051] The feature map block dynamic construction unit first obtains the optimal feature map block scheme for the current feature map partition under the constraints of the maximum load rate and the minimum hash mapping conflict rate by using a heuristic search algorithm according to the current feature map partition load rate information table and the feature map conflict rate information table, determines the optimal feature map block size and calculation method for the current partition, and then processes the feature map partition data in all channels in sequence.
[0052] As Figure 4 shown, the weight balanced loading module consists of two parts: the multi-dimensional load rate evaluation unit and the convolutional kernel grouping and merging unit. The multi-dimensional load rate evaluation unit includes two parts: the filter load rate evaluation unit and the channel load rate evaluation unit. By evaluating the weight load rates under different groupings, the optimal filter grouping and channel grouping strategies are determined. The convolutional kernel grouping and merging algorithm realizes the balanced loading of weights through convolutional kernel grouping and merging.
[0053] The multi-dimensional load rate evaluation unit first uses the filter load rate evaluation algorithm to statistically calculate the number of weight values contained in different filters of the entire network layer, and determines the optimal filter grouping strategy by evaluating the weight load rates under different filter groupings. Then it uses the channel load rate evaluation algorithm to statistically calculate the number of weight values in different channels under the current filter grouping, processes all filter groupings in sequence, and determines the optimal channel grouping strategy by evaluating the weight load rates under different channel groupings.
[0054] The convolutional kernel grouping and merging unit first reorders the convolutional kernels within the current channel grouping according to the number of weight values, then groups the convolutional kernels according to the hardware parallelism of the computing array, so that the number of weight values within the group is as aligned with the hardware parallelism as possible and the number of groups is as small as possible. Finally, the convolutional kernels within the group are merged and loaded in units of convolutional kernel groups to achieve the balanced loading of weights.
[0055] The hierarchical mapping calculation module includes a calculation array unit, a hierarchical storage unit, and a scheduling control unit. The calculation array unit is used for the calculation of ordinary convolution and depthwise separable convolution. The hierarchical storage unit is used for caching the calculation results and reducing the hash mapping conflict rate. The scheduling control unit is used to control the loading of activation values and weight values and the mapping of calculation results.
[0056] As Figure 5 and Figure 6 shown, the calculation array unit can be configured into an ordinary convolution calculation mode and a depthwise separable convolution calculation mode through configuration parameters. In the ordinary convolution calculation mode, a single-channel feature map input and a multi-channel convolution kernel input are adopted. In the depthwise separable convolution calculation mode, a multi-channel feature map input and a multi-channel convolution kernel input are adopted.
[0057] The hierarchical storage unit is composed of two-level buffers. The first-level buffer is configured into a multi-layer storage structure according to the feature map block size. Its single-layer storage capacity is determined by the current feature map block size, and the number of layers is determined by the proportion of the block in the current feature map partition. The multi-layer storage structure jointly completes the storage and mapping of the calculation results of the current feature map block, and at the same time reduces the hash mapping conflict rate. The second-level buffer is configured into a single-layer storage structure according to the feature map partition size. Its storage capacity is determined by the feature map partition size and is used for the storage and mapping of the convolution calculation results of the entire feature map partition.
[0058] The scheduling control unit includes three parts: a conflict detection unit, a data loading unit, and a mapping output unit. The conflict detection unit is used to count and mark the activation values with mapping conflicts and establish a hierarchical mapping table for the current feature map block. The data loading unit is used to implement the loading of activation values, weight values, and partial sum data. The mapping output unit is used to implement the mapping of the convolution calculation results of the current feature map block to the first-level buffer, the mapping of the convolution calculation results of multiple blocks within the current feature map partition to the second-level buffer, and then the mapping output of the calculation results of the current feature map partition to the external memory.
[0059] The conflict detection unit, according to the position mapping relationship from the input feature map to the output feature map, establishes an output mapping table by traversing the input feature map, and marks the storage positions of the data points with mapping conflicts as different layer storage spaces in the first-level buffer, realizing the uniform distribution of the mapping conflict data points in different layer storage spaces.
[0060] The data loading unit loads the feature map data, weight data, and partial sum data in the second-level buffer into the corresponding calculation units for multiplication and accumulation operations through local broadcasting according to the calculation array configuration mode.
[0061] The mapping output unit writes the calculation results of the current feature map block convolution into the multi-layer storage space of the primary buffer according to the output mapping table, and keeps waiting until the block calculation is completed. Then, it writes the data in the primary buffer into the corresponding positions of the secondary buffer and clears the primary buffer until all the block data in the current partition are processed. Finally, it writes the entire partition data stored in the secondary buffer into the external storage space to complete the mapping output of the current feature map partition.
[0062] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and its improved concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A sparse network load balancing scheduling method for MPSoC, characterized in that, Including: S1: Obtain the feature map of the current input layer; S2: Obtain the hardware configuration parameters of the computing platform, and achieve balanced loading of non-zero activation values by dynamically constructing feature map blocks; S3: Obtain the weight parameters of the current network layer, and achieve balanced loading of non-zero weight values by grouping and merging convolution kernels; S4: Based on the balanced loading of the non-zero activation values and the balanced loading of the non-zero weight values, adopt a hierarchical buffer mapping mechanism for convolution calculation and mapping output to achieve sparse network load balancing scheduling; The step S2 includes: S21: The hardware configuration parameters include buffer hardware configuration parameters and computing array hardware configuration parameters; obtain the buffer hardware configuration parameters and feature map partition parameters, perform partition division of the current feature map, and obtain multiple feature map partitions; S22: Statistically calculate the non-zero activation value parameters and convolution parameters according to the feature map partitions respectively; calculate the load rate of the current feature map partition according to the non-zero activation value parameters and the computing array hardware configuration parameters, and generate a load rate information table; S23: Statistically calculate the number and proportion of activation values with mapping conflicts according to the non-zero activation value parameters and convolution parameters respectively, and generate a conflict rate information table; S24: Calculate the optimal feature map block scheme of the current feature map partition according to the load rate information table and the conflict rate information table of the current feature map partition to achieve balanced loading of non-zero activation values; The step S3 includes: S31: Obtain the weight parameters, calculate the weight load rates under different filter groups, and determine the optimal filter group; after determining the optimal filter group, calculate the weight load rates under different channel groups and determine the optimal channel group; S32: Merge and load the convolution kernel according to the weight parameters, the computing array hardware configuration parameters, and the determined optimal channel group to achieve balanced loading of weights.
2. The sparse network load balancing scheduling method for MPSoC according to claim 1, wherein The adoption of the hierarchical buffer mapping mechanism for convolution calculation and mapping output to achieve sparse network load balancing scheduling includes: According to the computing array hardware configuration parameters, configure the computing array in the sparse network as a normal convolution calculation mode or a depthwise separable convolution calculation mode; the normal convolution calculation mode is single-channel feature map input and multi-channel convolution kernel input; the depthwise separable convolution calculation mode is multi-channel feature map input and multi-channel convolution kernel input; Input the calculation results output by the convolution calculation mode into a two-level buffer for hierarchical buffer storage; map and output the calculation results in the hierarchical buffer storage to achieve sparse network load balancing scheduling.
3. The sparse network load balancing scheduling method for MPSoC according to claim 2, wherein, Including: The two-level buffer includes a first-level buffer and a second-level buffer; The first-level buffer is a multi-layer storage structure for storing the convolution calculation results of the feature map blocks in the current feature map partition; The second-level buffer is a single-layer storage structure for storing the convolution calculation results of the entire feature map partition.
4. The sparse network load balancing scheduling method for MPSoC according to claim 3, characterized in that Mapping and outputting the feature map convolution calculation results in the hierarchical buffer storage to achieve sparse network load balancing scheduling includes: An output mapping table is established according to the storage location mapping relationship of the feature maps during the convolution calculation process; in the output mapping table, data points with the same mapping conflict location are marked in different levels of storage space; According to the output mapping table, the convolution calculation results of the current feature map in blocks are input into the first-level buffer; the convolution calculation results in blocks in the first-level buffer are input into the corresponding positions of the second-level buffer; All the data in the second-level buffer are output to the external storage space to complete the mapping output of the current feature map partition.
Citation Information
Patent Citations
A sparse convolutional neural network accelerator and an implementation method
CN109635944A
An acceleration method for realizing sparse convolutional neural network inference for hardware
CN109711532A