A simplified processing method for Concat operator in storage-computing integrated intelligent computing architecture
By reasonably planning the input and output channel segmentation and communication connection of the Concat operator in the integrated intelligent computing architecture, the data splicing problem caused by the Concat operator is solved, the hardware design is simplified and the communication bottleneck is eliminated, and efficient artificial intelligence algorithm deployment is achieved.
Patent Information
- Application Number
- CN202310748552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In the integrated intelligent computing architecture of storage and computing, the data splicing problems caused by the Concat operator in the convolutional neural network increase hardware overhead and scheduling difficulty, and may form a communication bottleneck.
By reasonably planning the channel segmentation method of the input and output convolutional layer of the Concat operator and the communication connection relationship between the convolutional layers at the task graph level, the data splicing introduced by the Concat operator is directly avoided, thereby simplifying hardware circuit design and eliminating communication bottlenecks.
This method greatly simplifies the hardware circuit design, reduces the hardware overhead and scheduling difficulty of the interconnect structure, and releases the bandwidth pressure introduced by multiple Concat input data, eliminating the potential for communication bottlenecks.
Smart Images

Figure CN116757252B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent technology, and specifically relates to a Concat operator simplified processing method in a storage-computation integrated intelligent computing architecture. Background Art
[0002] In recent years, artificial intelligence has developed rapidly. Artificial intelligence algorithms represented by neural networks are computationally and memory-intensive tasks. Most of the hardware resources in traditional CPUs (processors) are used for general control rather than computing, so they cannot provide sufficient computing power for the deployment of artificial intelligence algorithms. GPGPUs (general-purpose graphics processors) provide more computing resources at the expense of general control performance, providing an ideal deployment platform for artificial intelligence algorithms. However, GPGPUs still belong to the von Neumann architecture, and their computing and storage are separated. Therefore, in GPGPUs, data transfer between neural network layers must pass through memory. As the computing speed of GPGPUs increases, the time it takes to perform a calculation is getting smaller and smaller, which makes the proportion of the time spent on data exchange between computing units and memory continue to increase, even exceeding the computing time, forming a memory constraint. Memory constraints mean that memory access has become a bottleneck for computing. At this time, it is difficult to expand computing power simply by expanding the scale of computing units.
[0003] In order to eliminate memory constraints, it is necessary to reduce or avoid data movement between computing units and memory. However, in the von Neumann architecture where computing and storage are separated, memory cannot be bypassed by data. Therefore, it is necessary to go beyond the scope of the von Neumann architecture and explore a computing architecture that integrates storage and computing, that is, a storage-computing integrated architecture. In the storage-computing integrated architecture, the memory is given the function of computing, so the data generated by the memory can be calculated on the spot, thereby greatly reducing or even avoiding data movement during computing and eliminating memory constraints.
[0004] The current storage-computing integrated systems are mainly divided into three categories according to the type of in-memory computing devices they are based on: SRAM (static random access memory) storage-computing integrated, DRAM (dynamic random access memory) storage-computing integrated, and NVM (Non-Volatile Memory) storage-computing integrated. Due to the low storage density of SRAM, SRAM storage-computing integrated is difficult to achieve large capacity and is insufficient to support the deployment of large-scale artificial intelligence algorithms. DRAM has an extremely complex and closed hardware model and protocol stack, so only memory manufacturers have the conditions to conduct DRAM storage-computing integrated research. NVM, such as ReRAM, MRAM, PCM, and NOR-FLASH, not only has high storage density, but also has a simple and unified hardware model and strong universality. In addition, NVM also has the advantage of low power consumption. Therefore, how to design an NVM storage-computing integrated intelligent computing architecture to achieve an energy-efficient artificial intelligence algorithm deployment platform has become an open and valuable research topic.
[0005] In the NVM storage-computing integrated intelligent computing architecture (hereinafter referred to as the storage-computing integrated intelligent computing architecture), multiple computing units (Tiles) are usually combined into a computing array, and a special interconnection structure is provided in the computing array for communication between multiple computing units. When executing an algorithm task, each computing unit is responsible for performing a part of the matrix multiplication operation, and multiple computing units can send input data and transmit operation results in parallel, relying on the interconnection structure in the computing array, thereby realizing the forward reasoning process of the neural network. Summary of the invention
[0006] The Concat (Concatenate) operator in the convolutional neural network is used to concatenate data or vectors, which brings difficulties to the model deployment in the storage-computing integrated intelligent computing architecture. On the one hand, in order to support the Concat operator, it is usually necessary to set up a separate data concatenation unit and set up a separate receiving end buffer for each input of the Concat operator, which increases the hardware overhead and scheduling difficulty; on the other hand, when the Concat operator needs to receive multiple inputs (common in GoogleNet), the aggregate communication of multiple data will increase the local bandwidth pressure in the interconnection structure, which may form a communication bottleneck.
[0007] The present invention aims to propose a simplified processing method for the Concat operator in a storage-computing integrated intelligent computing architecture. By rationally planning the channel segmentation method of the input and output convolutional layers of the Concat operator and the communication connection relationship between the convolutional layers at the task graph level, the data splicing introduced by the Concat operator can be directly avoided, thereby greatly simplifying the hardware circuit design and eliminating the hidden danger of communication bottlenecks caused by local bandwidth pressure.
[0008] In order to achieve the above technical objectives, the present invention proposes the following solutions:
[0009] A simplified processing method for the Concat operator in a storage-computing integrated intelligent computing architecture includes the following two parts:
[0010] P1.Concat operator’s input and output channel segmentation method of the preceding and succeeding convolutional layers;
[0011] P2.Concat operator input and output communication connection planning process.
[0012] Furthermore, the division method of the input and output channels of the preceding and succeeding convolutional layers of the Concat operator includes the following contents:
[0013] The input and output channels of all the predecessor convolutional layers of the Concat operator are split according to the channel-by-channel Xbar mapping method, and the output channels of the successor convolutional layers of the Concat operator are split according to the conventional channel-by-channel Xbar mapping method. After mapping and splitting, the lengths of all input and output channel slices of the predecessor convolutional layer and all output channel slices of the successor convolutional layer are equal to K or less than K.
[0014] Furthermore, the input channels of the subsequent convolutional layer are not split according to the channel-by-channel Xbar mapping method, but are split according to the length of each output channel slice in each previous convolutional layer, that is, when splitting the input channels of the subsequent convolutional layer, the length of each slice is equal to the length of the output channel slice in the corresponding previous convolutional layer.
[0015] Furthermore, the Concat operator input and output communication connection planning process includes the following contents:
[0016] First, set up a two-layer nested loop. The loop variable of the outer loop is n, which is used to indicate the nth previous convolution layer of the Concat operator, and the loop range is [1, N]. The loop variable of the inner loop is t, which is used to indicate the tth output channel slice of the nth previous convolution layer of the Concat operator, and the loop range is [1, T n ], where T nis the number of output channel slices contained in the nth predecessor convolutional layer; two counters c and s are set. Initially, n, t, and c are all 1, and s is 0. Then the loop starts. At each step of the loop, the tth output channel slice of the nth predecessor convolutional layer is analyzed, and an input channel slice is cut out in the subsequent convolutional layer. The sequence number of the slice is assigned to c and recorded as the cth input channel slice of the subsequent convolutional layer. The channel range of the input channel slice is [s, s+m), where m is the length of the output channel slice of the current analyzed predecessor convolutional layer; at the same time, a communication connection is established between the output channel slice of the current analyzed predecessor convolutional layer and the input channel slice of the subsequent convolutional layer just cut out, then c is incremented by 1, s is incremented by m, and the next loop starts; this is repeated until all output channel slices of all predecessor convolutional layers establish communication connections with the subsequent convolutional layers.
[0017] Furthermore, the method for simplifying the processing of the Concat operator in the storage-computation integrated intelligent computing architecture includes the following steps:
[0018] (1) For the Concat operator in the neural network operator graph, search upstream from each input of the operator to obtain all the previous convolutional layers, and then search downstream from the output of the operator to obtain the subsequent convolutional layers;
[0019] (2) All the preceding convolutional layers of the Concat operator are divided into input and output channels according to the conventional channel-by-channel Xbar mapping method, and the output channels of the succeeding convolutional layers of the Concat operator are divided into output channels according to the conventional channel-by-channel Xbar mapping method. After mapping and division, the lengths of all input and output channel slices of the preceding convolutional layer and all output channel slices of the succeeding convolutional layer are equal to or less than K;
[0020] (3) Sort the N preceding convolutional layers of the Concat operator according to the concatenation order of the Concat operator required by the original model and label them 1, 2, ..., N in sequence;
[0021] (4) Run the communication connection planning program and set a two-layer nested loop. The loop variable of the outer loop is n, which is used to indicate the nth previous convolutional layer of the Concat operator, and the loop range is [1, N]. The loop variable of the inner loop is t, which is used to indicate the tth output channel slice of the nth previous convolutional layer of the Concat operator, and the loop range is [1, T n ], where T nis the number of output channel slices contained in the n-th predecessor convolutional layer; two counters c and s are set, and then the loop is started. At each step of the loop, the t-th output channel slice of the n-th predecessor convolutional layer is analyzed, and an input channel slice is cut out in the successor convolutional layer. The sequence number of the slice is assigned to c, which is recorded as the c-th input channel slice of the successor convolutional layer. The channel range of the input channel slice is [s, s+m), where m is the length of the output channel slice of the current successor convolutional layer being analyzed; at the same time, a communication connection is established between the output channel slice of the current successor convolutional layer being analyzed and the input channel slice of the successor convolutional layer just cut out, and then c is incremented by 1, s is incremented by m, and the next loop is started; this is repeated until all output channel slices of all successor convolutional layers establish communication connections with the successor convolutional layers.
[0022] Furthermore, in step (1), in the convolutional neural network, the weight of each convolutional layer is a matrix with a size of the number of input channels × the size of the convolution kernel × the number of output channels.
[0023] Furthermore, in step (1), in the convolutional neural network, the input channel and output channel of the convolution are divided into multiple slices of length K during mapping.
[0024] Furthermore, in step (1), in the convolutional neural network, each Tile contains an Xbar or a large logical Xbar composed of multiple Xbars.
[0025] Furthermore, in step (2), the number of input channels of the subsequent convolutional layer is equal to the sum of the number of output channels of all its predecessor convolutional layers, while the number of input channels of the predecessor convolutional layer and the number of output channels of the subsequent convolutional layer are arbitrary.
[0026] Furthermore, in step (4), initially n, t, c are all 1, and s is 0.
[0027] Compared with the prior art, the present invention has the following advantages:
[0028] The existing technology does not simplify or optimize the data splicing of the Concat operator of the convolutional neural network. The simplified processing method of the Concat operator proposed in the present invention can eradicate the data splicing introduced by the Concat operator in the storage-computing integrated intelligent computing architecture, thereby greatly reducing the hardware overhead and scheduling difficulty of the actual interconnection structure, and at the same time can release the bandwidth pressure introduced by multi-channel Concat input data and eliminate the hidden danger of communication bottleneck. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is an example diagram of a Concat operator with two preceding convolutional layers and one succeeding convolutional layer.
[0030] Figure 2 In order to split the input channels of Conv3 in a conventional channel-by-channel mapping manner, there is a data disassembly diagram.
[0031] Figure 3 The input channel of Conv3 is split according to the output channel slice length of Conv1 and Conv2, eliminating the data disassembly diagram. DETAILED DESCRIPTION
[0032] In an embodiment of the present invention, a method for simplifying the processing of Concat operators in a storage-computing integrated intelligent computing architecture includes the following steps:
[0033] (1) For the Concat operator in the neural network operator graph, search upstream from each input of the operator to obtain all the predecessor convolutional layers, and then search downstream from the output of the operator to obtain the successor convolutional layers.
[0034] (2) All the preceding convolutional layers of the Concat operator are divided into input and output channels according to the conventional channel-by-channel Xbar mapping method, and the output channels of the following convolutional layers of the Concat operator are divided into output channels according to the conventional channel-by-channel Xbar mapping method. After mapping and division, the lengths of all input and output channel slices of the preceding convolutional layer and all output channel slices of the following convolutional layer are equal to K or less than K (the remainder obtained when the number of output channels cannot be divided by K).
[0035] (3) According to the concatenation order of the Concat operator required by the original model, the N preceding convolutional layers of the Concat operator are sorted and labeled 1, 2, ..., N in sequence.
[0036] (4) Run the communication connection planning program and set a two-layer nested loop. The loop variable of the outer loop is n, which is used to indicate the nth previous convolutional layer of the Concat operator, and the loop range is [1, N]. The loop variable of the inner loop is t, which is used to indicate the tth output channel slice of the nth previous convolutional layer of the Concat operator, and the loop range is [1, T n ], where T nis the number of output channel slices contained in the nth predecessor convolution layer. Set two counters c and s. Initially, n, t, and c are all 1, and s is 0. Then the loop starts. At each step of the loop, the tth output channel slice of the nth predecessor convolution layer is analyzed, and an input channel slice is cut out in the subsequent convolution layer. The serial number of this slice is assigned to c and recorded as the cth input channel slice of the subsequent convolution layer. The channel range of this input channel slice is [s, s+m), where m is the length of the output channel slice of the current analyzed predecessor convolution layer. At the same time, a communication connection is established between the output channel slice of the current analyzed predecessor convolution layer and the input channel slice of the subsequent convolution layer just cut out. After that, c is incremented by 1, s is incremented by m, and the next loop starts. This process is repeated until all output channel slices of all predecessor convolution layers establish communication connections with the subsequent convolution layers.
[0037] In a convolutional neural network, the weight of each convolution layer is a matrix of size (number of input channels × convolution kernel size × number of output channels). Each Tile contains an Xbar (in-memory computation array) or a large logical Xbar composed of multiple Xbars. For an Xbar of size K×K, it can receive an input vector of length K during calculation, and obtain an output vector of length K after matrix multiplication. Taking the most common channel-by-channel Xbar mapping as an example, during mapping, the input channels and output channels of the convolution are divided into multiple slices of length K, respectively called input channel slices and output channel slices. Each Tile is responsible for one input channel slice and one output channel slice. The same input channel slice usually needs to be deployed to multiple tiles to realize the calculation of vectors at different positions in the convolution window, while the same output channel slice is deployed to a unique tile. For a certain convolution layer, the set of all tiles responsible for the i-th input channel slice and the j-th output channel slice is called a Tile cluster, denoted by TC. ij .
[0038] For a Concat operator in a neural network, the first convolution layer obtained by searching upstream from each input of the operator is the predecessor convolution layer of the Concat operator. Similarly, the first convolution layer obtained by searching downstream from the output of the operator along each branch (if there are multiple branches) is the successor convolution layer of the Concat operator. In general, each input of each Concat operator has a corresponding predecessor convolution layer, and since the result obtained by Concat may be sent to multiple subsequent convolution layers through branches, each Concat operator may have multiple successor convolution layers. For the case where the Concat operator has multiple successor convolution layers, the simplified processing method of the Concat operator proposed in the present invention requires the same processing method for each successor convolution layer. Therefore, for the convenience of discussion and without loss of generality, it is assumed that the Concat operator to be discussed has only one successor convolution layer.
[0039] The input and output channels of all the preceding convolutional layers of the Concat operator are split according to the conventional channel-by-channel Xbar mapping method, and the output channels of the subsequent convolutional layers of the Concat operator are split according to the conventional channel-by-channel Xbar mapping method. After mapping and splitting, the lengths of all input and output channel slices of the preceding convolutional layer and all output channel slices of the subsequent convolutional layer are equal to K or less than K (the remainder obtained when the number of output channels cannot be divided by K).
[0040] The key to the present invention lies in the way of segmenting the input channel of the subsequent convolutional layer: instead of segmenting according to the conventional channel-by-channel Xbar mapping method, segmenting is performed according to the length of each output channel slice in each previous convolutional layer. That is, when segmenting the input channel of the subsequent convolutional layer, the length of each slice is made equal to the length of the output channel slice in the corresponding previous convolutional layer. Such a segmentation method can avoid data disassembly caused by non-aligned segmentation.
[0041] For example, for Figure 1 The Concat operator shown in the figure requires that the number of input channels of the subsequent convolutional layer (Conv3) must be equal to the sum of the number of output channels of all its predecessor convolutional layers (Conv1 and Conv2), according to the requirement of the number of channels of the convolutional neural network, while the number of input channels of the predecessor convolutional layer and the number of output channels of the subsequent convolutional layer can be arbitrary.
[0042] When the input channels of the subsequent convolutional layer are split, if the split is performed in the conventional channel-by-channel mapping manner, data disconnection may occur when establishing a communication connection. Assuming K = 64, Figure 2As shown in the figure, the calculation result obtained in the first output channel slice of Conv2 needs to be split and sent to the second and third input channel slices of Conv3, while the calculation result obtained in the second output channel slice of Conv2 needs to be split and sent to the third and fourth input channel slices of Conv3, and the multi-channel data input needs to be spliced at the second and third input channel slices of Conv3. This shows that the input channel segmentation of the subsequent convolution layer in the conventional channel-by-channel mapping method cannot eliminate data splicing, because the segmentation boundary of the input channel of the subsequent convolution layer is not aligned with the output channel slice in the previous convolution layer.
[0043] The present invention proposes to segment the input channel of the subsequent convolutional layer according to the length of each output channel slice of each preceding convolutional layer, such as Figure 3 As shown. Different from Figure 2 The length of the second input channel slice of Conv3 is divided into 64 (i.e., K), and the method proposed in the present invention divides the length of the second input channel slice of Conv3 into 36 (100-64) in order to correspond to the second output channel slice of Conv1. This division method can ensure that each output channel slice in the previous convolution layer corresponds to a unique input channel slice in the subsequent convolution layer. In actual operation, the results of all tiles in each output channel slice of each previous convolution layer can be fused into a unique result and sent to all tiles in the input channel slice of the subsequent convolution layer in a multicast manner, thereby converting the original data splicing into multiple independent and parallel communication connections. Therefore, the simplified processing method of the Concat operator proposed in the present invention can eradicate the data splicing introduced by the Concat operator.
Claims
1. A simplified processing method for the Concat operator in a storage-computing integrated intelligent computing architecture. It is characterized in that It consists of the following two parts: P1.Concat operator’s input and output channel segmentation method of the preceding and succeeding convolutional layers; P2.Concat operator input and output communication connection planning process; The division method of the input and output channels of the preceding and succeeding convolutional layers of the Concat operator includes the following contents: All the preceding convolutional layers of the Concat operator are divided into input and output channels according to the channel-by-channel Xbar mapping method, where Xbar is an in-memory calculation array, and the output channels of the subsequent convolutional layers of the Concat operator are divided according to the conventional channel-by-channel Xbar mapping method. After mapping and division, the lengths of all input and output channel slices of the preceding convolutional layer and all output channel slices of the subsequent convolutional layer are equal to or less than K; For the input channels of the subsequent convolutional layer, they are not split according to the channel-by-channel Xbar mapping method, but are split according to the length of each output channel slice in each previous convolutional layer. That is, when splitting the input channels of the subsequent convolutional layer, the length of each slice is equal to the length of the output channel slice in the corresponding previous convolutional layer. The Concat operator input and output communication connection planning process includes the following contents: First, set up a two-layer nested loop. The loop variable of the outer loop is n, which is used to indicate the nth previous convolution layer of the Concat operator, and the loop range is [1, N]. The loop variable of the inner loop is t, which is used to indicate the tth output channel slice of the nth previous convolution layer of the Concat operator, and the loop range is [1, T n ], where T n is the number of output channel slices contained in the nth predecessor convolutional layer; two counters c and s are set. Initially, n, t, and c are all 1, and s is 0. Then the loop starts. At each step of the loop, the tth output channel slice of the nth predecessor convolutional layer is analyzed, and an input channel slice is cut out in the subsequent convolutional layer. The sequence number of the slice is assigned to c and recorded as the cth input channel slice of the subsequent convolutional layer. The channel range of the input channel slice is [s, s+m), where m is the length of the output channel slice of the current analyzed predecessor convolutional layer; at the same time, a communication connection is established between the output channel slice of the current analyzed predecessor convolutional layer and the input channel slice of the subsequent convolutional layer just cut out, then c is incremented by 1, s is incremented by m, and the next loop starts; this is repeated until all output channel slices of all predecessor convolutional layers establish communication connections with the subsequent convolutional layers.
2. According to the method for simplifying the processing of Concat operators in a storage-computation-integrated intelligent computing architecture according to claim 1, It is characterized in that The following steps are involved: (1) For the Concat operator in the neural network operator graph, search upstream from each input of the operator to obtain all the previous convolutional layers, and then search downstream from the output of the operator to obtain the subsequent convolutional layers; (2) All the previous convolutional layers of the Concat operator are divided into input and output channels according to the conventional channel-by-channel Xbar mapping method, and the output channels of the subsequent convolutional layers of the Concat operator are divided into output channels according to the conventional channel-by-channel Xbar mapping method. After mapping and division, the lengths of all input and output channel slices of the previous convolutional layer and all output channel slices of the subsequent convolutional layer are equal to K or less than K; (3) Sort the N preceding convolutional layers of the Concat operator according to the concatenation order of the Concat operator required by the original model and label them 1, 2, ..., N in sequence; (4) Run the communication connection planning program and set a two-layer nested loop. The loop variable of the outer loop is n, which is used to indicate the nth previous convolutional layer of the Concat operator, and the loop range is [1, N]. The loop variable of the inner loop is t, which is used to indicate the tth output channel slice of the nth previous convolutional layer of the Concat operator, and the loop range is [1, T n ], where T n is the number of output channel slices contained in the nth preceding convolutional layer; Two counters c and s are set, and then the loop is started. At each step of the loop, the tth output channel slice of the nth predecessor convolution layer is analyzed, and an input channel slice is cut out in the successor convolution layer. The serial number of the slice is assigned to c, which is recorded as the cth input channel slice of the successor convolution layer. The channel range of the input channel slice is [s, s+m), where m is the length of the output channel slice of the current successor convolution layer being analyzed; at the same time, a communication connection is established between the output channel slice of the current successor convolution layer being analyzed and the input channel slice of the successor convolution layer just cut out, and then c is incremented by 1, s is incremented by m, and the next loop starts; this is repeated until all output channel slices of all predecessor convolution layers establish communication connections with the successor convolution layers.
3. According to claim 2, the Concat operator simplified processing method in the storage and computing integrated intelligent computing architecture, It is characterized in that In step (1), in the convolutional neural network, the weight of each convolutional layer is a matrix with the size of the number of input channels × the size of the convolution kernel × the number of output channels.
4. According to claim 2, the Concat operator simplified processing method in the storage and computing integrated intelligent computing architecture, It is characterized in that In step (1), in the convolutional neural network, the input channel and output channel of the convolution are divided into multiple slices of length K during mapping.
5. According to claim 2, the Concat operator simplified processing method in the storage and computing integrated intelligent computing architecture, It is characterized in that In step (1), in the convolutional neural network, each Tile contains an Xbar or a large logical Xbar composed of multiple Xbars.
6. According to claim 2, the Concat operator simplified processing method in the storage and computing integrated intelligent computing architecture, It is characterized in that In step (2), the number of input channels of the subsequent convolutional layer is equal to the sum of the number of output channels of all its predecessor convolutional layers, while the number of input channels of the predecessor convolutional layer and the number of output channels of the subsequent convolutional layer are arbitrary.
Citation Information
Patent Citations
Data segmentation operation method of neural network based on NOR Flash module
CN111222626A
Neural network accelerator automation code generation method
CN114691108A