Neural Network Mapping Method, Device, and Equipment for In-Memory Computing Chips

By arranging the weight matrix and bias of the neural network layer in a serpentine sequence, the current noise problem caused by insufficient utilization of memory and large bias values in the memory and computing integrated chip is solved, and the computing accuracy and efficiency are improved.

CN113988277BActive Publication Date: 2025-07-25BEIJING ZHICUN (WITIN) TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111184060.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-07-25
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

Traditional CPU+GPU architectures have speed and power consumption bottlenecks when dealing with large-scale neural networks, and the existing technology fails to effectively utilize the memory and computing unit when mapping the neural network to the memory and computing integrated chip, resulting in large bias values and high current noise, which affects the computing accuracy.

Method used

The weight matrix and bias of each layer of the neural network is arranged in a serpentine order, and the weight matrix number is mapped and sorted according to the minimum number of Bias rows and the weight matrix number, and arranged in the main array and Bias array of the integrated memory chip to optimize the utilization and bias value distribution of the integrated memory cell.

Benefits of technology

The utilization efficiency of the memory and computing integrated unit array is improved, the bias value on a single memory and computing integrated unit is reduced, current noise is reduced, and computing accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113988277B_ABST
    Figure CN113988277B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a neural network mapping method, device, and equipment for a memory-computation integrated chip. The method includes: performing mapping sorting on each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix; according to the mapping sorting result, arranging the weight matrix corresponding to each layer into the main array of the memory-computation integrated chip in sequence, and arranging the corresponding Bias into the position corresponding to the column of the arrangement position in the Bias array of the memory-computation integrated chip according to the arrangement position of the weight matrix and the minimum number of rows of Bias; when arranging the weight matrix corresponding to each layer in sequence, arranging it in a snake-like order, effectively utilizing the memory-computation integrated unit. In addition, based on the minimum number of rows of Bias and the idle situation of the Bias array, the number of rows occupied by each Bias is extended, reducing the bias value on a single memory-computation integrated unit, reducing current noise, and improving the operation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor technology, and in particular, to a neural network mapping method, device, and equipment for in-memory computing chips. Background Art

[0002] In recent years, with the continuous development of the three dimensions of algorithms, computing power, and data volume scale, machine learning technology has continuously demonstrated powerful advantages in solving many problems. Among them, artificial neural networks have received extensive attention due to their outstanding performance in fields such as image recognition, object detection, and semantic segmentation. However, with the expansion of the scale of neural networks, the traditional mode of processing neural network algorithms with a CPU+GPU architecture has gradually encountered bottlenecks in speed and power consumption. The root cause is that the separation of memory and computing under the von Neumann architecture causes the data-centric neural network algorithm to bring too large a data transmission overhead to the computing system, reducing the speed while increasing the power consumption.

[0003] In-memory computing technology solves the problems caused by the separation of memory and computing. By storing the weights of the neural network on the conductance of the flash array nodes in the in-memory computing (NPU) chip, and then sending the data source represented by voltage into the array. According to Ohm's law, the current output by the array is the product of the voltage and the conductance, thus completing the matrix multiplication and addition operation of the data source and the network weights. Essentially, it is performing analog computing rather than traditional digital computing.

[0004] In the whole process from the design to the production of in-memory computing chips, the design of the tool chain is an important link. In the design of the tool chain for in-memory computing chips, the technology of automatically mapping the weight parameters of a specific neural network to the chip array according to requirements is a key technology. When mapping the trained neural network to the in-memory computing unit array of the in-memory computing chip, in the order of each layer of the neural network, the weights and biases are mapped to the in-memory computing chip array in sequence. On the one hand, it cannot effectively utilize the in-memory computing units, increasing the scale of the in-memory computing unit array. On the other hand, since the bias is directly mapped to the in-memory computing chip array, the larger the value of the bias, the larger the conductance of the corresponding in-memory computing unit. Under the same voltage, the current of the in-memory computing unit is larger, which further leads to larger noise and affects the operation accuracy. Summary of the Invention

[0005] Aiming at the problems in the prior art, the present invention provides a neural network mapping method, device, and equipment for in-memory computing chips, which can at least partially solve the problems existing in the prior art.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, a neural network mapping method for in-memory computing chips is provided, including:

[0008] Map and sort each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix;

[0009] According to the mapping and sorting results, arrange the weight matrices corresponding to each layer into the main array of the memory - computing integrated chip in sequence, and arrange the corresponding Bias into the position corresponding to the column of the arrangement position in the Bias array of the memory - computing integrated chip according to the arrangement position of the weight matrix and the minimum number of rows of Bias;

[0010] Among them, when arranging the weight matrices corresponding to each layer in sequence, arrange them in a snake - like order.

[0011] Furthermore, the main array contains multiple memory - computing integrated unit blocks distributed in an array, and the weight matrices corresponding to each layer are arranged on the memory - computing integrated unit blocks in the main array;

[0012] The arrangement of the weight matrices corresponding to each layer in a snake - like order includes: for the weight matrix corresponding to each layer of the neural network, sequentially poll the memory - computing integrated unit blocks in the main array in a snake - like order to find the arrangement position.

[0013] Furthermore, the mapping and sorting of each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix includes:

[0014] Map and sort each layer according to the minimum number of rows of Bias corresponding to each layer;

[0015] For the layers with the same minimum number of rows of Bias, map and sort them according to the number of columns of the weight matrix.

[0016] Furthermore, the neural network mapping method for the memory - computing integrated chip also includes;

[0017] Write the weight matrices and Bias of each layer of the neural network to be mapped to the memory - computing integrated chip according to the arrangement results.

[0018] Furthermore, the sequentially polling the memory - computing integrated unit blocks in the main array in a snake - like order to find the arrangement position includes:

[0019] Sequentially poll whether there is a position that meets the arrangement conditions of the current layer in the memory - computing integrated unit blocks in the main array in a snake - like order;

[0020] If so, arrange the weight matrix of the current layer to the position that meets the arrangement conditions of the current layer;

[0021] If not, continue to poll the next memory - computing integrated unit block until a position that meets the arrangement conditions is found;

[0022] Among them, the arrangement condition is: being arranged side by side with the weight matrix that has been arranged on the current memory-computation integrated unit block in the current arrangement period and being able to accommodate the weight matrix of the current layer.

[0023] Further, when arranging the weight matrix, start arranging from the next column of the non-idle column.

[0024] Further, the step of sequentially polling the memory-computation integrated unit blocks in the main array in a serpentine order to find the arrangement position further includes:

[0025] If no position that meets the arrangement condition is found after polling all the memory-computation integrated unit blocks in the serpentine order, return to the first memory-computation integrated unit block and enter the next arrangement period:

[0026] Judge whether the idle position of the first memory-computation integrated unit block can accommodate the weight matrix of the current layer;

[0027] If so, arrange the weight matrix of the current layer at the idle position of the first memory-computation integrated unit block;

[0028] If not, sequentially poll each memory-computation integrated unit block in the serpentine order until a memory-computation integrated unit block that can accommodate the weight matrix is found, and arrange the weight matrix of the current layer at the idle position of the memory-computation integrated unit block;

[0029] Among them, when arranging the weight matrix, start arranging from the next row of the non-idle row.

[0030] Further, when arranging the corresponding Bias to the position corresponding to the arrangement position column in the Bias array of the memory-computation integrated chip, arrange the Bias to the next row of the non-idle row in the column corresponding position.

[0031] Further, the neural network mapping method for the memory-computation integrated chip further includes:

[0032] Expand the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array to obtain the final Bias arrangement result.

[0033] Further, the step of expanding the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array includes:

[0034] Judge whether all the occupied rows of Bias can be doubled according to the number of idle rows in the Bias array;

[0035] If so, double the occupied rows of all Bias;

[0036] If not, select Bias according to a preset rule for expansion.

[0037] Furthermore, the neural network mapping method for the in-memory computing chip further includes:

[0038] Dividing the in-memory computing unit array of the in-memory computing chip into a main array and a Bias array;

[0039] Dividing the main array into multiple in-memory computing unit blocks.

[0040] Furthermore, the neural network mapping method for the in-memory computing chip further includes:

[0041] Obtaining the parameters of the neural network to be mapped and the parameters of the target in-memory computing chip, where the parameters of the neural network to be mapped include the weight matrix and Bias corresponding to each layer;

[0042] Obtaining the minimum number of rows of Bias corresponding to each layer according to the Bias of each layer and the parameters of the in-memory computing chip.

[0043] In a second aspect, an in-memory computing chip is provided, including: an in-memory computing unit array for performing neural network operations, where the in-memory computing unit array includes: a main array and a Bias array, and the weight matrix corresponding to each layer of the neural network is mapped in the main array; the Bias corresponding to each layer of the neural network is mapped in the Bias array;

[0044] The weight matrix corresponding to each layer is sorted based on the minimum number of rows of Bias and the number of columns of the weight matrix, and is arranged in a serpentine pattern on the main array according to the sorting result, and the Bias is arranged at the position corresponding to the column of the arrangement position of the corresponding weight matrix in the Bias array.

[0045] Furthermore, the main array includes multiple in-memory computing unit blocks distributed in an array, and the weight matrix corresponding to each layer is arranged on the in-memory computing unit blocks in the main array;

[0046] Among them, for the weight matrix corresponding to each layer of the neural network, the in-memory computing unit blocks in the main array are sequentially polled in a serpentine order to find the corresponding arrangement position.

[0047] Furthermore, the principle of the sorting is:

[0048] Performing mapping sorting on each layer according to the minimum number of rows of Bias corresponding to each layer;

[0049] For the layers with the same minimum number of rows of Bias, performing mapping sorting according to the number of columns of the weight matrix.

[0050] Further, based on the arrangement position column of the corresponding weight matrix in the Bias array, the Bias arrangement method further extends the number of rows occupied by each Bias based on the minimum number of rows of Bias and the idle situation of the Bias array.

[0051] In a third aspect, a memory - in - computing chip is provided, including: a memory - in - computing unit array for performing neural network operations, where the memory - in - computing unit array includes: a main array and a Bias array, and weight matrices corresponding to each layer of the neural network are mapped in the main array; Biases corresponding to each layer of the neural network are mapped in the Bias array;

[0052] The arrangement methods of the weight matrix and the corresponding Bias are generated according to the above - mentioned neural network mapping method.

[0053] In a fourth aspect, a neural network mapping device for a memory - in - computing chip is provided, including:

[0054] A sorting module that sorts each layer according to the minimum number of rows of Bias and the weight matrix corresponding to each layer of the neural network to be mapped;

[0055] An arrangement module that, according to the mapping and sorting result, sequentially arranges the weight matrices corresponding to each layer into the main array of the memory - in - computing chip, and arranges the corresponding Bias into the position in the Bias array of the memory - in - computing chip corresponding to the arrangement position column according to the arrangement position of the weight matrix and the minimum number of rows of Bias;

[0056] Among them, when sequentially arranging the weight matrices corresponding to each layer, they are arranged in a serpentine order.

[0057] In a fifth aspect, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above - mentioned neural network mapping method are implemented.

[0058] In a sixth aspect, a computer - readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above - mentioned neural network mapping method are implemented.

[0059] The neural network mapping method, device, and equipment for the in-memory computing chip provided by the embodiments of the present invention. The method includes: performing mapping sorting on each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix; according to the mapping sorting result, arranging the weight matrix corresponding to each layer into the main array of the in-memory computing chip in sequence, and arranging the corresponding Bias into the position corresponding to the column of the arrangement position in the Bias array of the in-memory computing chip according to the arrangement position of the weight matrix and the minimum number of rows of Bias; wherein, when arranging the weight matrix corresponding to each layer in sequence, it is arranged in a snake-like order. By sorting each layer and arranging it in a snake-like manner, the in-memory computing unit is effectively utilized, and the size of the in-memory computing unit array can be reduced under the same operation scale.

[0060] In addition, in the embodiments of the present invention, the number of rows occupied by each Bias is extended based on the minimum number of rows of Bias and the idle situation of the Bias array, reducing the bias value on a single in-memory computing unit, reducing current noise, and improving operation accuracy.

[0061] To make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings

[0062] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0063] Figure 1 Shows the flowchart of the neural network mapping method for the in-memory computing chip in the embodiments of the present invention;

[0064] Figure 2 Illustrates the partitioning method of the in-memory computing unit array in the embodiments of the present invention;

[0065] Figure 3 Shows the specific steps of step S100 in the embodiments of the present invention;

[0066] Figure 4 Illustrates the process of sorting each layer of the neural network in the embodiments of the present invention;

[0067] Figure 5 Illustrates the arrangement results of the weight matrix and the corresponding Bias in the embodiments of the present invention;

[0068] Figure 6Illustrated is the expansion process of the arrangement of Bias in the embodiments of the present invention;

[0069] Figure 7 It is a structural block diagram of a neural network mapping device for a memory - computing integrated chip in the embodiments of the present invention;

[0070] Figure 8 It is a structural diagram of an electronic device in the embodiments of the present invention. Detailed implementation manners

[0071] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0072] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk memory, CD - ROM, optical memory, etc.) containing computer - usable program code.

[0073] It should be noted that the terms "including" and "having" in the description and claims of the present application and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0074] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0075] Figure 1 Illustrated is a flowchart of a neural network mapping method for a memory - computing integrated chip in the embodiments of the present invention; as Figure 1 shown, the neural network mapping method for the memory - computing integrated chip may include the following:

[0076] Step S100: Perform mapping sorting on each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix;

[0077] For a specific neural network that has been trained, the weight matrix parameters of each layer, the bias parameters, the bias size, and the minimum number of rows occupied on the chip have been obtained. First, sort the neural network layers according to the minimum number of bias rows and the number of columns in the weight matrix as the mapping arrangement order.

[0078] Step S200: according to the mapping and sorting results, the weight matrices corresponding to each layer are sequentially arranged in the main array of the storage and computing integrated chip;

[0079] When arranging the weight matrices corresponding to each layer in sequence according to the order obtained in step S100, they are arranged in a serpentine order.

[0080] Specifically, in a scenario where the storage-computing integrated chip uses the rows of storage-computing integrated units as input and the columns of storage-computing integrated units as output (the drawings in the embodiments of the present invention are all based on this premise, and those skilled in the art can understand that the rows and columns of storage-computing integrated units are only a relative concept, and columns can also be used as input and rows as output in specific applications. The principle is the same and will not be repeated here), when arranging the weight matrix, for each block arranged in the main array (the main array is divided into multiple storage-computing integrated unit blocks, and the storage-computing integrated unit blocks are arranged in an array), the first row (block The first row follows the -X direction, the second row follows the -X direction, the third row follows the -X direction, the fourth row follows the -X direction, and so on in a serpentine manner. The first row may follow the -X direction, the second row follows the X direction, the third row follows the -X direction, the fourth row follows the X direction, and so on in a serpentine manner. Of course, it may also be from bottom to top, for example, starting from the last row, the last row follows the -X direction, the second to last row follows the X direction, the third to last row follows the -X direction, the fourth to last row follows the X direction, and so on in a serpentine manner.

[0081] Among them, the storage and computing integrated unit can be a flash memory unit.

[0082] Step S300: Arrange the corresponding Bias to the position corresponding to the arrangement position column in the Bias array of the storage and computing integrated chip according to the arrangement position of the weight matrix and the minimum number of Bias rows;

[0083] Specifically, due to the particularity of the calculation of the integrated storage and computing chip, the weight matrix and its corresponding Bias need to be aligned in columns. Therefore, the arrangement position of the Bias is aligned in columns with the arrangement position of the corresponding weight matrix, and the arrangement in the row direction is arranged according to the minimum number of Bias rows.

[0084] Through the above steps S100 to S300, a layout plan is generated. Subsequently, according to the layout plan, the parameters of each layer of the neural network are written into the memory computing unit array of the memory computing integrated chip through a compilation tool. In the application inference stage, according to the layout plan and combined with the control requirements, when performing the neural network operation of a certain layer, through the row-column decoder, the rows and columns corresponding to the weight matrix and Bias of the neural network operation of this layer are selected, and the input signal of this layer of the neural network is input into the corresponding row of the weight matrix, and matrix multiplication and addition operations are performed with the weight matrix, and then superimposed with the corresponding Bias to obtain the calculation result of this layer of the neural network in the corresponding column.

[0085] By adopting the above technical solution, after sorting the parameters of each layer of the neural network and arranging them in a snake shape, they are effectively complementary. For example, sorting by the number of rows in the order of 9, 8, 7, 6, 5, 4, 3, 2, 1, after the snake-shaped arrangement, the corresponding rows with fewer rows and more rows are paired, so that the number of occupied rows in the Bios array can be reduced, in order to provide more rows for expansion, increasing the utilization efficiency of the memory computing unit array. Under the same operation scale, the required memory computing unit array can be greatly reduced, meeting the requirements of chip miniaturization. In an alternative embodiment, the neural network mapping method for the memory computing integrated chip may further include: dividing the memory computing unit array of the memory computing integrated chip into a main array and a Bias array; dividing the main array into multiple memory computing unit blocks. Among them, the division can refer to the usage scenario and the corresponding neural network scale for division, which can ensure the usage performance on the basis of effectively improving the resource utilization rate.

[0086] Specifically, as Figure 2 shown in (a) of, the actual physical architecture of the chip consists of a main array and a Bias (bias) array. The applicant found that in the actual application process, since too large current in the analog computing process will have a significant impact on the calculation result, therefore, in the embodiments of the present invention, the array can be logically partitioned, as Figure 2 shown in (b) of, the main array can be divided into 2×4 blocks, only for example, in actual application, it can be divided according to needs. The Bias array can be divided or not divided. In the embodiments of the present invention, the case of not dividing is taken as an example for illustration. For the convenience of explanation, the blocks in the main array are marked with a two-dimensional array, starting from (0,0). If the chip is small, the main array can also not be divided, but due to the generally large network scale, in most scenarios in practice, division is necessary. It should be added that this division is based on the actual performance of the chip, and the size of each block that can be divided can be the same or different. The embodiments of the present invention do not limit this, but before division, consider the size of each layer of the neural network to ensure that the layer with the largest occupied space in the neural network can also be mapped in one block.

[0087] In an alternative embodiment, the neural network mapping method may further include: writing the weight matrices and Biases of each layer of the neural network to be mapped to the in-memory computing chip according to the layout result.

[0088] Specifically, the above mapping method is executed on a toolchain, which can be understood as a program running on a terminal device, a server, or a chip programming device. An arrangement plan is generated through the above mapping method, and then the weight matrices and Biases need to be written to the in-memory computing chip according to this arrangement plan. Then, the in-memory computing chip can be installed on the corresponding device circuit board for inference applications to implement neural network operations. For example, it can be installed on a toy for speech recognition. At this time, the neural network parameters written in the chip are the parameters corresponding to the speech recognition neural network. Of course, the in-memory computing chip can also be installed on a face recognition device, and the neural network parameters written in the chip are the parameters corresponding to the image recognition neural network. Of course, the above are only examples of several chip application scenarios, and the embodiments of the present invention do not limit the application scenarios of the chip, which can be various devices and scenarios that require neural network operations.

[0089] In an alternative embodiment, referring to Figure 3 , step S100 may include the following contents:

[0090] Step S110: Perform mapping sorting on each layer according to the minimum number of rows of the Bias corresponding to each layer;

[0091] Step S120: Perform mapping sorting on the layers with the same minimum number of rows of Bias according to the number of columns of the weight matrix.

[0092] Specifically, referring to Figure 4 (a) shows the parameters of a neural network. Among them, the five matrices in the upper row are the weight matrices corresponding to five neural network layers, and the five matrices in the lower row are the Biases corresponding to five neural network layers. Taking the first weight matrix as an example, 1024*128 represents the scale of this weight matrix, which is 1024 columns and 128 rows, and the corresponding Bias is 1024*2, indicating that this Bias is 1024 columns and the minimum number of rows it occupies is 2 rows. Among them, the number of columns of the weight matrix corresponding to one layer of the neural network is the same as the number of columns of the corresponding Bias, which is determined by the operation characteristics in the in-memory computing chip.

[0093] When sorting each layer of this neural network, the layers can be sorted first according to the minimum number of rows of the Bias corresponding to each layer; referring to Figure 4 (b), the sorting is performed in descending order of the minimum number of rows of Bias. Of course, in actual applications, it can also be sorted in ascending order, with the same principle.

[0094] In addition, for the two Biases of 1024*2 and 768*2, the minimum number of rows is the same. At this time, the layers with the same minimum number of rows of the Bias are sorted locally from large to small according to the number of columns of the weight matrix, and the final sorting result is as shown in Figure 4 shown in (b) of

[0095] It should be noted that in the embodiments of the present invention, for a specific neural network that has been trained, the size of the Bias of each layer and the minimum number of rows occupied on the chip are obtained. First, the layers are sorted from large to small according to the minimum number of rows occupied by the Bias of each layer, and then the layers with the same number of rows occupied by the Bias are sorted locally again from small to large according to the number of columns. There are two purposes for doing this. The first is that in the process of network mapping, the problem of reducing the size of the Bias of each physical row should be considered first, because this can improve the computing accuracy of the chip. And the input received by the embodiments of the present invention includes the minimum number of rows occupied by the Bias of each layer. After mapping according to this number of rows, there is generally still unmapped space in the Bias row on the array. If m rows are used to carry the bias value originally set to be carried by 1 row, the designed bias value becomes 1 / m of the original on each row. The second is that in the application, the number of input channels of each layer of the network is small, while the number of output channels is large. As shown in Figure 4 , for the illustrated network layer, the input is m rows and the output is n columns. Usually, m is less than n. Therefore, in order to better ensure that the entire network can be accommodated in a limited space area, the sizes of each layer in the X direction are sorted from large to small, and the layers of the network are placed with the size in the X direction as the second priority on the premise that the Bias size is the first priority.

[0096] In an optional embodiment, the main array includes multiple memory-computation integrated unit blocks distributed in an array. Refer to Figure 2 . When arranging the weight matrices corresponding to each layer on the memory-computation integrated unit blocks in the main array, the weight matrices corresponding to each layer are arranged in a snake-like order in turn, specifically including: for the weight matrix corresponding to each layer of the neural network, polling the memory-computation integrated unit blocks in the main array in a snake-like order in turn to find the arrangement position.

[0097] Specifically, poll in a snake-like order in turn whether there is a position in the memory-computation integrated unit block in the main array that meets the arrangement condition of the current layer; where the arrangement condition is: side by side with the weight matrix that has been arranged on the current memory-computation integrated unit block in the current arrangement cycle and can accommodate the weight matrix of the current layer.

[0098] If so, arrange the weight matrix of the current layer at the position that meets the arrangement condition of the current layer;

[0099] If not, continue to poll the next memory-computation integrated unit block until a position that meets the arrangement condition is found;

[0100] It should be noted that, in order to minimize resource idleness to the greatest extent, when a position that meets the conditions is found, the weight matrix is arranged starting from the next column of the non-idle column. Of course, in special application scenarios, or for the consideration of reducing interference, a certain number of columns can also be skipped, and the specific selection is based on actual application needs.

[0101] If no position that meets the arrangement conditions is found after polling all the memory-computation integrated unit blocks in the snake-like order, the first memory-computation integrated unit block is returned to enter the next arrangement cycle.

[0102] In a new arrangement cycle, first, it is judged whether the idle position of the first memory-computation integrated unit block can accommodate the weight matrix of the current layer; if so, the weight matrix of the current layer is arranged at the idle position of the first memory-computation integrated unit block; if not, each memory-computation integrated unit block is polled in the snake-like order until a memory-computation integrated unit block that can accommodate the weight matrix is found, and the weight matrix of the current layer is arranged at the idle position of the memory-computation integrated unit block.

[0103] It should be noted that, in order to minimize resource idleness to the greatest extent, when a position that meets the conditions is found, the weight matrix is arranged starting from the next row of the non-idle row. Of course, in special application scenarios, or for the consideration of reducing interference, a certain number of rows can also be skipped, and the specific selection is based on actual application needs.

[0104] By adopting the above technical solution, after arranging the sorted weight matrix on the main array, according to the principle that the weight matrix corresponds to the Bias column, when arranging the corresponding Bias to the position in the Bias array of the memory-computation integrated chip corresponding to the arranged position column, the Bias is arranged at the next row of the non-idle row in the column corresponding position.

[0105] To enable those skilled in the art to better understand the embodiments of the present invention, the following combines Figure 5 to elaborate in detail on the mapping arrangement process.

[0106] First, take a 20-layer neural network as an example. Map each layer of the network to the memory-compute unit array of the in-memory computing chip. After sorting, record each layer of the neural network in order as 0 - 19, number the weight matrices in order as D0~D19, and number the Bias in order as B0~B19. First, arrange the weight matrices D0~D19, and then, in this example, arrange Bias B0~B19 according to the arrangement result of the weight matrices. When arranging, in the main array, the first row of memory-compute unit blocks is arranged from left to right, the second row of memory-compute unit blocks is arranged from right to left, the third row of memory-compute unit blocks is arranged from left to right, the fourth row of memory-compute unit blocks is arranged from right to left, and then return to the first row, and perform polling in the above snake-shaped cyclic order;

[0107] For example, for D0, first poll block (0,0). There is a position that can accommodate the weight matrix of the current layer. At this time, arrange D0 at (0,0); for D1, first poll block (0,0). There is a position that meets the conditions. Arrange D1 at (0,0 and side by side with D0); for D2, first poll block (0,0). The position side by side with D0 and D1 on block (0,0) cannot accommodate D2. Then, in order, poll block (0,1), and then arrange D2 at (0,1); for D3, first poll block (0,0). The position side by side with D0 and D1 on block (0,0) cannot accommodate D3. Then poll block (0,1). The position side by side with D2 on block (0,1) also cannot accommodate D3. Then, according to the snake shape, poll block (1,1), and then arrange D3 at (1,1); for D4, first poll block (0,0). The position side by side with D0 and D1 on block (0,0) cannot accommodate D4. Then poll block (0,1). The position side by side with D2 on block (0,1) also cannot accommodate D4. Then, according to the snake shape, poll block (1,1), and then arrange D4 at (1,1 and side by side with D3); for D5, first poll block (0,0). There is a position on block (0,0) side by side with D0 and D1 that can accommodate D5. At this time, arrange D5 at (0,0 and side by side with D0 and D1); and so on. After arranging D15, for block D16, by polling sequentially starting from (0,0), there is no position that meets the conditions (a position that can accommodate D16 and is side by side with the matrix before this arrangement cycle) on each block. Then jump to (0,0) and enter the next arrangement cycle.

[0108] In a new arrangement cycle, first judge whether the idle position of (0,0) can accommodate D16; if so, arrange D16 at (0,0), and when arranging, start arranging from the next row of the non-idle row. In this example, start arranging from the next row of D1. The arrangement of D17~D19 refers to D1~D5 and will not be elaborated here.

[0109] It should be noted that in an optional embodiment, D17 first polls whether the idle position at (0,0) can accommodate D17. Among them, the idle position includes polling the positions in the horizontal row where D0 is located (i.e., the horizontal row in the previous layout cycle), polling in sequence, and then polling whether there is a position on the right side of D16 in the current layout cycle that can accommodate D17, and so on. The principle of arrangement is that in one round of mapping, in an array block, the two layers of the network should be adjacent in the X direction and cannot be staggered, but in the Y direction, for each layer of the network from top to bottom, as long as it can be placed, it is placed. Therefore, every empty position on the array is within the consideration range.

[0110] In another optional embodiment, D17 first polls whether the idle position on the right side of D16 in (0,0) can accommodate D17. Among them, when polling, the position in the horizontal row where D0 is located (i.e., the horizontal row in the previous layout cycle) is not polled, and only the position in the horizontal row of the current cycle is polled. It should be noted that on a block, the matrix with a higher sorting order is in front of the matrix with a lower sorting order, and the front and back are determined according to the polling order. For example, for the first row of blocks (0,1) and (0,1), polling is performed in the order from left to right, so the left position is in front of the right position. In a block, the matrix with a higher sorting order is on the left side of the matrix with a lower sorting order. For example, D0 is on the left side of D1; for the second row of blocks (1,0) and (1,1), polling is performed in the order from right to left, so the right position is in front of the left position. In a block in the second row, the matrix with a higher sorting order is on the right side of the matrix with a lower sorting order. For example, D3 is on the right side of D4.

[0111] It should be noted that during mapping, first consider the positions of the weights of each layer on the main array, and then consider the positions of the Biases of each layer on the Bias array. The general principle of mapping is as follows: starting from block (0,0), in even-row blocks, map from left to right (the 0th row block is considered an even-row block), in odd-row blocks, map from right to left, and overall, map in a snake-like order from top to bottom. After one round, return to block (0,0) and repeat the above process until the entire neural network is mapped. (The above general principle is only for example. In the actual process, it is also possible to map from right to left in the 0th block row or overall in a snake-like order from the bottom block row of the array to the top block row, but the snake-like principle must be followed, and it must be a snake-like arrangement by row, not by column. This is the principle in the case where the row is the input and the column is the output. Those skilled in the art can understand that if the row is the output and the column is the input, the situation is reversed, and the embodiments of the present invention will not elaborate on this.) In the above steps, arrange the layers of the neural network in descending order. The snake-like arrangement basically ensures that the layers with the most and least number of Bias rows occupied correspond in the Y direction, and the layers with more and fewer Bias rows occupied also correspond in the Y direction. In this way, finally, the Biases of each layer on the Bias array are distributed upwards, or if mapped in the order from bottom to top, they are distributed downwards, providing more space for subsequent Bias expansion and higher computational accuracy for the calculation results of the chip.

[0112] The following several issues need to be noted during the mapping process of the main array. Unless there are special requirements, adjacent blocks should be seamlessly connected. The purpose of seamless connection is to make full use of the array space as much as possible in the early stage of mapping. Since each layer must be placed entirely in the array, if there are gaps between layers and these gaps are difficult to utilize, it may be difficult to place them in the later stage of mapping. In one round of mapping, each layer should start from block (0,0) and find a position in a snake-like order. If the current block position is not enough, consider the next block. In one round of mapping, if there is a layer that cannot be placed in all blocks, enter a new round of mapping. In the Bias mapping, only need to map from top to bottom on the Bias array according to the corresponding weight position and the pre-given minimum number of rows occupied. Figure 5 The numbers in parentheses after each Bias represent the number of rows it occupies. For example, B1(9) means that the number of rows occupied by B1 is 9.

[0113] In an optional embodiment, the neural network mapping method for the memory-computation integrated chip may further include: obtaining the parameters of the neural network to be mapped and the parameters of the target memory-computation integrated chip. The parameters of the neural network to be mapped include the weight matrix and Bias corresponding to each layer; and then obtaining the minimum number of rows of Bias corresponding to each layer according to the Bias of each layer and the parameters of the memory-computation integrated chip.

[0114] Specifically, after the neural network model is determined and the weight arrays and bias values of each layer are known, the minimum number of rows of Bias for each layer is calculated based on the bias values and the parameters of the target chip (the attributes of each row of Bias).

[0115] The minimum number of rows of Bias can be preset by engineers in the circuit field according to the accuracy requirements, generally based on meeting the worst accuracy requirements, or it can be calculated according to preset rules. This is not elaborated in the embodiments of the present invention.

[0116] In an optional embodiment, the neural network mapping method for the memory - in - computing chip may further include: expanding the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array to obtain the final Bias arrangement result.

[0117] Specifically, after the arrangement is completed, there may be idle rows in the Bias array. To minimize the Bias values stored in all or part of the memory - in - computing units in the Bias array, it is necessary to use the idle rows in the Bias array to perform a secondary expansion of the Bias arrangement to obtain the final arrangement scheme, and then write the neural network parameters into the memory - in - computing chip according to the final arrangement scheme.

[0118] Since the memory - in - computing chip essentially uses analog computing, the larger the bias value of each memory - in - computing unit on the Bias array, the greater the final calculated noise. The excessive noise introduced by excessive Bias will have a decisive impact on the calculation accuracy. Therefore, according to the array size, the actual number of rows of the Bias array occupied by logically one - row Bias can be expanded as much as possible. If the actual number of occupied rows is m, the size of the Bias stored in each row is 1 / m of the logical Bias size, thereby improving the calculation accuracy.

[0119] In an optional embodiment, expanding the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array includes:

[0120] Judging whether all the occupied rows of Bias can be doubled according to the number of idle rows in the Bias array;

[0121] If so, double all the occupied rows of Bias;

[0122] If not, select Bias for expansion according to preset rules.

[0123] Specifically, while mapping the weights of each layer of the neural network to the main array, the preliminary mapping of the Bias of each layer is completed on the Bias array according to the minimum number of rows occupied by the Bias of each layer given in advance. For the mapping result, see Figure 6(a) therein. Then, first expand according to the largest integer multiple that the mapped Bias rows can be expanded. For example, if the total number of rows in the Bias array is 10 rows and 3 rows have been mapped, then each row is expanded to 3 times the original. If 4 rows have been mapped, then each row is expanded to 2 times the original. For the Figure 6 array arrangement shown in (a) therein, it can be completely expanded by two times, and the result is as shown in Figure 6 (b) therein. After that, if there are still idle rows in the Bias array, select the number of mapped Bias rows equal to the idle rows from bottom to top, and expand these Bias rows into two rows. As shown in Figure 6 (c) therein. If there are still 2 idle rows, then expand the third and fourth rows from the bottom up into two rows respectively. Among them, the idle rows refer to the rows where no Bias has been mapped for the whole row, and the rows where Bias has been partially mapped are considered as mapped Bias rows.

[0124] Generally speaking, the principle of expansion is to make the best use of the resources of the Bias array as much as possible, reduce the idle rows, expand the Bias corresponding to all layers if possible, expand the Bias corresponding to some layers if not possible to expand the Bias corresponding to all layers, and expand the Bias of some rows if not possible to expand the Bias of some layers.

[0125] Among them, when expanding the Bias corresponding to some layers, it can be expanded from front to back or from back to front according to the above sorting order of each layer, or the Bias corresponding to some layers can be expanded according to the preset priority; for example, the Bias corresponding to the layer with a greater impact on the accuracy can be preferentially expanded, or the Bias corresponding to the layer with a larger Bias value corresponding to a single memory-computation unit after preliminary mapping can be expanded, and it is specifically selected according to the actual application requirements.

[0126] In summary, the embodiment of the present invention provides a method for automatically mapping neural network weights to a memory-computation integrated chip, which can be integrated into the tool chain of the memory-computation integrated chip design and provides convenient conditions for the users of the memory-computation integrated chip.

[0127] Among them, by expanding the arrangement of Bias, on the one hand, the memory-computation integrated units are fully utilized, the resource utilization rate is improved, and resource idleness is reduced. On the other hand, the Bias values stored in the memory-computation integrated units in part or all of the Bias arrays are reduced. The smaller the Bias value stored on a single memory-computation integrated unit, the smaller the conductance. Under the same voltage, the smaller the current, the smaller the current noise, and the higher the operation accuracy. In addition, after sorting the neural network parameters of each layer and arranging them in a snake shape, effective complementarity is achieved, and the utilization efficiency of the memory-computation integrated unit array is increased. Under the same operation scale, the required memory-computation integrated unit array can be greatly reduced to meet the requirements of chip miniaturization.

[0128] An embodiment of the present invention further provides a memory-computation integrated chip, including: a memory-computation integrated unit array for performing neural network operations, where the memory-computation integrated unit array includes: a main array and a Bias array, and the weight matrices corresponding to each layer of the neural network are mapped in the main array; the Bias corresponding to each layer of the neural network is mapped in the Bias array;

[0129] The weight matrices corresponding to each layer are sorted based on the minimum number of rows of Bias and the number of columns of the weight matrix, and are arranged in a snake shape on the main array according to the sorting result, and the Bias is arranged at the position corresponding to the column of the arrangement position of the corresponding weight matrix in the Bias array.

[0130] In an optional embodiment, the main array includes a plurality of memory-computation integrated unit blocks distributed in an array, and the weight matrices corresponding to each layer are arranged on the memory-computation integrated unit blocks in the main array;

[0131] Among them, for the weight matrix corresponding to each layer of the neural network, the memory-computation integrated unit blocks in the main array are sequentially polled in a snake shape to find the corresponding arrangement position.

[0132] In an optional embodiment, the sorting principle is: mapping and sorting each layer according to the minimum number of rows of Bias corresponding to each layer; for the layers with the same minimum number of rows of Bias, mapping and sorting according to the number of columns of the weight matrix.

[0133] In an optional embodiment, on the basis of arranging the Bias arrangement method at the position corresponding to the column of the arrangement position of the corresponding weight matrix in the Bias array, the number of rows occupied by each Bias is also expanded based on the minimum number of rows of Bias and the idle situation of the Bias array.

[0134] For the memory-computation integrated chip provided by the embodiment of the present invention, the arrangement methods of the weight matrix and the corresponding Bias are generated according to the above neural network mapping method.

[0135] It should be noted that the in-memory computing chip provided by the embodiments of the present invention can be applied to various electronic devices, such as: smart phones, tablet electronic devices, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), vehicle-mounted devices, smart wearable devices, toys, smart home control devices, pipeline equipment controllers, etc. Among them, the smart wearable devices may include smart glasses, smart watches, smart bracelets, etc.

[0136] Based on the same inventive concept, an embodiment of the present application also provides a neural network mapping device for an in-memory computing chip, which can be used to implement the method described in the above embodiments, as described in the following embodiments. The principle of solving problems by the neural network mapping device for the in-memory computing chip is similar to the above method. Therefore, the implementation of the neural network mapping device for the in-memory computing chip can refer to the implementation of the above method, and the repeated parts will not be described again. As used below, the term "unit" or "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0137] Figure 7 is a structural block diagram of the neural network mapping device for the in-memory computing chip in the embodiments of the present invention. The neural network mapping device for the in-memory computing chip includes: a sorting module 10, a weight arrangement module 20, and a bias arrangement module 30.

[0138] The sorting module maps and sorts each layer according to the minimum number of rows of Bias corresponding to each layer of the neural network to be mapped and the weight matrix;

[0139] The weight arrangement module sequentially arranges the weight matrices corresponding to each layer into the main array of the in-memory computing chip according to the mapping and sorting results;

[0140] The bias arrangement module arranges the corresponding Bias to the position corresponding to the column of the arrangement position in the Bias array of the in-memory computing chip according to the arrangement position of the weight matrix and the minimum number of rows of Bias;

[0141] Among them, when sequentially arranging the weight matrices corresponding to each layer, they are arranged in a snake-like order.

[0142] The devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is an electronic device. Specifically, the electronic device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0143] In a typical example, the electronic device specifically includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned neural network mapping method for the memory-computation integrated chip are implemented.

[0144] The following refers to Figure 8 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present application.

[0145] As Figure 8 shown, the electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0146] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that the computer program read from it can be installed in the storage section 608 as required.

[0147] Specifically, according to the embodiments of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present invention include a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned neural network mapping method for the memory-computation integrated chip are implemented.

[0148] In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 609, and / or installed from the removable medium 611.

[0149] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0150] For convenience of description, when describing the above device, it is divided into various units according to functions for separate description. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0151] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufacture including an instruction device that implements the functions specified in Figure 1 one or more flows and / or Figure 1 blocks.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0154] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0155] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0156] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0157] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0158] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A neural network mapping method for in-memory computing chips, characterized in that, include: Map and sort each layer according to the minimum number of Bias rows and the number of columns of the weight matrix corresponding to each layer of the neural network to be mapped; According to the mapping sorting result, the weight matrices corresponding to each layer are arranged in the main array of the storage-computing integrated chip in turn, and the corresponding Bias is arranged to the position corresponding to the arrangement position column in the Bias array of the storage-computing integrated chip according to the arrangement position of the weight matrix and the minimum number of Bias rows, wherein the arrangement position of the Bias is aligned with the arrangement position column of the corresponding weight matrix, and the arrangement method in the row direction is arranged according to the minimum number of Bias rows; Among them, the weight matrices corresponding to each layer are arranged in a serpentine order.

2. The neural network mapping method for the in-memory computing chip according to claim 1, wherein The main array includes a plurality of storage-computing integrated unit blocks distributed in an array, and the weight matrices corresponding to each layer are arranged on the storage-computing integrated unit blocks in the main array; The weight matrices corresponding to each layer are arranged in a serpentine order, including: for the weight matrix corresponding to each layer of the neural network, the storage and computing integrated unit blocks in the main array are polled in a serpentine order to find the arrangement position.

3. The neural network mapping method for the in-memory computing chip according to claim 2, wherein The mapping and sorting of each layer according to the minimum number of Bias rows and the weight matrix corresponding to each layer of the neural network to be mapped includes: Sort the mapping of each layer according to the minimum number of Bias rows corresponding to each layer; For layers with the same minimum number of Bias rows, they are mapped and sorted according to the number of columns in the weight matrix.

4. The neural network mapping method for the in-memory computing chip according to claim 2, wherein, Also includes; According to the arrangement results, the weight matrix and bias of each layer of the neural network to be mapped are written into the storage and computing integrated chip.

5. The neural network mapping method for the in-memory computing chip according to claim 2, wherein The step of sequentially polling the storage-computation integrated unit blocks in the main array in a serpentine order to find the arrangement position includes: polling the storage-computation integrated unit blocks in the main array in a serpentine order to see whether there are positions that meet the current layer arrangement conditions; If so, arrange the weight matrix of the current layer to a position that satisfies the arrangement conditions of the current layer; If not, continue to poll the next storage-computing integrated unit block until a position that meets the arrangement condition is found; Among them, the arrangement condition is: to be arranged side by side with the weight matrix that has been arranged on the current storage-computation integrated unit block in the current arrangement cycle and to be able to accommodate the weight matrix of the current layer.

6. The neural network mapping method for the in-memory computing chip according to claim 5, wherein When arranging the weight matrix, start from the next column of the non-free column.

7. The neural network mapping method for the in-memory computing chip according to claim 5, characterized in that The step of sequentially polling the storage-computation integrated unit blocks in the main array in a serpentine order to find the arrangement position further includes: If no position that meets the arrangement condition is found after polling all storage-computing integrated unit blocks in a serpentine order, the first storage-computing integrated unit block is returned to enter the next arrangement cycle: Determine whether the free space of the first storage-computation-in-one unit block can accommodate the weight matrix of the current layer; If yes, arrange the weight matrix of the current layer in the free position of the first storage-computation integrated unit block; If not, poll each storage-computation-in-one unit block in a serpentine order until a storage-computation-in-one unit block that can accommodate the weight matrix is found, and the weight matrix of the current layer is arranged in an idle position of the storage-computation-in-one unit block; Among them, when arranging the weight matrix, the arrangement starts from the next row of the non-idle row.

8. The neural network mapping method for the in-memory computing chip according to claim 1, wherein When arranging the corresponding Bias to the position corresponding to the arrangement position column in the Bias array of the memory-computation integrated chip, arrange the Bias to the next row of the non-idle row in the position corresponding to the column.

9. The neural network mapping method for the in-memory computing chip according to claim 1, wherein It further includes: Expand the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array to obtain the final Bias arrangement result.

10. The neural network mapping method for the in-memory computing chip according to claim 9, wherein The expanding the arrangement of Bias according to the Bias arrangement result and the idle situation of the Bias array includes: Judge whether all Bias occupancy rows can be doubled according to the number of idle rows in the Bias array; If so, double-expand all Bias occupancy rows; If not, select Bias for expansion according to the preset rules.

11. The neural network mapping method for the in-memory computing chip according to any one of claims 1 to 9, characterized in that, It further includes: Divide the memory-computation integrated unit array of the memory-computation integrated chip into a main array and a Bias array; Divide the main array into multiple memory-computation integrated unit blocks.

12. The neural network mapping method for the in-memory computing chip according to any one of claims 1 to 9, characterized in that, It further includes: Obtain the parameters of the neural network to be mapped and the parameters of the target memory-computation integrated chip, where the parameters of the neural network to be mapped include the weight matrices and Bias corresponding to each layer; Obtain the minimum number of rows of Bias corresponding to each layer according to the Bias of each layer and the parameters of the memory-computation integrated chip.

13. An in-memory computing chip, characterized in that, It includes: A memory-computation integrated unit array for performing neural network operations, the memory-computation integrated unit array includes: a main array and a Bias array, where the weight matrices corresponding to each layer of the neural network are mapped in the main array; the Bias corresponding to each layer of the neural network is mapped in the Bias array; The weight matrices corresponding to each layer are sorted based on the minimum number of rows of Bias and the number of columns of the weight matrix, and are arranged in a serpentine pattern on the main array according to the sorting result. The Bias is arranged to the position corresponding to the arrangement position column of the corresponding weight matrix in the Bias array, where the arrangement position of the Bias is aligned with the arrangement position column of the corresponding weight matrix, and the arrangement method in the row direction is arranged according to the minimum number of rows of Bias.

14. The in-memory computing chip according to claim 13, wherein The main array includes multiple memory-computation integrated unit blocks distributed in an array, and the weight matrices corresponding to each layer are arranged on the memory-computation integrated unit blocks in the main array; Among them, for the weight matrix corresponding to each layer of the neural network, sequentially poll the memory-computation integrated unit blocks in the main array in a serpentine order to find the corresponding arrangement position.

15. The in-memory computing chip according to claim 13, characterized in that, The principle of the sorting is: Perform mapping sorting on each layer according to the minimum number of rows of Bias corresponding to each layer; Perform mapping sorting on the layers with the same minimum number of rows of Bias according to the number of columns of the weight matrix.

16. The in-memory computing chip according to claim 13, characterized in that, Based on the above-mentioned arrangement method of Bias to the position corresponding to the arrangement position column of the corresponding weight matrix in the Bias array, the number of rows occupied by each Bias is further expanded based on the minimum number of rows of Bias and the idle situation of the Bias array.

17. An in-memory computing chip, comprising: A memory-computation integrated unit array for performing neural network operations, the memory-computation integrated unit array includes: a main array and a Bias array, where the weight matrices corresponding to each layer of the neural network are mapped in the main array; the Bias corresponding to each layer of the neural network is mapped in the Bias array; The arrangement of the weight matrix and the corresponding Bias is generated according to the neural network mapping method described in any one of claims 1 to 12.

18. A neural network mapping device for a memory-compute integrated chip, characterized in that, It includes: A sorting module that maps and sorts each layer according to the minimum number of rows of the Bias corresponding to each layer of the neural network to be mapped and the weight matrix; A weight arrangement module that sequentially arranges the weight matrix corresponding to each layer into the main array of the memory-compute integrated chip according to the mapping and sorting result; A bias arrangement module that arranges the corresponding Bias to the position corresponding to the column of the arrangement position in the Bias array of the memory-compute integrated chip according to the arrangement position of the weight matrix and the minimum number of rows of the Bias; Among them, when sequentially arranging the weight matrix corresponding to each layer, it is arranged in a serpentine order.

19. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the neural network mapping method described in any one of claims 1 to 12.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the neural network mapping method described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Digital-analog hybrid storage and calculation integrated chip and calculation device

    CN111241028A

  • Neural network weight matrix adjustment method, writing control method, and related device

    WO2021163866A1