Fine-grained degree modulus hybrid weight deployment method suitable for storage and calculation integrated equipment

By performing fine-grained digital-analog hybrid weight deployment on the weight matrix in neural network calculation, it is split into a matrix suitable for CIM computing array and digital computing core, and deployed according to energy efficiency values, the problem of inconsistent weight matrix size and CIM computing array is solved, and computing energy efficiency and space utilization are improved.

CN119990204APending Publication Date: 2025-05-13INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411310098.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In neural network calculation, the size of the weight matrix is ​​difficult to match the size of the CIM calculation array, which makes it difficult for each CIM calculation array to achieve peak computing force and peak energy efficiency during calculation, and the utilization rate of the array area is low.

Method used

A fine-grained digital-analog hybrid weight deployment method is proposed. By splitting the first weight matrix, multiple second weight matrices are obtained, and their energy efficiency values ​​on the CIM calculation array and the digital calculation core are obtained respectively, and the deployment method of the second weight matrix is ​​determined based on the energy efficiency value.

Benefits of technology

The space utilization, computing energy efficiency, and computing parallelism of the CIM computing array are improved, and the system-level energy efficiency and resource allocation freedom are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990204A_ABST
    Figure CN119990204A_ABST
Patent Text Reader

Abstract

The invention discloses a fine grain degree modulus mixed weight deployment method suitable for storage and calculation integrated equipment. The method comprises the steps of obtaining a first weight matrix; under the condition that the row number and / or the column number of the first weight matrix are / is greater than the row number and / or the column number of the CIM calculation array, splitting the first weight matrix to obtain a plurality of second weight matrixes; respectively acquiring energy efficiency values of each second weight matrix on the CIM calculation array and the digital calculation core, and recording the energy efficiency values as a first energy efficiency value and a second energy efficiency value; and determining a deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix. According to the deployment method, the first weight matrix is split, and the second weight matrix is distributed to the digital calculation core or the CIM calculation array for calculation according to the energy efficiency value of each split second weight matrix in the CIM calculation array and the digital calculation core. Therefore, the space utilization rate, the calculation energy efficiency and the parallelism degree of the CIM calculation array can be improved, and the system-level energy efficiency and the resource allocation freedom degree can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a fine-grained digital-analog hybrid weight deployment method applicable to a storage-computing integrated device, a fine-grained digital-analog hybrid weight deployment device applicable to a storage-computing integrated device, and an electronic device. Background Art

[0002] When performing neural network calculations, it is usually necessary to convert the weight values ​​of each layer into a weight matrix and deploy it to the CIM (Compute In Memory) computing array. The relevant technology usually directly deploys the weight matrix to be calculated to the CIM computing array. When the size of the weight matrix is ​​larger than the CIM computing array, the weight matrix will be evenly split and deployed to multiple CIM computing arrays. Since the size of the weight matrix is ​​difficult to be consistent with the size of the CIM computing array, it makes it difficult for each CIM computing array to achieve peak computing power and peak energy efficiency during calculation, and the utilization rate of the array area is low. Summary of the invention

[0003] The present invention aims to solve one of the technical problems in the related art at least to a certain extent. To this end, the first purpose of the present invention is to propose a fine-grained digital-analog hybrid weight deployment method suitable for storage-computing integrated devices, obtain a first weight matrix; when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, split the first weight matrix to obtain multiple second weight matrices; respectively obtain the energy efficiency value of each second weight matrix on the CIM computing array and the digital computing core, recorded as the first energy efficiency value and the second energy efficiency value; determine the deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix. The deployment method of the present invention adaptively splits and evaluates the energy efficiency of the first weight matrix according to the hardware characteristics of the CIM computing array and the digital computing core, and allocates it to the digital computing core or the CIM computing array for calculation according to the actual performance of each split second weight matrix calculation. In this way, the space utilization, computing energy efficiency, and computing parallelism of the CIM computing array can be improved, and the system-level energy efficiency and resource allocation freedom can also be improved.

[0004] The second object of the present invention is to propose a fine-grained digital-analog hybrid weight deployment device suitable for storage and computing integrated devices.

[0005] A third objective of the present invention is to provide an electronic device.

[0006] To achieve the above-mentioned objectives, the first aspect of the present invention proposes a fine-grained digital-analog hybrid weight deployment method suitable for storage and computing integrated devices, and the deployment method includes: obtaining a first weight matrix; when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, splitting the first weight matrix to obtain multiple second weight matrices; respectively obtaining the energy efficiency value of each second weight matrix on the CIM computing array and the digital computing core, recorded as the first energy efficiency value and the second energy efficiency value; determining the deployment method of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

[0007] According to one embodiment of the present invention, the deployment mode of the second weight matrix is ​​determined according to the first energy efficiency value and the second energy efficiency value of each second weight matrix, including: when the first energy efficiency value is greater than the second energy efficiency value, the second weight matrix is ​​deployed on the CIM computing array for calculation; when the first energy efficiency value is less than the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core for calculation; when the first energy efficiency value is equal to the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core or the CIM computing array for calculation.

[0008] According to one embodiment of the present invention, when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array, the first weight matrix is ​​split to obtain multiple second weight matrices, including: splitting the first weight matrix based on the number of rows and / or columns of the CIM calculation array.

[0009] According to one embodiment of the present invention, the above-mentioned deployment method also includes: when the number of rows of the first weight matrix is ​​greater than the number of rows of the CIM computing array and the number of columns of the first weight matrix is ​​less than or equal to the number of columns of the CIM computing array, the first weight matrix is ​​split based on the number of rows of the CIM computing array until the remaining number of rows of the first weight matrix is ​​less than the number of rows of the CIM computing array.

[0010] According to one embodiment of the present invention, the above-mentioned deployment method also includes: when the number of columns of the first weight matrix is ​​greater than the number of columns of the CIM computing array and the number of rows of the first weight matrix is ​​less than or equal to the number of rows of the CIM computing array, the first weight matrix is ​​split according to the number of columns of the CIM computing array until the remaining number of columns of the first weight matrix is ​​less than the number of columns of the CIM computing array.

[0011] According to one embodiment of the present invention, the above-mentioned deployment method also includes: when the number of rows and columns of the first weight matrix is ​​greater than the number of rows and columns of the CIM computing array, the first weight matrix is ​​split according to the number of rows and columns of the CIM computing array until the remaining number of rows and columns of the first weight matrix is ​​less than the number of rows and columns of the CIM computing array.

[0012] According to an embodiment of the present invention, obtaining a first energy efficiency value of a second weight matrix on a CIM computing array includes: determining the first energy efficiency value according to chip design parameters of the CIM computing array and a dimension of the second weight matrix.

[0013] According to one embodiment of the present invention, obtaining a second energy efficiency value of a second weight matrix on a digital computing core includes: determining the second energy efficiency value according to chip design parameters of the digital computing core and the dimension of the second weight matrix.

[0014] To achieve the above-mentioned objectives, the second aspect of the present invention proposes a fine-grained digital-analog hybrid weight deployment device suitable for a storage-computing integrated device, the device comprising: a first acquisition module, used to obtain a first weight matrix; a splitting module, used to split the first weight matrix to obtain multiple second weight matrices when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array; a second acquisition module, used to obtain a first energy efficiency value of the second weight matrix on the CIM computing array; a third acquisition module, used to obtain a second energy efficiency value of the second weight matrix on the digital computing core; a determination module, used to determine the deployment method of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

[0015] To achieve the above-mentioned purpose, the third aspect embodiment of the present invention proposes an electronic device, including a memory, a processor, and a fine-grained digital-analog hybrid weight deployment program suitable for storage-computing integrated devices, which is stored in the memory and can be run on the processor. When the processor executes the fine-grained digital-analog hybrid weight deployment program suitable for storage-computing integrated devices, the aforementioned fine-grained digital-analog hybrid weight deployment method suitable for storage-computing integrated devices is implemented.

[0016] According to the fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to an embodiment of the present invention, a first weight matrix is ​​obtained; when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, the first weight matrix is ​​split to obtain multiple second weight matrices; the energy efficiency value of each second weight matrix on the CIM computing array and the digital computing core is obtained respectively, recorded as the first energy efficiency value and the second energy efficiency value; the deployment mode of the second weight matrix is ​​determined according to the first energy efficiency value and the second energy efficiency value of each second weight matrix. The deployment method of the present invention adaptively splits and evaluates the energy efficiency of the first weight matrix according to the hardware characteristics of the CIM computing array and the digital computing core, and allocates it to the digital computing core or the CIM computing array for calculation according to the actual performance of each split second weight matrix calculation. In this way, the space utilization, computing energy efficiency, and computing parallelism of the CIM computing array can be improved, and the system-level energy efficiency and resource allocation freedom can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of the structure of a CIM according to some embodiments of the present invention;

[0018] Figure 2 A schematic diagram of the operation principle of a single row device of a CIM according to some embodiments of the present invention;

[0019] Figure 3 A flowchart of a fine-grained digital-analog hybrid weight deployment method applicable to a storage-computing integrated device according to some embodiments of the present invention;

[0020] Figure 4 A schematic diagram of weight deployment of a traditional deployment method according to some embodiments of the present invention and a deployment method of the present invention;

[0021] Figure 5 A schematic diagram of splitting the first weight matrix based on the number of rows of a CIM calculation array according to some embodiments of the present invention;

[0022] Figure 6 A schematic diagram of weight deployment according to some embodiments of the present invention;

[0023] Figure 7 A schematic diagram of splitting the first weight matrix based on the number of columns of the CIM calculation array according to some embodiments of the present invention;

[0024] Figure 8 A schematic diagram of weight deployment according to some other embodiments of the present invention;

[0025] Fig. 9 A schematic diagram of splitting the first weight matrix based on the number of rows and columns of a CIM calculation array according to some embodiments of the present invention;

[0026] Fig.10 A schematic diagram of weight deployment according to some other embodiments of the present invention;

[0027] Fig.11 A flowchart of a fine-grained digital-analog hybrid weight deployment method applicable to a storage-computing integrated device according to some other embodiments of the present invention;

[0028] Fig.12 A block diagram of a fine-grained digital-analog hybrid weight deployment device applicable to a storage-computing integrated device according to some embodiments of the present invention;

[0029] Fig.13 is a block diagram of an electronic device according to some embodiments of the present invention. DETAILED DESCRIPTION

[0030] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0031] The following describes in detail the fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to an embodiment of the present invention with reference to the accompanying drawings.

[0032] In some embodiments, CIM is an efficient analog computing device that can be used to implement large-scale matrix-vector multiplication and addition operations, and is generally widely used in artificial intelligence and neural network calculations. The basic computing unit of its circuit is usually a conductance or charge modulated circuit device, such as RRAM (Resistance RAM), PCM (Phase-Change Memory), MRAM (Magnetoresistive Random Access Memory), Flash (floating gate transistor), DRAM (Dynamic Random Access Memory), SRAM (Static Random-Access Memory), etc. By gating the row and column switches, the current is accumulated and sampled in the analog domain to achieve efficient matrix multiplication operations.

[0033] For example, refer to Figure 1 , taking the RRAM storage-computation integrated array for neural network reasoning (matrix multiplication) as an example for explanation, but not as a limitation of the present invention. Generally speaking, the weight information is stored in the storage-computation integrated device in the form of conductance G, and the row and column switches are selected by using WL (Word Line) and DAC (Digital-to-Analog Converter) to apply different levels of voltage on BL (Bit Line) to obtain the current I tot After collecting the current at the SL (Source Line) end and sampling it through the ADC (Analog-to-Digital Converter), multiple sets of matrix multiplication and addition operation results can be obtained in parallel.

[0034] The specific calculation process is as follows Figure 2 As shown, Figure 2 The weight values ​​of the neural network are quantized into integers and stored in the RRAM unit. For example, the weight information is stored in the form of conductance G1, G2, ..., G nThe WL controls all transistor switches in the row circuit. BL is used as the input voltage of each RRAM. The voltage value V is adjusted to correspond to different input signals. For example, the input signal can be V1, V2, ..., V n The input signal is a neuron or feature map quantized to an integer. According to Kirchhoff's current law, the current through each RRAM is V×G. All currents flowing through the RRAM will converge on the SL and be sampled as digital signals by the back-end ADC, completing a multiplication and addition operation in the analog domain. The current formula in the analog domain is as follows:

[0035] I tot =V1×G1+V2×G2+…+V n ×G n

[0036] Among them, I tot [A] represents the total current, V i(i=1,2,…,n) [V] represents the voltage applied to each memristor, G i(i=1,2,…,n) [S] represents the conductance value of each memristor, and the corresponding dimension is in square brackets.

[0037] When performing neural network inference, the input data X and weight value W need to be mapped to the input voltage V and the conductivity value G respectively. When performing neural network calculations, it is usually necessary to convert the weight value of each layer into a two-dimensional matrix and deploy it on the CIM computing array. At the same time, the computational parallelism of the CIM computing array depends on the number of rows and columns of the weight matrix, and its computational energy efficiency is linearly related to the number of rows and columns. The CIM computing array can only achieve peak computational energy efficiency when the number of rows and columns of the matrix is ​​consistent with the maximum parallelism of the array. When the size of the weight matrix is ​​larger than the CIM computing array, the weight matrix will be evenly split and deployed on multiple arrays, but the size of the weight matrix is ​​difficult to be consistent with the size of the CIM computing array, and when the weight matrix is ​​small, due to insufficient computational parallelism, the actual computational energy efficiency of the CIM computing array will be low, even lower than that of the digital computing core.

[0038] It should be noted that both digital computing cores and CIM computing arrays can perform matrix multiplication calculations. But generally speaking, CIM computing arrays have higher peak energy efficiency and peak computing power. However, the actual computing energy efficiency of CIM devices is usually proportional to their parallelism. Only when the parallelism is 100% can the peak energy efficiency and peak computing power be exerted. For most computing layers of neural networks, due to the different sizes of weights, the average parallelism is often low (for example, less than 50%), so it is difficult to exert the high computing power and high energy efficiency advantages of CIM. Although the peak energy efficiency of the digital computing core is lower than that of the CIM computing array, because it has a more flexible design space and a more sophisticated computing resource scheduling scheme, when facing weights of different sizes, it can usually make full use of its own computing power and provide more stable energy efficiency indicators than the CIM computing array. In some small-weight computing scenarios, it can even achieve a higher energy efficiency level than CIM.

[0039] Based on this, the deployment method of the present invention splits the weight matrix when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, and performs energy efficiency evaluation on each split weight matrix to obtain the energy efficiency value of each second weight matrix in the CIM computing array and the digital computing core after splitting, and distributes the split weight matrix to the digital computing core or the CIM computing array for calculation according to the energy efficiency value. In this way, the space utilization, computing energy efficiency, and computing parallelism of the CIM computing array can be improved, and the system-level energy efficiency and resource allocation freedom can also be improved.

[0040] Figure 3 Flow chart of a fine-grained digital-analog hybrid weight deployment method applicable to a storage-computing integrated device according to some embodiments of the present invention. Figure 3 The fine-grained digital-analog hybrid weight deployment method applicable to the storage-computing integrated device of the embodiment of the present application may include the following steps:

[0041] S110, obtaining a first weight matrix.

[0042] Exemplarily, when using the CIM computing array to perform neural network inference (matrix multiplication), a corresponding deep learning framework, such as PyTorch, can be used to obtain a weight matrix of a layer in the neural network as the first weight matrix.

[0043] S120, when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array, split the first weight matrix to obtain a plurality of second weight matrices.

[0044] Specifically, both the digital computing core and the CIM computing array can realize the multiplication calculation of the second weight matrix, among which the CIM computing array has higher peak energy efficiency and peak computing power. Therefore, when performing weight deployment, priority is given to deploying the first weight matrix on the CIM computing array for calculation. However, in general, the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, resulting in the first weight matrix being unable to be deployed on the CIM computing array. Therefore, it is necessary to split the first weight matrix to obtain multiple second weight matrices so that each second weight matrix can be deployed on the CIM computing array for calculation.

[0045] It should be noted that in addition to splitting the first weight matrix when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM computing array, it is also necessary to consider various reasons such as task requirements, accuracy requirements, computing delay requirements, etc. to determine whether to split the first weight matrix.

[0046] S130, respectively obtaining energy efficiency values ​​of each second weight matrix on the CIM computing array and the digital computing core, recorded as a first energy efficiency value and a second energy efficiency value.

[0047] Specifically, both the digital computing core and the CIM computing array can realize the multiplication calculation of the second weight matrix, but the computing energy efficiency of the CIM computing array is usually proportional to its parallelism. If the parallelism after the second weight matrix is ​​deployed on the CIM computing array is low, the computing energy efficiency of the CIM computing array may be lower than the computing energy efficiency of deploying the second weight matrix on the digital computing core. Therefore, before deploying the second weight matrix, it is necessary to obtain the first energy efficiency value of the second weight matrix on the CIM computing array, and at the same time obtain the second energy efficiency value of the second weight matrix on the digital computing core.

[0048] S140: Determine a deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

[0049] Specifically, after obtaining the first energy efficiency value and the second energy efficiency value of each second weight matrix, the deployment method of each second weight matrix can be determined by comparing the first energy efficiency value and the second energy efficiency value. For example, if the first energy efficiency value is relatively large, the second weight matrix is ​​deployed on the CIM computing array for calculation; if the second energy efficiency value is relatively large, the second weight matrix is ​​deployed on the digital computing core for calculation.

[0050] The deployment method of the present invention adaptively splits and evaluates the energy efficiency of the first weight matrix according to the hardware characteristics of the CIM computing array and the digital computing core, and allocates each split second weight matrix to the digital computing core or the CIM computing array for calculation according to the actual performance of the calculation. In this way, the space utilization, computing energy efficiency, and computing parallelism of the CIM computing array can be improved, and the system-level energy efficiency and resource allocation freedom can also be improved.

[0051] In some embodiments, the deployment method of the second weight matrix is ​​determined according to the first energy efficiency value and the second energy efficiency value of each second weight matrix, including: when the first energy efficiency value is greater than the second energy efficiency value, the second weight matrix is ​​deployed on the CIM computing array for calculation; when the first energy efficiency value is less than the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core for calculation; when the first energy efficiency value is equal to the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core or the CIM computing array for calculation.

[0052] Exemplarily, if the first energy efficiency value is greater than the second energy efficiency value, the second weight matrix is ​​deployed on the CIM computing array for calculation; if the first energy efficiency value is less than the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core for calculation; if the first energy efficiency value is equal to the second energy efficiency value, the second weight matrix can be deployed on the CIM computing array for calculation, or it can be deployed on the digital computing core for calculation. The specific deployment method can be determined according to actual conditions.

[0053] In some embodiments, in addition to determining the deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value, the deployment mode of the second weight matrix may also be determined in combination with the calculated delay requirement and / or accuracy requirement. For example, the deployment mode of the second weight matrix is ​​determined based on the first energy efficiency value, the second energy efficiency value, the delay requirement, and the accuracy requirement.

[0054] In some embodiments, when the number of rows and columns of the first weight matrix are both less than or equal to the number of rows and columns of the CIM computing array, it indicates that the first weight matrix can be directly deployed on the CIM computing array, and there is no need to split the first weight matrix.

[0055] It should be noted that the deployment method of the first weight matrix that does not need to be split is the same as the deployment method of each second weight matrix, which will not be described in detail here.

[0056] In some embodiments, when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array, the first weight matrix is ​​split to obtain multiple second weight matrices, including: splitting the first weight matrix based on the number of rows and / or columns of the CIM calculation array.

[0057] Specifically, generally speaking, when the number of rows and columns of the weight matrix is ​​the same as the number of rows and columns of the CIM computing array, 100% CIM array utilization can be achieved, reaching peak computing energy efficiency and peak computing power. Therefore, when splitting the first weight matrix, the first weight matrix can be split based on the number of rows of the CIM computing array, the number of columns of the CIM computing array, or the number of rows and columns of the CIM computing array to obtain a higher CIM array utilization.

[0058] For example, it is assumed that the dimension of the first weight matrix is ​​6 rows and 4 columns, and the dimension of the CIM calculation array is 4 rows and 4 columns. Figure 4 , usually an average splitting method is adopted to split the first weight matrix into two second weight matrices with dimensions of 3 rows and 4 columns (for example, second weight matrix 1 and second weight matrix 2), and respectively deploy them in two CIM computing arrays. At this time, the utilization rate of the two CIM computing arrays is about 50%, and the peak computing energy efficiency and peak computing power cannot be achieved.

[0059] The deployment method of the present invention is directed to this situation, and further reference is made to Figure 4 , split the first weight matrix in units of the number of rows of the CIM computing array, and split the first weight matrix into a second weight matrix 1 with a dimension of 4 rows and 4 columns and a second weight matrix 2 with a dimension of 2 rows and 4 columns. Then, obtain the first energy efficiency value and the second energy efficiency value of the second weight matrix 1 respectively. If the first energy efficiency value is greater than the second energy efficiency value, deploy the second weight matrix 1 on the CIM computing array. At this time, the number of rows and columns of the second weight matrix 1 is the same as the number of rows and columns of the CIM computing array, and 100% CIM array utilization can be achieved. Similarly, obtain the first energy efficiency value and the second energy efficiency value of the second weight matrix 2 respectively. If the first energy efficiency value is less than the second energy efficiency value, deploy the second weight matrix 2 on the digital computing core.

[0060] In summary, the part of the first weight matrix that is the same size as the CIM computing array is split out and deployed in one CIM computing array to achieve 100% CIM array utilization, reaching peak computing energy efficiency and peak computing power. At the same time, for the remaining weight matrices, since their energy efficiency values ​​in the CIM computing array are not as good as those of the digital computing core, the digital computing core will be used to calculate them. In this way, compared with the traditional deployment method, the deployment method of the present invention can obtain higher computing energy efficiency and higher CIM computing array utilization.

[0061] In some embodiments, the above-mentioned deployment method also includes: when the number of rows of the first weight matrix is ​​greater than the number of rows of the CIM computing array and the number of columns of the first weight matrix is ​​less than or equal to the number of columns of the CIM computing array, the first weight matrix is ​​split according to the number of rows of the CIM computing array until the remaining number of rows of the first weight matrix is ​​less than the number of rows of the CIM computing array.

[0062] For example, refer to Figure 5 , assuming that the dimension of the first weight matrix is ​​6 rows and 4 columns, and the dimension of the CIM computing array is 4 rows and 4 columns, that is, the number of rows of the first weight matrix is ​​greater than the number of rows of the CIM computing array and the number of columns of the first weight matrix is ​​less than or equal to the number of columns of the CIM computing array, at this time, the first weight matrix is ​​split into the second weight matrix 1 and the second weight matrix 2 based on the number of rows of the CIM computing array. Among them, the dimension of the second weight matrix 1 is 4 rows and 4 columns, and the dimension of the second weight matrix 2 is 2 rows and 4 columns. As the number of rows of the first weight matrix gradually increases, the utilization rate of the CIM computing array and the system-level energy efficiency change as shown in the figure. Figure 6 shown.

[0063] In some embodiments, the above-mentioned deployment method also includes: when the number of columns of the first weight matrix is ​​greater than the number of columns of the CIM computing array and the number of rows of the first weight matrix is ​​less than or equal to the number of rows of the CIM computing array, the first weight matrix is ​​split according to the number of columns of the CIM computing array until the remaining number of columns of the first weight matrix is ​​less than the number of columns of the CIM computing array.

[0064] For example, refer to Figure 7 , assuming that the dimension of the first weight matrix is ​​4 rows and 6 columns, and the dimension of the CIM computing array is 4 rows and 4 columns, that is, the number of columns of the first weight matrix is ​​greater than the number of columns of the CIM computing array and the number of rows of the first weight matrix is ​​less than or equal to the number of rows of the CIM computing array, at this time, the first weight matrix is ​​split into the second weight matrix 1 and the second weight matrix 2 based on the number of columns of the CIM computing array. Among them, the dimension of the second weight matrix 1 is 4 rows and 4 columns, and the dimension of the second weight matrix 2 is 4 rows and 2 columns. As the number of columns of the first weight matrix gradually increases, the utilization rate of the CIM computing array and the system-level energy efficiency change as shown in the figure. Figure 8 shown.

[0065] In some embodiments, the above-mentioned deployment method also includes: when the number of rows and columns of the first weight matrix is ​​greater than the number of rows and columns of the CIM computing array, the first weight matrix is ​​split according to the number of rows and columns of the CIM computing array until the remaining number of rows and columns of the first weight matrix is ​​less than the number of rows and columns of the CIM computing array.

[0066] For example, refer to Fig. 9, assuming that the dimension of the first weight matrix is ​​6 rows and 6 columns, and the dimension of the CIM calculation array is 4 rows and 4 columns, that is, the number of rows and columns of the first weight matrix is ​​greater than the number of rows and columns of the CIM calculation array. At this time, the first weight matrix is ​​split into the second weight matrix 1, the second weight matrix 2, the second weight matrix 3, and the second weight matrix 4 based on the number of rows and columns of the CIM calculation array. Among them, the dimension of the second weight matrix 1 is 4 rows and 4 columns, the dimension of the second weight matrix 2 is 2 rows and 4 columns, the dimension of the second weight matrix 3 is 4 rows and 2 columns, and the dimension of the second weight matrix 4 is 2 rows and 2 columns. When the number of rows and columns of the first weight matrix is ​​relatively large, the deployment diagram of the first weight matrix is ​​as follows Fig.10 shown.

[0067] In some embodiments, obtaining a first energy efficiency value of the second weight matrix on the CIM computing array includes: determining the first energy efficiency value according to chip design parameters of the CIM computing array and a dimension of the second weight matrix.

[0068] Specifically, the chip design parameters of the CIM computing array include but are not limited to the design parameters of DAC, the design parameters of ADC, the size of the CIM computing array, the chip energy consumption parameters, the row parallelism and column parallelism of the chip computing delay parameters, etc. When obtaining the first energy efficiency value of the second weight matrix on the CIM computing array, the chip design parameters of the CIM computing array and the dimensions of the second weight matrix can be input into the automatic tool for weight deployment to calculate and obtain the first energy efficiency value.

[0069] In some embodiments, obtaining a second energy efficiency value of a second weight matrix on a digital computing core includes: determining the second energy efficiency value according to chip design parameters of the digital computing core and a dimension of the second weight matrix.

[0070] Specifically, the chip design parameters of the digital computing core include but are not limited to the chip's PPA, the structure of adders and multipliers, and the level and structure of caches, etc. When obtaining the first energy efficiency value of the second weight matrix on the digital computing core, the chip design parameters of the digital computing core and the dimensions of the second weight matrix can be input into the automated tool for weight deployment to calculate and obtain the second energy efficiency value.

[0071] As a specific example, see Fig.11 , the fine-grained digital-analog hybrid weight deployment method applicable to the storage-computing integrated device may also include the following steps:

[0072] S401, start.

[0073] S402: Obtain a first weight matrix.

[0074] S403, determine whether the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array. If yes, execute S405; otherwise, execute S404.

[0075] S404: There is no need to split the first weight matrix.

[0076] S405 , splitting the first weight matrix into multiple second weight matrices based on the number of rows and / or columns of the CIM calculation array as units.

[0077] S406, determining whether the first energy efficiency value of the weight matrix on the CIM computing array is greater than the second energy efficiency value on the digital computing core. If yes, execute S408; otherwise, execute S407.

[0078] S407, deploying the weight matrix on the digital computing core for calculation.

[0079] S408, deploying the weight matrix on the CIM computing array for calculation.

[0080] S409, end.

[0081] In summary, the deployment method of the present invention adaptively splits and evaluates the energy efficiency of the first weight matrix according to the hardware characteristics of the CIM computing array and the digital computing core, and allocates each split second weight matrix to the digital computing core or the CIM computing array for calculation according to the actual performance of the calculation. In this way, the space utilization, computing energy efficiency, and computing parallelism of the CIM computing array can be improved, and the system-level energy efficiency and resource allocation freedom can also be improved.

[0082] Corresponding to the above embodiments, the present application also proposes a fine-grained digital-analog hybrid weight deployment device suitable for storage and computing integrated devices.

[0083] Reference Fig.12 The fine-grained digital-analog hybrid weight deployment device 200 suitable for storage and computing integrated devices includes: a first acquisition module 210, a splitting module 220, a second acquisition module 230, a third acquisition module 240 and a determination module 250.

[0084] Among them, the first acquisition module 210 is used to obtain the first weight matrix. The splitting module 220 is used to split the first weight matrix to obtain multiple second weight matrices when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array. The second acquisition module 230 is used to obtain the first energy efficiency value of the second weight matrix on the CIM calculation array. The third acquisition module 240 is used to obtain the second energy efficiency value of the second weight matrix on the digital computing core. The determination module 250 is used to determine the deployment method of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

[0085] According to one embodiment of the present invention, the determination module 250 is specifically used to deploy the second weight matrix on the CIM computing array for calculation when the first energy efficiency value is greater than the second energy efficiency value; deploy the second weight matrix on the digital computing core for calculation when the first energy efficiency value is less than the second energy efficiency value; and deploy the second weight matrix on the digital computing core or the CIM computing array for calculation when the first energy efficiency value is equal to the second energy efficiency value.

[0086] According to an embodiment of the present invention, the splitting module 220 is specifically configured to split the first weight matrix based on the number of rows and / or columns of the CIM calculation array.

[0087] According to one embodiment of the present invention, the splitting module 220 is also used to split the first weight matrix by the number of rows of the CIM calculation array, until the remaining number of rows of the first weight matrix is ​​less than the number of rows of the CIM calculation array, when the number of rows of the first weight matrix is ​​greater than the number of rows of the CIM calculation array and the number of columns of the first weight matrix is ​​less than or equal to the number of columns of the CIM calculation array.

[0088] According to one embodiment of the present invention, the splitting module 220 is also used to split the first weight matrix by the number of columns of the CIM calculation array, until the remaining number of columns of the first weight matrix is ​​less than the number of columns of the CIM calculation array, when the number of columns of the first weight matrix is ​​greater than the number of columns of the CIM calculation array and the number of rows of the first weight matrix is ​​less than or equal to the number of rows of the CIM calculation array.

[0089] According to one embodiment of the present invention, the splitting module 220 is also used to split the first weight matrix in units of the number of rows and columns of the CIM calculation array, when the number of rows and columns of the first weight matrix is ​​greater than the number of rows and columns of the CIM calculation array, until the remaining number of rows and columns of the first weight matrix is ​​less than the number of rows and columns of the CIM calculation array.

[0090] According to an embodiment of the present invention, the second acquisition module 230 is specifically configured to determine the first energy efficiency value according to the chip design parameters of the CIM calculation array and the dimension of the second weight matrix.

[0091] According to an embodiment of the present invention, the third acquisition module 240 is specifically configured to determine the second energy efficiency value according to the chip design parameters of the digital computing core and the dimension of the second weight matrix.

[0092] Corresponding to the above embodiments, the present application also proposes a computer-readable storage medium.

[0093] The computer-readable storage medium of the present application stores a fine-grained digital-analog hybrid weight deployment program applicable to a storage-computing integrated device. When the fine-grained digital-analog hybrid weight deployment program applicable to the storage-computing integrated device is executed by a processor, the aforementioned fine-grained digital-analog hybrid weight deployment method applicable to the storage-computing integrated device is implemented.

[0094] It should be pointed out that the above-mentioned explanation of the embodiments and beneficial effects of the fine-grained digital-analog hybrid weight deployment method of the storage and computing integrated device is also applicable to the computer-readable storage medium of the embodiment of the present invention. In order to avoid redundancy, it will not be elaborated here.

[0095] Corresponding to the above embodiment, the present application also proposes an electronic device.

[0096] See also Fig.13 As shown, the electronic device 300 of the present application includes a memory 310, a processor 320, and a fine-grained digital-analog hybrid weight deployment program suitable for a storage-in-one device, which is stored on the memory 310 and can be run on the processor 320. When the processor executes the fine-grained digital-analog hybrid weight deployment program suitable for the storage-in-one device, the aforementioned fine-grained digital-analog hybrid weight deployment method suitable for the storage-in-one device is implemented.

[0097] It should be pointed out that the above-mentioned explanation of the embodiments and beneficial effects of the fine-grained digital-analog hybrid weight deployment method applicable to storage and computing integrated devices is also applicable to the electronic devices of the embodiments of the present invention. To avoid redundancy, they will not be elaborated here.

[0098] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0099] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0100] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0101] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0102] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0103] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A fine-grained digital-analog hybrid weight deployment method suitable for storage and computing integrated devices, characterized in that: The deployment method includes: Obtain a first weight matrix; When the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array, splitting the first weight matrix to obtain a plurality of second weight matrices; Respectively obtain energy efficiency values ​​of each of the second weight matrices on the CIM computing array and the digital computing core, recorded as first energy efficiency values ​​and second energy efficiency values; The deployment mode of the second weight matrix is ​​determined according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

2. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 1 is characterized in that: Determining a deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix includes: When the first energy efficiency value is greater than the second energy efficiency value, deploying the second weight matrix on the CIM computing array for calculation; When the first energy efficiency value is less than the second energy efficiency value, deploying the second weight matrix on the digital computing core for calculation; When the first energy efficiency value is equal to the second energy efficiency value, the second weight matrix is ​​deployed on the digital computing core or the CIM computing array for calculation.

3. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 1 is characterized in that: When the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array, the first weight matrix is ​​split to obtain a plurality of second weight matrices, including: The first weight matrix is ​​split based on the number of rows and / or columns of the CIM calculation array.

4. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 3 is characterized in that: The deployment method further comprises: When the number of rows of the first weight matrix is ​​greater than the number of rows of the CIM calculation array and the number of columns of the first weight matrix is ​​less than or equal to the number of columns of the CIM calculation array, the first weight matrix is ​​split based on the number of rows of the CIM calculation array until the number of remaining rows of the first weight matrix is ​​less than the number of rows of the CIM calculation array.

5. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 3 is characterized in that: The deployment method further comprises: When the number of columns of the first weight matrix is ​​greater than the number of columns of the CIM calculation array and the number of rows of the first weight matrix is ​​less than or equal to the number of rows of the CIM calculation array, the first weight matrix is ​​split based on the number of columns of the CIM calculation array until the number of remaining columns of the first weight matrix is ​​less than the number of columns of the CIM calculation array.

6. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 3 is characterized in that: The deployment method further comprises: When the number of rows and columns of the first weight matrix is ​​greater than the number of rows and columns of the CIM calculation array, the first weight matrix is ​​split according to the number of rows and columns of the CIM calculation array until the remaining number of rows and columns of the first weight matrix is ​​less than the number of rows and columns of the CIM calculation array.

7. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 1 is characterized in that: Obtaining a first energy efficiency value of the second weight matrix on the CIM calculation array includes: The first energy efficiency value is determined according to the chip design parameters of the CIM computing array and the dimension of the second weight matrix.

8. The fine-grained digital-analog hybrid weight deployment method applicable to storage-computing integrated devices according to claim 1 is characterized in that: Obtaining a second energy efficiency value of the second weight matrix on the digital computing core includes: The second energy efficiency value is determined according to the chip design parameters of the digital computing core and the dimension of the second weight matrix.

9. A fine-grained digital-analog hybrid weight deployment device suitable for storage and computing integrated devices, characterized in that: The device comprises: A first acquisition module, used to acquire a first weight matrix; A splitting module, configured to split the first weight matrix to obtain a plurality of second weight matrices when the number of rows and / or columns of the first weight matrix is ​​greater than the number of rows and / or columns of the CIM calculation array; A second acquisition module, used for acquiring a first energy efficiency value of the second weight matrix on the CIM calculation array; A third acquisition module, used to acquire a second energy efficiency value of the second weight matrix on the digital computing core; A determination module is used to determine the deployment mode of the second weight matrix according to the first energy efficiency value and the second energy efficiency value of each second weight matrix.

10. An electronic device, characterized in that: It includes a memory, a processor, and a fine-grained digital-analog hybrid weight deployment program suitable for storage-computing integrated devices, which is stored in the memory and can be run on the processor. When the processor executes the fine-grained digital-analog hybrid weight deployment program suitable for storage-computing integrated devices, it implements the fine-grained digital-analog hybrid weight deployment method suitable for storage-computing integrated devices according to any one of claims 1-8.